StoryResearch
Harvey open-sources the Legal Agent Benchmark: 1,200+ tasks graded by lawyers
Harvey released LAB, an open benchmark of more than 1,200 agent tasks across 24 practice areas, graded by over 75,000 expert-written criteria. Test files hide planted issues to see if agents catch them. NVIDIA, OpenAI, Anthropic, Mistral and DeepMind backed the launch.
- A task counts as done only if every criterion passes.
- Anyone can run agents on it; Harvey invites labs, startups and firms to audit the rubrics.
- Head of Applied Research Niko Grupen: agents decompose a task, execute it and use data and tools.