StoryResearch

Harvey open-sources the Legal Agent Benchmark: 1,200+ tasks graded by lawyers

HarveyHarvey agents

Harvey released LAB, an open benchmark of more than 1,200 agent tasks across 24 practice areas, graded by over 75,000 expert-written criteria. Test files hide planted issues to see if agents catch them. NVIDIA, OpenAI, Anthropic, Mistral and DeepMind backed the launch.

  • A task counts as done only if every criterion passes.
  • Anyone can run agents on it; Harvey invites labs, startups and firms to audit the rubrics.
  • Head of Applied Research Niko Grupen: agents decompose a task, execute it and use data and tools.
Read the original · Press

More on Harvey agents

Primer