Skip to main content
ModelScale

Agentic benchmark

AA Harvey LAB leaderboard

Artificial Analysis Harvey LAB-AA. Every model the catalog carries a published AA Harvey LAB value for, ranked by that value.

CategoryAgentic
MeasureTask success rate
TasksLegal agent tasks
DifficultyProfessional legal work

An independently evaluated legal-agent benchmark from Artificial Analysis.

AA Harvey LAB ranking

13 models with a published AA Harvey LAB value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis Harvey LAB-AA value
RankModelProviderTask success rate
1Kimi K3Moonshot AI94.6
2Claude Fable 5Anthropic93.6
3Claude Opus 5Anthropic93.5
4Claude Fable 5.1Anthropic93.0
5Grok 4.5xAI92.4
6Gemini 3.7 FlashGoogle90.7
7MiniMax M3MiniMax88.4
8GPT-5.6 LunaOpenAI87.9
9GPT-5.6 SolOpenAI87.2
10GPT-5.6 TerraOpenAI85.2
11Nemotron 3 UltraNVIDIA81.7
12Mistral Medium 3.5 128BMistral69.1
13GPT-OSS 120BOpenAI13.9

Evidence key: Observed

Rows are ordered by the value Artificial Analysis Harvey LAB-AA Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard