Skip to main content
ModelScale

Agentic benchmark

GDPval-AA leaderboard

Every model the catalog carries a published GDPval-AA value for, ranked by that value.

CategoryAgentic
MeasureElo
TasksAgentic real-world work tasks
DifficultyProfessional agentic workflows

An agentic real-world work-task evaluation reported as an Elo score in DeepSeek-V4 thinking-mode evaluations.

GDPval-AA ranking

101 models with a published GDPval-AA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published GDPval-AA value
RankModelProviderElo
1Claude Opus 5Anthropic1862
2GLM-5.3-FlashZ.AI1773
3GLM-5.3Z.AI1769
4Muse Spark 1.3Meta1754
5Claude Fable 5Anthropic1747
6GPT-5.6 SolOpenAI1735
7Claude Fable 5.1Anthropic1724
8Hy4 previewTencent1678
9Qwen3.8-Flash-NextAlibaba1648
10Grok 4.6xAI1643
11DeepSeek V4.1 FlashDeepSeek1632
12Muse Spark 1.2Meta1631
13Qwen3.8 Max PreviewAlibaba1630
14Claude Sonnet 5Anthropic1603
15Claude Opus 4.8Anthropic1593
16Atria Dawn PreviewShanghai Artificial Intelligence Laboratory1583
16GPT-5.6 TerraOpenAI1583
18GPT-5.6 LunaOpenAI1582
19GPT-6 AstraOpenAI1580
20Kimi K3Moonshot AI1548
21Gemini 3.8 FlashGoogle1545
22Gemini 3.7 FlashGoogle1525
23Qwen3.8-27BAlibaba1463
24Grok 4.5xAI1430
25Gemini 3.6 FlashGoogle1423
26GLM-5.2Z.AI1418
27Claude Opus 4.7 (Adaptive)Anthropic1396
27GPT-5.5OpenAI1396
29Muse Spark 1.1Meta1375
30Gemini 3.5 FlashGoogle1345
31GPT-5.4OpenAI1307
32DeepSeek V4 Pro 0813DeepSeek1306
33MiniMax M3MiniMax1304
34APApodex 1.1Apodex1273
34APApodex 1.1 MiniApodex1273
36MiMo-V2.5-ProXiaomi1265
37MCQuasar 438BMultiverse Computing1239
38TMInkling-SmallThinking Machines Lab1191
39Qwen3.7 MaxAlibaba1190
40DeepSeek V4 Flash 0731DeepSeek1189
41GLM-5.1Z.AI1181
42TMInklingThinking Machines Lab1165
43Gemini 3.5 Flash-LiteGoogle1139
44Hy3Tencent1136
44Hy3 PreviewTencent1136
46Kimi K2.6Moonshot AI1115
47Kimi K2.7 CodeMoonshot AI1114
48Ling 3.0 FlashInclusionAI1107
49GLM-4.7Z.AI1096
50GPT-5.4 miniOpenAI1095
51Nemotron 3 UltraNVIDIA1091
52MiniMax M2.7MiniMax1087
53Muse SparkMeta1076
54Qwen3.6-27BAlibaba1069
55Qwen3.6 PlusAlibaba1066
56Ling 3.0 Flash FP8InclusionAI1036
57GPT-5.4 nanoOpenAI1035
58Grok 4.3xAI1018
59GPT-5 (high)OpenAI1015
60Qwen3.6-35B-A3BAlibaba992
61Step 3.7 FlashStepFun954
62Kimi K2.5Moonshot AI936
62Kimi K2.5 (Reasoning)Moonshot AI936
64GPT-5.1OpenAI930
65Qwen3.5-122B-A10BAlibaba925
66Gemini 3.1 ProGoogle904
67Muse Glimmer 30BMeta893
68Qwen3.7 PlusAlibaba886
69Mistral Medium 3.5 128BMistral875
70Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA865
71MiMo-V2-FlashXiaomi788
72Gemma 4 31BGoogle755
73GPT-OSS 120BOpenAI745
74Gemma 4 26B A4BGoogle713
75Command A+Cohere658
76Granite 4.2 8BIBM646
77Nemotron 3 Super 100BNVIDIA644
78Gemini 2.5 ProGoogle616
79Gemma 4 12BGoogle591
80Mistral Large 3Mistral588
81Ling 2.6 FlashInclusionAI550
82K-ExaoneLG AI Research540
83Mistral Small 4Mistral537
83Mistral Small 4 (Reasoning)Mistral537
85GPT-OSS 20BOpenAI514
86Trinity-Large-PreviewArcee AI512
86Trinity-Large-ThinkingArcee AI512
88CECeleris-1Celeris483
89GPT-4.1 miniOpenAI453
90Nemotron 3 Nano 30BNVIDIA438
91Nemotron 3 Nano Omni 30B A3BNVIDIA416
92LFM2.5-2.6BLiquidAI202
93GPT-4o miniOpenAI188
94DeepSeek V3DeepSeek183
95Gemma 4 E4BGoogle177
96Llama 4 ScoutMeta60
97FAUltravox v0.6 Llama 3.3 70BFixie AI50
98Gemma 4 E2BGoogle36
99GPT-4.1 nanoOpenAI13
100Llama 4 MaverickMeta-45
101Gemma 3 27BGoogle-171

Evidence key: Observed

Rows are ordered by the value DeepSeek-V4 Technical Report published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard