Skip to main content
ModelScale

Agentic benchmark

GDPval-AA leaderboard

GDPval-AA normalized. Every model the catalog carries a published GDPval-AA value for, ranked by that value.

CategoryAgentic
MeasureNormalized score
TasksEconomically valuable tasks
DifficultyProfessional agentic workflows

A display-only Artificial Analysis normalized score for economically valuable tasks.

GDPval-AA ranking

108 models with a published GDPval-AA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published GDPval-AA normalized value
RankModelProviderNormalized score
1Claude Opus 5Anthropic61.8
2Claude Fable 5.1Anthropic61.2
3Muse Spark 1.3Meta60.2
4Qwen3.8 Max PreviewAlibaba58.2
5Qwen3.8-Flash-NextAlibaba57.4
6Grok 4.6xAI57.1
7GLM-5.3Z.AI56.7
8Claude Fable 5Anthropic56.6
8DeepSeek V4.1 FlashDeepSeek56.6
10DeepSeek V4 Pro 0813DeepSeek54.5
11GPT-5.6 SolOpenAI54.3
12GPT-6 AstraOpenAI54.0
13Kimi K3Moonshot AI52.4
14Muse Spark 1.2Meta51.2
15Claude Sonnet 5Anthropic50.0
16Claude Opus 4.8Anthropic49.5
17GPT-5.6 TerraOpenAI48.8
18Gemini 3.8 FlashGoogle48.2
18Qwen3.8-27BAlibaba48.2
20GPT-5.6 LunaOpenAI47.8
21DeepSeek V4 Flash 0731DeepSeek47.1
22Gemini 3.7 FlashGoogle46.7
23Grok 4.5xAI46.5
24GLM-5.2Z.AI45.3
25Claude Opus 4.7 (Adaptive)Anthropic44.8
25GPT-5.5OpenAI44.8
27Gemini 3.5 FlashGoogle42.2
28Gemini 3.6 FlashGoogle41.6
29GPT-5.4OpenAI40.3
30MiniMax M3MiniMax40.2
31Muse Spark 1.1Meta39.7
32APApodex 1.1Apodex38.7
32APApodex 1.1 MiniApodex38.7
34MCQuasar 438BMultiverse Computing37.0
35Ling 3.0 Flash VLInclusionAI36.2
36Hy3 PreviewTencent35.8
37TMInkling-SmallThinking Machines Lab34.6
38Qwen3.7 MaxAlibaba34.5
39MiMo-V2.5-ProXiaomi34.3
40GLM-5.1Z.AI34.0
41Solar Pro 4Upstage33.6
42TMInklingThinking Machines Lab33.3
43Nemotron 3 UltraNVIDIA33.1
44Hy3Tencent31.8
45Kimi K2.7 CodeMoonshot AI30.7
46GLM-4.7Z.AI29.8
46Kimi K2.6Moonshot AI29.8
48GPT-5.4 miniOpenAI29.7
49MiniMax M2.7MiniMax29.4
50Grok 4.3xAI29.2
51Muse SparkMeta28.8
52Qwen3.6-27BAlibaba28.4
53Qwen3.6 PlusAlibaba28.3
54Gemini 3.5 Flash-LiteGoogle28.2
55STA.X K2SK Telecom27.2
56GPT-5.4 nanoOpenAI26.8
57Ling 3.0 FlashInclusionAI25.9
57Ling 3.0 Flash FP8InclusionAI25.9
59Step 3.7 FlashStepFun25.8
60GPT-5 (high)OpenAI25.7
61Qwen3.6-35B-A3BAlibaba24.5
62Kimi K2.5Moonshot AI21.8
62Kimi K2.5 (Reasoning)Moonshot AI21.8
64GPT-5.1OpenAI21.5
65Qwen3.5-122B-A10BAlibaba21.3
66K-EXAONE 2.0LG AI Research20.9
67Gemini 3.1 ProGoogle20.2
68Muse Glimmer 30BMeta19.6
69Qwen3.7 PlusAlibaba19.3
70Mistral Medium 3.5 128BMistral18.8
71OPMiniCPM5-2BOpenBMB16.4
72MiMo-V2-FlashXiaomi14.4
73Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA13.3
74Gemma 4 31BGoogle12.8
75GPT-OSS 120BOpenAI12.2
76Granite 4.2 30BIBM11.1
77Ling 3.0 TinyInclusionAI10.8
78Gemma 4 26B A4BGoogle9.7
79Command A+Cohere7.9
80Granite 4.2 8BIBM7.3
81Nemotron 3 Super 100BNVIDIA7.2
82Gemini 2.5 ProGoogle5.8
83Gemma 4 12BGoogle4.5
84Mistral Large 3Mistral4.4
85K-ExaoneLG AI Research2.0
86Mistral Small 4Mistral1.8
86Mistral Small 4 (Reasoning)Mistral1.8
88GPT-OSS 20BOpenAI0.7
89Trinity-Large-PreviewArcee AI0.6
89Trinity-Large-ThinkingArcee AI0.6
91CECeleris-1Celeris0.0
91DeepSeek V3DeepSeek0.0
91Gemma 3 27BGoogle0.0
91Gemma 4 E2BGoogle0.0
91Gemma 4 E4BGoogle0.0
91GPT-4.1 miniOpenAI0.0
91GPT-4.1 nanoOpenAI0.0
91GPT-4o miniOpenAI0.0
91Granite 4.2 3BIBM0.0
91LFM2.5-2.6BLiquidAI0.0
91Ling 2.6 FlashInclusionAI0.0
91Llama 4 MaverickMeta0.0
91Llama 4 ScoutMeta0.0
91Nemotron 3 Nano 30BNVIDIA0.0
91Nemotron 3 Nano Omni 30B A3BNVIDIA0.0
91North Mini CodeCohere0.0
91Solar Pro 3Upstage0.0
91FAUltravox v0.6 Llama 3.3 70BFixie AI0.0

Evidence key: Observed

Rows are ordered by the value Artificial Analysis model benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard