Skip to main content
ModelScale

Agentic benchmark

AA Agentic Index leaderboard

Artificial Analysis Agentic Index. Every model the catalog carries a published AA Agentic Index value for, ranked by that value.

CategoryAgentic
MeasureAggregated model score
TasksCross-benchmark agentic index
DifficultyDisplay-only external reference

A display-only Artificial Analysis agentic index.

AA Agentic Index ranking

72 models with a published AA Agentic Index value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis Agentic Index value
RankModelProviderAggregated model score
1Claude Fable 5.1Anthropic58.0
2Claude Opus 5Anthropic56.2
3Muse Spark 1.3Meta55.7
4GLM-5.3Z.AI53.4
5Grok 4.6xAI53.4
6GPT-6 AstraOpenAI51.5
7Claude Fable 5Anthropic51.0
8Kimi K3Moonshot AI50.6
9GPT-5.6 SolOpenAI50.5
10Qwen3.8 Max PreviewAlibaba49.6
11DeepSeek V4 Pro 0813DeepSeek49.6
12Qwen3.8-27BAlibaba46.5
13Claude Sonnet 5Anthropic44.3
14Muse Spark 1.2Meta44.0
15GPT-5.6 TerraOpenAI43.7
16GPT-5.6 LunaOpenAI42.7
17Claude Opus 4.8Anthropic42.6
18Grok 4.5xAI42.1
19DeepSeek V4 Flash 0731DeepSeek41.7
20Gemini 3.8 FlashGoogle41.1
21Claude Opus 4.7 (Adaptive)Anthropic39.5
22GLM-5.2Z.AI39.4
23GPT-5.5OpenAI37.3
24Gemini 3.7 FlashGoogle36.4
25MCQuasar 438BMultiverse Computing32.7
26MiniMax M3MiniMax30.8
27Gemini 3.6 FlashGoogle30.1
28Muse Spark 1.1Meta27.5
29Gemini 3.5 FlashGoogle27.3
30Hy3Tencent25.6
30Hy3 PreviewTencent25.6
32GLM-5.1Z.AI25.2
33TMInkling-SmallThinking Machines Lab24.9
34TMInklingThinking Machines Lab24.3
35Qwen3.7 MaxAlibaba23.9
36MiMo-V2.5-ProXiaomi22.7
37Kimi K2.7 CodeMoonshot AI22.5
38Kimi K2.6Moonshot AI22.1
39Nemotron 3 UltraNVIDIA21.7
40Ling 3.0 FlashInclusionAI21.0
40Ling 3.0 Flash FP8InclusionAI21.0
42Qwen3.6-27BAlibaba20.1
43Qwen3.7 PlusAlibaba19.7
44GPT-5.4 miniOpenAI19.6
45GPT-5.4 nanoOpenAI17.7
46Grok 4.3xAI17.2
47MiniMax M2.7MiniMax16.8
48Gemini 3.5 Flash-LiteGoogle15.9
49Qwen3.6-35B-A3BAlibaba15.0
50Muse Glimmer 30BMeta10.5
51Gemini 3.1 ProGoogle10.3
52Qwen3.5-122B-A10BAlibaba9.6
53Mistral Medium 3.5 128BMistral9.3
54Gemma 4 31BGoogle6.7
55GPT-OSS 120BOpenAI6.2
56Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA6.1
57Nemotron 3 Super 100BNVIDIA4.1
58Granite 4.2 8BIBM3.7
59Command A+Cohere3.6
60Gemini 2.5 ProGoogle3.5
61Mistral Large 3Mistral2.4
62Mistral Small 4Mistral1.4
62Mistral Small 4 (Reasoning)Mistral1.4
64GPT-OSS 20BOpenAI1.4
65Trinity-Large-PreviewArcee AI1.2
65Trinity-Large-ThinkingArcee AI1.2
67Nemotron 3 Nano 30BNVIDIA1.0
68DeepSeek V3DeepSeek0.8
69CECeleris-1Celeris0.7
70Llama 4 MaverickMeta0.6
71Llama 4 ScoutMeta0.6
72Gemma 3 27BGoogle0.1

Evidence key: Observed

Rows are ordered by the value Artificial Analysis model leaderboards published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard