Skip to main content
ModelScale

Coding benchmark

AA Coding Index leaderboard

Artificial Analysis Coding Index. Every model the catalog carries a published AA Coding Index value for, ranked by that value.

CategoryCoding
MeasureAggregated model score
TasksCross-benchmark coding index
DifficultyDisplay-only external reference

A display-only Artificial Analysis coding index.

AA Coding Index ranking

102 models with a published AA Coding Index value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis Coding Index value
RankModelProviderAggregated model score
1Claude Fable 5.1Anthropic81.6
2Claude Opus 5Anthropic78.0
3GPT-5.6 SolOpenAI77.4
4GPT-6 AstraOpenAI76.9
5Grok 4.6xAI76.8
6GPT-5.6 TerraOpenAI76.7
7Claude Fable 5Anthropic76.5
8Gemini 3.8 FlashGoogle76.3
9Kimi K3Moonshot AI76.2
10Gemini 3.7 FlashGoogle76.1
11Muse Spark 1.3Meta75.8
12GPT-5.5OpenAI74.9
13GLM-5.3Z.AI74.8
14Claude Opus 4.8Anthropic74.3
15Claude Opus 4.7 (Adaptive)Anthropic73.6
16Qwen3.8-Flash-NextAlibaba73.0
17Grok 4.5xAI72.5
18Muse Spark 1.2Meta72.2
19Qwen3.8 Max PreviewAlibaba71.8
20Claude Sonnet 5Anthropic71.5
21GPT-5.6 LunaOpenAI71.5
22Muse Spark 1.1Meta71.3
23GPT-5.4OpenAI71.0
24Gemini 3.5 FlashGoogle70.1
25Gemini 3.6 FlashGoogle69.2
26DeepSeek V4 Flash 0731DeepSeek69.1
27DeepSeek V4 Pro 0813DeepSeek68.8
27Gemini 3.1 ProGoogle68.8
29GLM-5.2Z.AI68.8
30Qwen3.8-27BAlibaba68.1
31Qwen3.7 MaxAlibaba66.0
32Kimi K2.6Moonshot AI61.8
33MCQuasar 438BMultiverse Computing61.2
34APApodex 1.1Apodex60.8
34APApodex 1.1 MiniApodex60.8
34Kimi K2.7 CodeMoonshot AI60.8
37MiMo-V2.5-ProXiaomi60.2
38Hy3Tencent58.8
38Hy3 PreviewTencent58.8
40Muse SparkMeta58.6
41MiniMax M3MiniMax58.6
42GPT-5.4 miniOpenAI56.1
43GPT-5.4 nanoOpenAI56.1
44Qwen3.7 PlusAlibaba55.9
45GLM-5.1Z.AI55.8
46Qwen3.6 PlusAlibaba54.5
47Qwen3.6-27BAlibaba53.7
48TMInkling-SmallThinking Machines Lab53.0
49MiniMax M2.7MiniMax52.6
50TMInklingThinking Machines Lab52.1
51Ling 3.0 FlashInclusionAI50.6
51Ling 3.0 Flash FP8InclusionAI50.6
53MiMo-V2-FlashXiaomi49.8
54GPT-5.1OpenAI49.4
55Gemini 3.5 Flash-LiteGoogle49.3
56Nemotron 3 UltraNVIDIA49.3
57Muse Glimmer 30BMeta49.0
58Mistral Medium 3.5 128BMistral46.9
59Kimi K2.5Moonshot AI46.8
59Kimi K2.5 (Reasoning)Moonshot AI46.8
61Qwen3.5-122B-A10BAlibaba45.7
62GLM-4.7Z.AI45.3
63Gemma 4 31BGoogle43.4
64Grok 4.3xAI42.3
65Qwen3.6-35B-A3BAlibaba41.9
66o1OpenAI39.7
67Step 3.7 FlashStepFun39.6
68Gemma 4 26B A4BGoogle39.3
69GPT-5 (high)OpenAI37.8
70Nemotron 3 Super 100BNVIDIA37.7
71o1-previewOpenAI34.0
72Gemini 2.5 ProGoogle33.3
73K-ExaoneLG AI Research32.1
74Gemma 4 12BGoogle31.0
75GPT-OSS 120BOpenAI30.4
76Command A+Cohere27.9
77Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA26.8
78Mistral Small 4Mistral26.6
78Mistral Small 4 (Reasoning)Mistral26.6
80Trinity-Large-PreviewArcee AI25.8
80Trinity-Large-ThinkingArcee AI25.8
82Ling 2.6 FlashInclusionAI25.3
83Gemini 1.5 ProGoogle23.6
84DeepSeek V3DeepSeek23.0
85Granite 4.2 8BIBM22.4
86GPT-4 TurboOpenAI21.5
87GPT-OSS 20BOpenAI20.7
88GPT-4.1 miniOpenAI20.2
89Mistral Large 3Mistral20.1
90Claude 3 OpusAnthropic19.5
91Llama 4 MaverickMeta16.3
92CECeleris-1Celeris14.4
93Nemotron 3 Nano 30BNVIDIA14.4
94Nemotron 3 Nano Omni 30B A3BNVIDIA13.8
95FAUltravox v0.6 Llama 3.3 70BFixie AI11.9
96GPT-4o miniOpenAI11.4
97GPT-4.1 nanoOpenAI11.1
98Gemma 3 27BGoogle10.1
99Gemma 4 E4BGoogle9.4
100Llama 4 ScoutMeta8.2
101LFM2.5-2.6BLiquidAI7.7
102Gemma 4 E2BGoogle7.2

Evidence key: Observed

Rows are ordered by the value Artificial Analysis model leaderboards published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard