Skip to main content
ModelScale

Coding benchmark

AA-SciCode leaderboard

Artificial Analysis SciCode. Every model the catalog carries a published AA-SciCode value for, ranked by that value.

CategoryCoding
MeasureTask success rate
TasksScientific coding subproblems
DifficultyScientific programming

A display-only Artificial Analysis SciCode score.

AA-SciCode ranking

89 models with a published AA-SciCode value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis SciCode value
RankModelProviderTask success rate
1Claude Fable 5.1Anthropic63.1
2Claude Fable 5Anthropic61.0
3Kimi K3Moonshot AI59.5
4GLM-5.3Z.AI59.0
5Muse Spark 1.1Meta58.8
5Muse Spark 1.3Meta58.8
7Gemini 3.1 ProGoogle58.7
8Muse Spark 1.2Meta57.4
9Gemini 3.7 FlashGoogle57.2
10GPT-5.6 SolOpenAI57.1
11Gemini 3.8 FlashGoogle56.6
12GPT-6 AstraOpenAI56.5
12Grok 4.6xAI56.5
14Claude Opus 5Anthropic56.4
15GPT-5.5OpenAI55.8
16GPT-5.6 TerraOpenAI55.0
16Grok 4.5xAI55.0
18Claude Opus 4.8Anthropic54.4
19Claude Sonnet 5Anthropic54.3
20Gemini 3.5 FlashGoogle53.9
21GPT-5.6 LunaOpenAI53.6
22Gemini 3.6 FlashGoogle53.4
23GPT-5.4 miniOpenAI52.1
23Qwen3.8 Max PreviewAlibaba52.1
25DeepSeek V4.1 FlashDeepSeek51.9
26GLM-5.3-FlashZ.AI51.6
27Kimi K2.6Moonshot AI51.5
28GLM-5.2Z.AI51.2
29DeepSeek V4 Pro 0813DeepSeek51.0
30MiMo-V2.5-ProXiaomi50.6
30Qwen3.8-Flash-NextAlibaba50.6
32DeepSeek V4 Flash 0731DeepSeek50.3
33MiniMax M2.7MiniMax50.1
34TMInkling-SmallThinking Machines Lab49.7
35Qwen3.7 MaxAlibaba49.5
36Hy3Tencent48.6
36Hy3 PreviewTencent48.6
38Grok 4.3xAI48.3
39MCQuasar 438BMultiverse Computing48.1
40Kimi K2.7 CodeMoonshot AI47.8
41GPT-5.4 nanoOpenAI47.2
42MiniMax M3MiniMax47.1
43TMInklingThinking Machines Lab47.0
44Qwen3.8-27BAlibaba46.6
45Gemini 2.5 ProGoogle46.3
46Qwen3.7 PlusAlibaba46.1
47APApodex 1.1Apodex45.5
47APApodex 1.1 MiniApodex45.5
47Gemma 4 31BGoogle45.5
50Muse Glimmer 30BMeta44.9
51GLM-5.1Z.AI44.8
52Solar Pro 4Upstage44.6
53Ling 3.0 Flash VLInclusionAI44.2
54Step 3.7 FlashStepFun43.9
55Qwen3.6-27BAlibaba42.8
56K-EXAONE 2.0LG AI Research42.0
56Ling 3.0 FlashInclusionAI42.0
56Ling 3.0 Flash FP8InclusionAI42.0
59Gemini 3.5 Flash-LiteGoogle41.3
60STA.X K2SK Telecom41.0
61Trinity-Large-PreviewArcee AI40.6
61Trinity-Large-ThinkingArcee AI40.6
63Nemotron 3 UltraNVIDIA40.3
64Mistral Medium 3.5 128BMistral40.2
65Gemma 4 26B A4BGoogle40.0
66Qwen3.5-122B-A10BAlibaba39.7
67GPT-OSS 20BOpenAI38.9
68Mistral Small 4Mistral38.8
68Mistral Small 4 (Reasoning)Mistral38.8
68North Mini CodeCohere38.8
71Command A+Cohere38.5
72Granite 4.2 30BIBM37.8
73Mistral Large 3Mistral36.6
73Qwen3.6-35B-A3BAlibaba36.6
75Nemotron 3 Super 100BNVIDIA36.2
76DeepSeek V3DeepSeek35.8
77GPT-OSS 120BOpenAI34.0
78Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA32.1
79Llama 4 MaverickMeta31.7
80Granite 4.2 8BIBM31.5
81Nemotron 3 Nano 30BNVIDIA30.6
82OPMiniCPM5-2BOpenBMB26.3
83Solar Pro 3Upstage25.5
84Granite 4.2 3BIBM25.3
85Ling 3.0 TinyInclusionAI24.2
86Gemma 3 27BGoogle23.3
87CECeleris-1Celeris21.6
88Llama 4 ScoutMeta21.3
89LFM2.5-2.6BLiquidAI14.4

Evidence key: Observed

Rows are ordered by the value Artificial Analysis SciCode Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard