Coding benchmark
AA-SciCode leaderboard
Artificial Analysis SciCode. Every model the catalog carries a published AA-SciCode value for, ranked by that value.
A display-only Artificial Analysis SciCode score.
AA-SciCode ranking
89 models with a published AA-SciCode value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Task success rate |
|---|---|---|---|
| 1 | Anthropic | 63.1 | |
| 2 | Anthropic | 61.0 | |
| 3 | Moonshot AI | 59.5 | |
| 4 | Z.AI | 59.0 | |
| 5 | Meta | 58.8 | |
| 5 | Meta | 58.8 | |
| 7 | 58.7 | ||
| 8 | Meta | 57.4 | |
| 9 | 57.2 | ||
| 10 | OpenAI | 57.1 | |
| 11 | 56.6 | ||
| 12 | OpenAI | 56.5 | |
| 12 | xAI | 56.5 | |
| 14 | Anthropic | 56.4 | |
| 15 | OpenAI | 55.8 | |
| 16 | OpenAI | 55.0 | |
| 16 | xAI | 55.0 | |
| 18 | Anthropic | 54.4 | |
| 19 | Anthropic | 54.3 | |
| 20 | 53.9 | ||
| 21 | OpenAI | 53.6 | |
| 22 | 53.4 | ||
| 23 | OpenAI | 52.1 | |
| 23 | Alibaba | 52.1 | |
| 25 | DeepSeek | 51.9 | |
| 26 | Z.AI | 51.6 | |
| 27 | Moonshot AI | 51.5 | |
| 28 | Z.AI | 51.2 | |
| 29 | DeepSeek | 51.0 | |
| 30 | Xiaomi | 50.6 | |
| 30 | Alibaba | 50.6 | |
| 32 | DeepSeek | 50.3 | |
| 33 | MiniMax | 50.1 | |
| 34 | TMInkling-Small | Thinking Machines Lab | 49.7 |
| 35 | Alibaba | 49.5 | |
| 36 | Tencent | 48.6 | |
| 36 | Tencent | 48.6 | |
| 38 | xAI | 48.3 | |
| 39 | MCQuasar 438B | Multiverse Computing | 48.1 |
| 40 | Moonshot AI | 47.8 | |
| 41 | OpenAI | 47.2 | |
| 42 | MiniMax | 47.1 | |
| 43 | TMInkling | Thinking Machines Lab | 47.0 |
| 44 | Alibaba | 46.6 | |
| 45 | 46.3 | ||
| 46 | Alibaba | 46.1 | |
| 47 | APApodex 1.1 | Apodex | 45.5 |
| 47 | APApodex 1.1 Mini | Apodex | 45.5 |
| 47 | 45.5 | ||
| 50 | Meta | 44.9 | |
| 51 | Z.AI | 44.8 | |
| 52 | Upstage | 44.6 | |
| 53 | InclusionAI | 44.2 | |
| 54 | StepFun | 43.9 | |
| 55 | Alibaba | 42.8 | |
| 56 | LG AI Research | 42.0 | |
| 56 | InclusionAI | 42.0 | |
| 56 | InclusionAI | 42.0 | |
| 59 | 41.3 | ||
| 60 | STA.X K2 | SK Telecom | 41.0 |
| 61 | Arcee AI | 40.6 | |
| 61 | Arcee AI | 40.6 | |
| 63 | NVIDIA | 40.3 | |
| 64 | Mistral | 40.2 | |
| 65 | 40.0 | ||
| 66 | Alibaba | 39.7 | |
| 67 | OpenAI | 38.9 | |
| 68 | Mistral | 38.8 | |
| 68 | Mistral | 38.8 | |
| 68 | Cohere | 38.8 | |
| 71 | Cohere | 38.5 | |
| 72 | IBM | 37.8 | |
| 73 | Mistral | 36.6 | |
| 73 | Alibaba | 36.6 | |
| 75 | NVIDIA | 36.2 | |
| 76 | DeepSeek | 35.8 | |
| 77 | OpenAI | 34.0 | |
| 78 | NVIDIA | 32.1 | |
| 79 | Meta | 31.7 | |
| 80 | IBM | 31.5 | |
| 81 | NVIDIA | 30.6 | |
| 82 | OPMiniCPM5-2B | OpenBMB | 26.3 |
| 83 | Upstage | 25.5 | |
| 84 | IBM | 25.3 | |
| 85 | InclusionAI | 24.2 | |
| 86 | 23.3 | ||
| 87 | CECeleris-1 | Celeris | 21.6 |
| 88 | Meta | 21.3 | |
| 89 | LiquidAI | 14.4 |
Evidence key: Observed
Rows are ordered by the value Artificial Analysis SciCode Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.