Coding benchmark
LiveCodeBench (Vals) leaderboard
LiveCodeBench, Vals AI run. Every model the catalog carries a published LiveCodeBench (Vals) value for, ranked by that value.
Vals AI’s independent implementation of the LiveCodeBench code-generation benchmark, run under one fixed harness across the models it tracks.
LiveCodeBench (Vals) ranking
48 models with a published LiveCodeBench (Vals) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Pass@1 accuracy |
|---|---|---|---|
| 1 | Anthropic | 90.5 | |
| 2 | Anthropic | 89.8 | |
| 3 | 89.5 | ||
| 4 | Anthropic | 89.0 | |
| 5 | 88.7 | ||
| 6 | 88.5 | ||
| 7 | xAI | 88.2 | |
| 8 | 88.1 | ||
| 9 | Alibaba | 87.9 | |
| 10 | Anthropic | 87.8 | |
| 11 | 87.6 | ||
| 12 | DeepSeek | 87.5 | |
| 13 | xAI | 87.4 | |
| 14 | DeepSeek | 87.3 | |
| 15 | Moonshot AI | 87.2 | |
| 16 | Alibaba | 87.1 | |
| 17 | Moonshot AI | 86.8 | |
| 18 | NVIDIA | 86.0 | |
| 18 | Alibaba | 86.0 | |
| 20 | OpenAI | 85.9 | |
| 20 | TMInkling-Small | Thinking Machines Lab | 85.9 |
| 20 | Meta | 85.9 | |
| 23 | 85.6 | ||
| 24 | TMInkling | Thinking Machines Lab | 85.5 |
| 25 | OpenAI | 85.3 | |
| 26 | Anthropic | 85.1 | |
| 27 | xAI | 84.5 | |
| 28 | xAI | 84.3 | |
| 29 | OpenAI | 84.0 | |
| 29 | InclusionAI | 84.0 | |
| 29 | Alibaba | 84.0 | |
| 32 | OpenAI | 82.6 | |
| 33 | Anthropic | 82.4 | |
| 34 | MiniMax | 82.2 | |
| 35 | Anthropic | 82.1 | |
| 36 | OpenAI | 81.5 | |
| 36 | Xiaomi | 81.5 | |
| 38 | Z.AI | 81.4 | |
| 38 | Xiaomi | 81.4 | |
| 40 | Z.AI | 80.5 | |
| 40 | Z.AI | 80.5 | |
| 42 | 80.1 | ||
| 43 | MiniMax | 79.9 | |
| 44 | 79.0 | ||
| 45 | Z.AI | 69.5 | |
| 46 | Poolside | 68.1 | |
| 47 | Poolside | 67.8 | |
| 48 | Anthropic | 41.2 |
Evidence key: Observed
Rows are ordered by the value Vals AI LiveCodeBench, Vals AI run leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.