Coding benchmark
SWE-bench (Vals) leaderboard
SWE-bench, Vals AI run. Every model the catalog carries a published SWE-bench (Vals) value for, ranked by that value.
Vals AI’s independent run of the public SWE-bench issue set, reported by human time-to-fix bucket. Vals removed SWE-bench Verified from its index as saturated on 2026-05-04; this board is the standalone SWE-bench run.
SWE-bench (Vals) ranking
50 models with a published SWE-bench (Vals) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Resolved rate |
|---|---|---|---|
| 1 | Anthropic | 97.0 | |
| 2 | DeepSeek | 96.4 | |
| 3 | OpenAI | 96.2 | |
| 4 | xAI | 95.6 | |
| 5 | Z.AI | 95.4 | |
| 5 | OpenAI | 95.4 | |
| 7 | Anthropic | 95.0 | |
| 8 | Moonshot AI | 93.4 | |
| 9 | OpenAI | 93.0 | |
| 10 | Z.AI | 92.0 | |
| 11 | DeepSeek | 88.8 | |
| 12 | Anthropic | 88.6 | |
| 13 | xAI | 86.6 | |
| 13 | Meta | 86.6 | |
| 15 | Alibaba | 86.0 | |
| 16 | Alibaba | 85.6 | |
| 17 | Z.AI | 82.8 | |
| 18 | OpenAI | 82.6 | |
| 19 | TMInkling-Small | Thinking Machines Lab | 82.2 |
| 20 | Anthropic | 82.0 | |
| 20 | Meta | 82.0 | |
| 22 | 80.8 | ||
| 23 | 80.0 | ||
| 24 | Anthropic | 79.6 | |
| 24 | 79.6 | ||
| 26 | 78.8 | ||
| 26 | 78.8 | ||
| 28 | TMInkling | Thinking Machines Lab | 77.6 |
| 29 | Anthropic | 77.4 | |
| 30 | Z.AI | 76.4 | |
| 31 | Moonshot AI | 76.2 | |
| 32 | 75.0 | ||
| 32 | 75.0 | ||
| 32 | MiniMax | 75.0 | |
| 35 | Xiaomi | 74.0 | |
| 36 | MiniMax | 73.8 | |
| 37 | Alibaba | 73.4 | |
| 38 | OpenAI | 73.0 | |
| 39 | xAI | 72.2 | |
| 40 | xAI | 71.4 | |
| 41 | Xiaomi | 71.0 | |
| 42 | OpenAI | 69.8 | |
| 43 | NVIDIA | 69.0 | |
| 44 | Alibaba | 68.8 | |
| 45 | Anthropic | 66.6 | |
| 46 | Mistral | 66.4 | |
| 47 | InclusionAI | 65.2 | |
| 48 | 62.8 | ||
| 49 | Poolside | 57.6 | |
| 50 | Poolside | 55.2 |
Evidence key: Observed
Rows are ordered by the value Vals AI SWE-bench, Vals AI run leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.