Coding benchmark
SWE-bench Pro leaderboard
Every model the catalog carries a published SWE-bench Pro value for, ranked by that value.
A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.
SWE-bench Pro ranking
70 models with a published SWE-bench Pro value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Repository task completion |
|---|---|---|---|
| 1 | Anthropic | 81.2 | |
| 2 | Anthropic | 80.3 | |
| 3 | Anthropic | 80 | |
| 4 | Anthropic | 79.2 | |
| 5 | SASakana Fugu-Ultra | Sakana AI | 73.7 |
| 6 | Anthropic | 69.2 | |
| 7 | Alibaba | 67.7 | |
| 8 | Tencent | 65.7 | |
| 9 | OAOrnith-1.5-397B | Ornith AI | 65.1 |
| 10 | xAI | 64.7 | |
| 11 | OpenAI | 64.6 | |
| 12 | Anthropic | 64.3 | |
| 13 | OpenAI | 63.4 | |
| 14 | Alibaba | 63.3 | |
| 15 | Anthropic | 63.2 | |
| 16 | OpenAI | 62.7 | |
| 17 | Alibaba | 62.5 | |
| 18 | DAOrnith-1.0-397B | DeepReinforce AI | 62.2 |
| 19 | Z.AI | 62.1 | |
| 20 | Alibaba | 61.7 | |
| 21 | Meta | 61.5 | |
| 22 | DSdots3-note Preview | Dots Studio | 61 |
| 23 | Alibaba | 60.6 | |
| 24 | Shanghai Artificial Intelligence Laboratory | 59.6 | |
| 24 | OAOrnith-1.5-35B-A3B | Ornith AI | 59.6 |
| 26 | Poolside | 59.4 | |
| 27 | MiniMax | 59 | |
| 27 | SASakana Fugu | Sakana AI | 59 |
| 29 | OpenAI | 58.6 | |
| 29 | Moonshot AI | 58.6 | |
| 31 | Z.AI | 58.4 | |
| 32 | OpenAI | 57.7 | |
| 33 | Alibaba | 57.6 | |
| 34 | Alibaba | 57.3 | |
| 35 | Xiaomi | 57.2 | |
| 36 | Anthropic | 57.1 | |
| 37 | OpenAI | 56.8 | |
| 38 | InclusionAI | 56.6 | |
| 38 | Alibaba | 56.6 | |
| 40 | StepFun | 56.3 | |
| 41 | MiniMax | 56.2 | |
| 42 | Xiaomi | 56.1 | |
| 43 | TMInkling-Small | Thinking Machines Lab | 55.9 |
| 44 | OpenAI | 55.6 | |
| 45 | DeepSeek | 55.4 | |
| 46 | 55.1 | ||
| 46 | Z.AI | 55.1 | |
| 48 | TMInkling | Thinking Machines Lab | 54.3 |
| 49 | 54.2 | ||
| 50 | Alibaba | 53.5 | |
| 51 | Anthropic | 53.4 | |
| 52 | Microsoft | 52.8 | |
| 53 | DeepSeek | 52.6 | |
| 54 | Meta | 52.4 | |
| 55 | xAI | 51.8 | |
| 56 | Meta | 51.2 | |
| 57 | Alibaba | 50.9 | |
| 58 | Moonshot AI | 50.7 | |
| 59 | DAOrnith-1.0-35B | DeepReinforce AI | 50.4 |
| 60 | Alibaba | 49.5 | |
| 61 | Poolside | 49.2 | |
| 62 | Poolside | 47.6 | |
| 63 | OAOrnith-1.5-9B | Ornith AI | 47.5 |
| 64 | Poolside | 46.3 | |
| 65 | DAOrnith-1.0-9B | DeepReinforce AI | 42.9 |
| 66 | Meituan | 40.63 | |
| 67 | IBM | 33.29 | |
| 68 | InclusionAI | 30.1 | |
| 69 | IBM | 19.11 | |
| 70 | OPMiniCPM5-2B | OpenBMB | 14.4 |
Evidence key: Observed
Rows are ordered by the value SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.