Coding benchmark
SWE-bench Verified leaderboard
Software Engineering Benchmark Verified. Every model the catalog carries a published SWE-bench Verified value for, ranked by that value.
A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.
SWE-bench Verified ranking
75 models with a published SWE-bench Verified value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Code patch generation |
|---|---|---|---|
| 1 | Anthropic | 96 | |
| 2 | Anthropic | 95.5 | |
| 3 | Anthropic | 95 | |
| 4 | Anthropic | 88.6 | |
| 5 | Anthropic | 87.6 | |
| 6 | OAOrnith-1.5-397B | Ornith AI | 86 |
| 7 | Anthropic | 85.2 | |
| 8 | OpenAI | 85 | |
| 9 | DAOrnith-1.0-397B | DeepReinforce AI | 82.4 |
| 10 | Anthropic | 80.9 | |
| 11 | Anthropic | 80.84 | |
| 12 | DeepSeek | 80.6 | |
| 13 | MiniMax | 80.5 | |
| 14 | Alibaba | 80.4 | |
| 15 | TMInkling-Small | Thinking Machines Lab | 80.2 |
| 15 | Moonshot AI | 80.2 | |
| 17 | OpenAI | 80 | |
| 18 | Anthropic | 79.6 | |
| 19 | DeepSeek | 79 | |
| 19 | OAOrnith-1.5-35B-A3B | Ornith AI | 79 |
| 21 | Alibaba | 78.8 | |
| 22 | BTBTL-4 | Bad Theory Labs | 78.4 |
| 22 | DSdots3-note Preview | Dots Studio | 78.4 |
| 24 | Xiaomi | 78 | |
| 25 | Z.AI | 77.8 | |
| 26 | APApodex 1.1 | Apodex | 77.7 |
| 26 | Alibaba | 77.7 | |
| 28 | TMInkling | Thinking Machines Lab | 77.6 |
| 28 | Mistral | 77.6 | |
| 30 | Meta | 77.4 | |
| 31 | Anthropic | 77.2 | |
| 31 | Alibaba | 77.2 | |
| 33 | Moonshot AI | 76.8 | |
| 33 | Moonshot AI | 76.8 | |
| 35 | xAI | 76.7 | |
| 36 | Alibaba | 76.2 | |
| 37 | Meta | 76 | |
| 38 | DAOrnith-1.0-35B | DeepReinforce AI | 75.6 |
| 39 | Xiaomi | 74.8 | |
| 40 | Poolside | 74.6 | |
| 41 | Anthropic | 74.5 | |
| 42 | Tencent | 74.4 | |
| 43 | Z.AI | 73.8 | |
| 44 | Microsoft | 73.5 | |
| 45 | Xiaomi | 73.4 | |
| 45 | Alibaba | 73.4 | |
| 47 | Anthropic | 73.3 | |
| 48 | Anthropic | 72.7 | |
| 49 | Microsoft | 72.6 | |
| 50 | Alibaba | 72.4 | |
| 51 | Alibaba | 72 | |
| 52 | NVIDIA | 71.9 | |
| 53 | Poolside | 70.9 | |
| 54 | xAI | 70.8 | |
| 55 | OAOrnith-1.5-9B | Ornith AI | 70.6 |
| 55 | Upstage | 70.6 | |
| 57 | Upstage | 70.4 | |
| 58 | Poolside | 69.9 | |
| 59 | DAOrnith-1.0-9B | DeepReinforce AI | 69.4 |
| 60 | Alibaba | 69.2 | |
| 61 | LG AI Research | 68.2 | |
| 61 | Meituan | 68.2 | |
| 63 | 63.8 | ||
| 64 | PMTernary Bonsai 2 27B | Prism ML | 60.8 |
| 65 | IBM | 57 | |
| 66 | OpenAI | 54.6 | |
| 67 | ZYZAYA1-74B-Preview | Zyphra | 53.2 |
| 68 | NVIDIA | 52.8 | |
| 69 | OpenAI | 49.3 | |
| 70 | InclusionAI | 49.28 | |
| 71 | Anthropic | 49 | |
| 72 | IBM | 47.67 | |
| 73 | OPMiniCPM5-2B | OpenBMB | 46.4 |
| 74 | DeepSeek | 42 | |
| 75 | OpenAI | 23.6 |
Evidence key: Observed
Rows are ordered by the value SWE-bench: Can Language Models Resolve Real-World GitHub Issues? published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.