Coding benchmark
DeepSWE leaderboard
Every model the catalog carries a published DeepSWE value for, ranked by that value.
A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.
DeepSWE ranking
28 models with a published DeepSWE value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Pass@1 with confidence interval, cost, time, and token metadata |
|---|---|---|---|
| 1 | Meta | 75.4 | |
| 2 | DeepSeek | 74.2 | |
| 3 | OpenAI | 74.1 | |
| 4 | UNPareto 26.9 | Unbiased | 74.0 |
| 5 | 73.8 | ||
| 6 | Cognition | 73.0 | |
| 7 | OpenAI | 72.7 | |
| 8 | OpenAI | 69.6 | |
| 9 | Anthropic | 68.8 | |
| 10 | Moonshot AI | 67.5 | |
| 11 | Anthropic | 67.4 | |
| 12 | OpenAI | 67.2 | |
| 13 | Z.AI | 66.9 | |
| 14 | xAI | 65.9 | |
| 15 | 65.3 | ||
| 16 | Tencent | 64.3 | |
| 17 | Z.AI | 63.4 | |
| 18 | DeepSeek | 62.7 | |
| 19 | Meta | 59.3 | |
| 20 | Alibaba | 58.7 | |
| 21 | Alibaba | 57.8 | |
| 22 | Alibaba | 56.6 | |
| 23 | OAOrnith-1.5-397B | Ornith AI | 56.0 |
| 24 | DeepSeek | 54.4 | |
| 25 | 49.0 | ||
| 26 | Alibaba | 42.2 | |
| 27 | Poolside | 40.4 | |
| 28 | OAOrnith-1.5-35B-A3B | Ornith AI | 22.0 |
Evidence key: Observed
Rows are ordered by the value DeepSWE benchmark blog published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.