Skip to main content
ModelScale

Coding benchmark

DeepSWE leaderboard

Every model the catalog carries a published DeepSWE value for, ranked by that value.

CategoryCoding
MeasurePass@1 with confidence interval, cost, time, and token metadata
Tasks113 software engineering tasks across 91 repositories and 5 languages
DifficultyLong-horizon software engineering

A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.

DeepSWE ranking

28 models with a published DeepSWE value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published DeepSWE value
RankModelProviderPass@1 with confidence interval, cost, time, and token metadata
1Muse Spark 1.3Meta75.4
2DeepSeek V4.1 FlashDeepSeek74.2
3GPT-6 AstraOpenAI74.1
4UNPareto 26.9Unbiased74.0
5Gemini 3.8 FlashGoogle73.8
6SWE-2Cognition73.0
7GPT-5.6 SolOpenAI72.7
8GPT-5.6 TerraOpenAI69.6
9Claude Opus 5Anthropic68.8
10Kimi K3Moonshot AI67.5
11Claude Fable 5.1Anthropic67.4
12GPT-5.6 LunaOpenAI67.2
13GLM-5.3Z.AI66.9
14Grok 4.6xAI65.9
15Gemini 3.7 FlashGoogle65.3
16Hy4 previewTencent64.3
17GLM-5.3-FlashZ.AI63.4
18DeepSeek V4 Pro 0813DeepSeek62.7
19Muse Spark 1.2Meta59.3
20Qwen3.8-Flash-NextAlibaba58.7
21Qwen3.8-Omni-FlashAlibaba57.8
22Qwen3.8 MaxAlibaba56.6
23OAOrnith-1.5-397BOrnith AI56.0
24DeepSeek V4 Flash 0731DeepSeek54.4
25Gemini 3.6 FlashGoogle49.0
26Qwen3.8-27BAlibaba42.2
27Laguna S 2.1Poolside40.4
28OAOrnith-1.5-35B-A3BOrnith AI22.0

Evidence key: Observed

Rows are ordered by the value DeepSWE benchmark blog published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard