Skip to main content
ModelScale

Coding benchmark

SWE-bench Pro leaderboard

Every model the catalog carries a published SWE-bench Pro value for, ranked by that value.

CategoryCoding
MeasureRepository task completion
Tasks1,865 repository problems
DifficultyLong-horizon professional engineering

A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

SWE-bench Pro ranking

70 models with a published SWE-bench Pro value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published SWE-bench Pro value
RankModelProviderRepository task completion
1Claude Fable 5.1Anthropic81.2
2Claude Mythos 5Anthropic80.3
3Claude Fable 5Anthropic80
4Claude Opus 5Anthropic79.2
5SASakana Fugu-UltraSakana AI73.7
6Claude Opus 4.8Anthropic69.2
7Qwen3.8 MaxAlibaba67.7
8Hy4 previewTencent65.7
9OAOrnith-1.5-397BOrnith AI65.1
10Grok 4.5xAI64.7
11GPT-5.6 SolOpenAI64.6
12Claude Opus 4.7 (Adaptive)Anthropic64.3
13GPT-5.6 TerraOpenAI63.4
14Qwen3.8-Omni-FlashAlibaba63.3
15Claude Sonnet 5Anthropic63.2
16GPT-5.6 LunaOpenAI62.7
17Qwen3.8-Flash-NextAlibaba62.5
18DAOrnith-1.0-397BDeepReinforce AI62.2
19GLM-5.2Z.AI62.1
20Qwen3.8-27BAlibaba61.7
21Muse Spark 1.1Meta61.5
22DSdots3-note PreviewDots Studio61
23Qwen3.7 MaxAlibaba60.6
24Atria Dawn PreviewShanghai Artificial Intelligence Laboratory59.6
24OAOrnith-1.5-35B-A3BOrnith AI59.6
26Laguna S 2.1Poolside59.4
27MiniMax M3MiniMax59
27SASakana FuguSakana AI59
29GPT-5.5OpenAI58.6
29Kimi K2.6Moonshot AI58.6
31GLM-5.1Z.AI58.4
32GPT-5.4OpenAI57.7
33Qwen3.7 PlusAlibaba57.6
34Qwen 3.6 Max (preview)Alibaba57.3
35MiMo-V2.5-ProXiaomi57.2
36Claude Opus 4.5Anthropic57.1
37GPT-5.3 CodexOpenAI56.8
38Ling 3.0 FlashInclusionAI56.6
38Qwen3.6 PlusAlibaba56.6
40Step 3.7 FlashStepFun56.3
41MiniMax M2.7MiniMax56.2
42MiMo-V2.5Xiaomi56.1
43TMInkling-SmallThinking Machines Lab55.9
44GPT-5.2OpenAI55.6
45DeepSeek V4 Pro 0813DeepSeek55.4
46Gemini 3.5 FlashGoogle55.1
46GLM-5Z.AI55.1
48TMInklingThinking Machines Lab54.3
49Gemini 3.5 Flash-LiteGoogle54.2
50Qwen3.6-27BAlibaba53.5
51Claude Opus 4.6Anthropic53.4
52MAI-Thinking-1Microsoft52.8
53DeepSeek V4 Flash 0731DeepSeek52.6
54Muse SparkMeta52.4
55Grok 4.20xAI51.8
56Muse Glimmer 30BMeta51.2
57Qwen3.5 397BAlibaba50.9
58Kimi K2.5Moonshot AI50.7
59DAOrnith-1.0-35BDeepReinforce AI50.4
60Qwen3.6-35B-A3BAlibaba49.5
61Laguna M.1Poolside49.2
62Laguna XS 2.1Poolside47.6
63OAOrnith-1.5-9BOrnith AI47.5
64Laguna XS.2Poolside46.3
65DAOrnith-1.0-9BDeepReinforce AI42.9
66LongCat-Flash-Lite-SparseMeituan40.63
67Granite 4.2 30BIBM33.29
68LLaDA2.2-flashInclusionAI30.1
69Granite 4.2 8BIBM19.11
70OPMiniCPM5-2BOpenBMB14.4

Evidence key: Observed

Rows are ordered by the value SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard