Skip to main content
ModelScale

Coding benchmark

SWE-bench Verified leaderboard

Software Engineering Benchmark Verified. Every model the catalog carries a published SWE-bench Verified value for, ranked by that value.

CategoryCoding
MeasureCode patch generation
Tasks500 verified issues
DifficultyProfessional software engineering

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

SWE-bench Verified ranking

75 models with a published SWE-bench Verified value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Software Engineering Benchmark Verified value
RankModelProviderCode patch generation
1Claude Opus 5Anthropic96
2Claude Mythos 5Anthropic95.5
3Claude Fable 5Anthropic95
4Claude Opus 4.8Anthropic88.6
5Claude Opus 4.7 (Adaptive)Anthropic87.6
6OAOrnith-1.5-397BOrnith AI86
7Claude Sonnet 5Anthropic85.2
8GPT-5.3 CodexOpenAI85
9DAOrnith-1.0-397BDeepReinforce AI82.4
10Claude Opus 4.5Anthropic80.9
11Claude Opus 4.6Anthropic80.84
12DeepSeek V4 Pro 0813DeepSeek80.6
13MiniMax M3MiniMax80.5
14Qwen3.7 MaxAlibaba80.4
15TMInkling-SmallThinking Machines Lab80.2
15Kimi K2.6Moonshot AI80.2
17GPT-5.2OpenAI80
18Claude Sonnet 4.6Anthropic79.6
19DeepSeek V4 Flash 0731DeepSeek79
19OAOrnith-1.5-35B-A3BOrnith AI79
21Qwen3.6 PlusAlibaba78.8
22BTBTL-4Bad Theory Labs78.4
22DSdots3-note PreviewDots Studio78.4
24MiMo-V2-ProXiaomi78
25GLM-5Z.AI77.8
26APApodex 1.1Apodex77.7
26Qwen3.7 PlusAlibaba77.7
28TMInklingThinking Machines Lab77.6
28Mistral Medium 3.5 128BMistral77.6
30Muse SparkMeta77.4
31Claude Sonnet 4.5Anthropic77.2
31Qwen3.6-27BAlibaba77.2
33Kimi K2.5Moonshot AI76.8
33Kimi K2.5 (Reasoning)Moonshot AI76.8
35Grok 4.20xAI76.7
36Qwen3.5 397BAlibaba76.2
37Muse Glimmer 30BMeta76
38DAOrnith-1.0-35BDeepReinforce AI75.6
39MiMo-V2-OmniXiaomi74.8
40Laguna M.1Poolside74.6
41Claude 4.1 OpusAnthropic74.5
42Hy3 PreviewTencent74.4
43GLM-4.7Z.AI73.8
44MAI-Thinking-1Microsoft73.5
45MiMo-V2-FlashXiaomi73.4
45Qwen3.6-35B-A3BAlibaba73.4
47Claude Haiku 4.5Anthropic73.3
48Claude 4 SonnetAnthropic72.7
49MAI-Code-1.1-FlashMicrosoft72.6
50Qwen3.5-27BAlibaba72.4
51Qwen3.5-122B-A10BAlibaba72
52Nemotron 3 UltraNVIDIA71.9
53Laguna XS 2.1Poolside70.9
54Grok Code Fast 1xAI70.8
55OAOrnith-1.5-9BOrnith AI70.6
55Solar Pro 4Upstage70.6
57Solar Open 2Upstage70.4
58Laguna XS.2Poolside69.9
59DAOrnith-1.0-9BDeepReinforce AI69.4
60Qwen3.5-35B-A3BAlibaba69.2
61K-EXAONE 2.0LG AI Research68.2
61LongCat-Flash-Lite-SparseMeituan68.2
63Gemini 2.5 ProGoogle63.8
64PMTernary Bonsai 2 27BPrism ML60.8
65Granite 4.2 30BIBM57
66GPT-4.1OpenAI54.6
67ZYZAYA1-74B-PreviewZyphra53.2
68Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA52.8
69o3-miniOpenAI49.3
70LLaDA2.2-flashInclusionAI49.28
71Claude 3.5 SonnetAnthropic49
72Granite 4.2 8BIBM47.67
73OPMiniCPM5-2BOpenBMB46.4
74DeepSeek V3DeepSeek42
75GPT-4.1 miniOpenAI23.6

Evidence key: Observed

Rows are ordered by the value SWE-bench: Can Language Models Resolve Real-World GitHub Issues? published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard