Skip to main content
ModelScale

Coding benchmark

SWE-bench (Vals) leaderboard

SWE-bench, Vals AI run. Every model the catalog carries a published SWE-bench (Vals) value for, ranked by that value.

CategoryCoding
MeasureResolved rate
TasksReal repository issues by human time bucket
DifficultyFrontier coding agents

Vals AI’s independent run of the public SWE-bench issue set, reported by human time-to-fix bucket. Vals removed SWE-bench Verified from its index as saturated on 2026-05-04; this board is the standalone SWE-bench run.

SWE-bench (Vals) ranking

50 models with a published SWE-bench (Vals) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published SWE-bench, Vals AI run value
RankModelProviderResolved rate
1Claude Opus 5Anthropic97.0
2DeepSeek V4 Pro 0813DeepSeek96.4
3GPT-5.6 SolOpenAI96.2
4Grok 4.6xAI95.6
5GLM-5.3Z.AI95.4
5GPT-5.6 TerraOpenAI95.4
7Claude Fable 5Anthropic95.0
8Kimi K3Moonshot AI93.4
9GPT-5.6 LunaOpenAI93.0
10GLM-5.3-FlashZ.AI92.0
11DeepSeek V4 Flash 0731DeepSeek88.8
12Claude Opus 4.8Anthropic88.6
13Grok 4.5xAI86.6
13Muse Spark 1.2Meta86.6
15Qwen3.8-27BAlibaba86.0
16Qwen3.8 MaxAlibaba85.6
17GLM-5.2Z.AI82.8
18GPT-5.5OpenAI82.6
19TMInkling-SmallThinking Machines Lab82.2
20Claude Opus 4.7Anthropic82.0
20Muse Spark 1.1Meta82.0
22Gemini 3.7 FlashGoogle80.8
23Gemini 3.8 FlashGoogle80.0
24Claude Sonnet 5Anthropic79.6
24Gemini 3.6 FlashGoogle79.6
26Gemini 3.1 ProGoogle78.8
26Gemini 3.5 FlashGoogle78.8
28TMInklingThinking Machines Lab77.6
29Claude Sonnet 4.6Anthropic77.4
30GLM-5.1Z.AI76.4
31Kimi K2.6Moonshot AI76.2
32Gemini 3 FlashGoogle75.0
32Gemini 3.5 Flash-LiteGoogle75.0
32MiniMax M3MiniMax75.0
35MiMo-V2.5-ProXiaomi74.0
36MiniMax M2.7MiniMax73.8
37Qwen3.6 PlusAlibaba73.4
38GPT-5.4 miniOpenAI73.0
39Grok 4.20xAI72.2
40Grok 4.3xAI71.4
41MiMo-V2.5Xiaomi71.0
42GPT-5.4 nanoOpenAI69.8
43Nemotron 3 UltraNVIDIA69.0
44Qwen3.7 MaxAlibaba68.8
45Claude Haiku 4.5Anthropic66.6
46Mistral Medium 3.5 128BMistral66.4
47Ling 3.0 FlashInclusionAI65.2
48Gemini 3.1 Flash-LiteGoogle62.8
49Laguna M.1Poolside57.6
50Laguna XS.2Poolside55.2

Evidence key: Observed

Rows are ordered by the value Vals AI SWE-bench, Vals AI run leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard