Skip to main content
ModelScale

Coding benchmark

SWE-Rebench leaderboard

Every model the catalog carries a published SWE-Rebench value for, ranked by that value.

CategoryCoding
MeasureCode patch generation
TasksFresh GitHub issues (rolling window)
DifficultyProfessional software engineering

A continuously updated software engineering benchmark by Nebius using fresh GitHub issues to avoid contamination. Models are evaluated 5 times per problem under a fixed ReAct scaffolding; the Resolved Rate (best pass@1) is reported.

SWE-Rebench ranking

13 models with a published SWE-Rebench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published SWE-Rebench value
RankModelProviderCode patch generation
1Claude Opus 4.6Anthropic65.3
2GLM-5Z.AI62.8
3GLM-5.1Z.AI62.7
4DeepSeek V3.2DeepSeek60.9
5Claude Sonnet 4.6Anthropic60.7
6Qwen3.5-27BAlibaba58.9
7GLM-4.7Z.AI58.7
8Kimi K2.5Moonshot AI58.5
9GPT-5.3 CodexOpenAI58.2
10Composer 2Cursor58
11Qwen3.5-35B-A3BAlibaba53.7
12MiniMax M2.7MiniMax51.9
13Gemma 4 31BGoogle41.6

Evidence key: Observed

Rows are ordered by the value SWE-Rebench: Contamination-Free Evaluation of Software Engineering Agents published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard