Skip to main content
ModelScale

Coding benchmark

FrontierSWE v2 leaderboard

Every model the catalog carries a published FrontierSWE v2 value for, ranked by that value.

CategoryCoding
MeasureFive-trial mean task score (Mean@5), 0-100
Tasks34 ultra-long-horizon engineering and research tasks
DifficultyUltra-long-horizon frontier software engineering
Published byFrontierSWE v2

A 34-task expansion of FrontierSWE for ultra-long-horizon engineering and research work that remains far from saturation.

FrontierSWE v2 ranking

12 models with a published FrontierSWE v2 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published FrontierSWE v2 value
RankModelProviderFive-trial mean task score (Mean@5), 0-100
1Claude Fable 5.1Anthropic56.3
2Claude Opus 5Anthropic52.0
3Claude Fable 5Anthropic47.0
4GPT-5.6 SolOpenAI32.2
5GLM-5.3Z.AI30.2
6Kimi K3Moonshot AI25.9
7Grok 4.6xAI25.3
8Gemini 3.7 FlashGoogle20.3
9Gemini 3.8 FlashGoogle19.6
10Qwen3.8 MaxAlibaba15.8
11Muse Spark 1.2Meta12.0
12TMInklingThinking Machines Lab4.1

Evidence key: Observed

Rows are ordered by the value FrontierSWE v2 published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard