Skip to main content
ModelScale

Coding benchmark

SWE Multilingual leaderboard

Every model the catalog carries a published SWE Multilingual value for, ranked by that value.

CategoryCoding
MeasureRepository task completion
TasksMultilingual software-engineering tasks
DifficultyProfessional software engineering

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

SWE Multilingual ranking

41 models with a published SWE Multilingual value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published SWE Multilingual value
RankModelProviderRepository task completion
1Claude Opus 5Anthropic89.5
2Claude Fable 5.1Anthropic89.1
3Claude Opus 4.8Anthropic84.4
4Hy4 previewTencent82.9
5Qwen3.8-Flash-NextAlibaba81
6Qwen3.8-Omni-FlashAlibaba80.5
7Composer 2.5Cursor79.8
8OAOrnith-1.5-397BOrnith AI79.6
9DAOrnith-1.0-397BDeepReinforce AI78.9
10Laguna S 2.1Poolside78.5
11Claude Sonnet 5Anthropic78.3
11Qwen3.7 MaxAlibaba78.3
13Grok 4.5xAI78
14SWE-1.7Cognition77.8
15Claude Opus 4.5Anthropic77.5
16Kimi K2.6Moonshot AI76.7
17MiniMax M2.7MiniMax76.5
18DeepSeek V4 Pro 0813DeepSeek76.2
19Qwen3.7 PlusAlibaba75.8
20DSdots3-note PreviewDots Studio75.7
21Qwen3.6 PlusAlibaba73.8
22Composer 2Cursor73.7
23DeepSeek V4 Flash 0731DeepSeek73.3
23GLM-5Z.AI73.3
25Kimi K2.5Moonshot AI73
26Ling 3.0 FlashInclusionAI72.4
27OAOrnith-1.5-35B-A3BOrnith AI71.4
28Qwen3.6-27BAlibaba71.3
29DAOrnith-1.0-35BDeepReinforce AI69.3
30Nemotron 3 UltraNVIDIA67.7
31Qwen3.6-35B-A3BAlibaba67.2
32Laguna M.1Poolside63.1
32Laguna XS 2.1Poolside63.1
34LongCat-Flash-Lite-SparseMeituan59.33
35Laguna XS.2Poolside57.7
36OAOrnith-1.5-9BOrnith AI54.4
37DAOrnith-1.0-9BDeepReinforce AI52
38Granite 4.2 30BIBM41.89
39Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA36.47
40Granite 4.2 8BIBM30.78
41LLaDA2.2-flashInclusionAI25

Evidence key: Observed

Rows are ordered by the value MiniMax M2.7: Early Echoes of Self-Evolution published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard