Skip to main content
ModelScale

Coding benchmark

NL2Repo leaderboard

Every model the catalog carries a published NL2Repo value for, ranked by that value.

CategoryCoding
MeasureRepository understanding benchmark
TasksNatural language to repository tasks
DifficultySystem-level software comprehension

A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.

NL2Repo ranking

29 models with a published NL2Repo value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published NL2Repo value
RankModelProviderRepository understanding benchmark
1DeepSeek V4.1 FlashDeepSeek65.4
2DeepSeek V4 Pro 0813DeepSeek61.5
3OAOrnith-1.5-397BOrnith AI59.5
4Hy4 previewTencent58.9
5GLM-5.3Z.AI58
6GLM-5.3-FlashZ.AI56.3
7Qwen3.8 MaxAlibaba55.9
8DeepSeek V4 Flash 0731DeepSeek54.2
9DSdots3-note PreviewDots Studio49.8
10GLM-5.2Z.AI48.9
10Qwen3.8-Omni-FlashAlibaba48.9
12DAOrnith-1.0-397BDeepReinforce AI48.2
13Qwen3.8-Flash-NextAlibaba48.1
14Qwen3.7 MaxAlibaba47.2
15Seed 2.1 ProByteDance47
16OAOrnith-1.5-35B-A3BOrnith AI46.2
17Seed 2.1 TurboByteDance43.7
18Claude Opus 4.5Anthropic43.2
19Qwen 3.6 Max (preview)Alibaba42.9
20GLM-5.1Z.AI42.7
21Qwen3.8-27BAlibaba42.3
22MiniMax M3MiniMax42.13
23Qwen3.7 PlusAlibaba41.1
24MiniMax M2.7MiniMax39.8
25Qwen3.6-27BAlibaba36.2
26DAOrnith-1.0-35BDeepReinforce AI34.6
27OAOrnith-1.5-9BOrnith AI32.4
28Qwen3.6-35B-A3BAlibaba29.4
29DAOrnith-1.0-9BDeepReinforce AI27.2

Evidence key: Observed

Rows are ordered by the value MiniMax M2.7: Early Echoes of Self-Evolution published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard