Coding benchmark
NL2Repo leaderboard
Every model the catalog carries a published NL2Repo value for, ranked by that value.
A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.
NL2Repo ranking
29 models with a published NL2Repo value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Repository understanding benchmark |
|---|---|---|---|
| 1 | DeepSeek | 65.4 | |
| 2 | DeepSeek | 61.5 | |
| 3 | OAOrnith-1.5-397B | Ornith AI | 59.5 |
| 4 | Tencent | 58.9 | |
| 5 | Z.AI | 58 | |
| 6 | Z.AI | 56.3 | |
| 7 | Alibaba | 55.9 | |
| 8 | DeepSeek | 54.2 | |
| 9 | DSdots3-note Preview | Dots Studio | 49.8 |
| 10 | Z.AI | 48.9 | |
| 10 | Alibaba | 48.9 | |
| 12 | DAOrnith-1.0-397B | DeepReinforce AI | 48.2 |
| 13 | Alibaba | 48.1 | |
| 14 | Alibaba | 47.2 | |
| 15 | ByteDance | 47 | |
| 16 | OAOrnith-1.5-35B-A3B | Ornith AI | 46.2 |
| 17 | ByteDance | 43.7 | |
| 18 | Anthropic | 43.2 | |
| 19 | Alibaba | 42.9 | |
| 20 | Z.AI | 42.7 | |
| 21 | Alibaba | 42.3 | |
| 22 | MiniMax | 42.13 | |
| 23 | Alibaba | 41.1 | |
| 24 | MiniMax | 39.8 | |
| 25 | Alibaba | 36.2 | |
| 26 | DAOrnith-1.0-35B | DeepReinforce AI | 34.6 |
| 27 | OAOrnith-1.5-9B | Ornith AI | 32.4 |
| 28 | Alibaba | 29.4 | |
| 29 | DAOrnith-1.0-9B | DeepReinforce AI | 27.2 |
Evidence key: Observed
Rows are ordered by the value MiniMax M2.7: Early Echoes of Self-Evolution published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.