Coding benchmark
LiveCodeBench leaderboard
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. Every model the catalog carries a published LiveCodeBench value for, ranked by that value.
A continuously updated coding benchmark built from newly collected LeetCode, AtCoder, and Codeforces problems. Fresh problem windows reduce one contamination path, but results still need a release and setup check.
LiveCodeBench ranking
7 models with a published LiveCodeBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Competitive-programming evaluation |
|---|---|---|---|
| 1 | Alibaba | 91.6 | |
| 2 | Alibaba | 89.6 | |
| 3 | Upstage | 87.8 | |
| 4 | Z.AI | 84.9 | |
| 5 | Alibaba | 83.9 | |
| 6 | Alibaba | 80.4 | |
| 7 | DeepSeek | 37.6 |
Evidence key: Observed
Rows are ordered by the value LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.