Skip to main content
ModelScale

Coding benchmark

LiveCodeBench leaderboard

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. Every model the catalog carries a published LiveCodeBench value for, ranked by that value.

CategoryCoding
MeasureCompetitive-programming evaluation
TasksContinuously updated contest problems
DifficultyCompetitive programming level

A continuously updated coding benchmark built from newly collected LeetCode, AtCoder, and Codeforces problems. Fresh problem windows reduce one contamination path, but results still need a release and setup check.

LiveCodeBench ranking

7 models with a published LiveCodeBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code value
RankModelProviderCompetitive-programming evaluation
1Qwen3.7 MaxAlibaba91.6
2Qwen3.7 PlusAlibaba89.6
3Solar Pro 4Upstage87.8
4GLM-4.7Z.AI84.9
5Qwen3.6-27BAlibaba83.9
6Qwen3.6-35B-A3BAlibaba80.4
7DeepSeek V3DeepSeek37.6

Evidence key: Observed

Rows are ordered by the value LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard