Skip to main content
ModelScale

Coding benchmark

LiveCodeBench (Vals) leaderboard

LiveCodeBench, Vals AI run. Every model the catalog carries a published LiveCodeBench (Vals) value for, ranked by that value.

CategoryCoding
MeasurePass@1 accuracy
TasksCompetitive programming problems (easy, medium, hard)
DifficultyFrontier coding

Vals AI’s independent implementation of the LiveCodeBench code-generation benchmark, run under one fixed harness across the models it tracks.

LiveCodeBench (Vals) ranking

48 models with a published LiveCodeBench (Vals) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published LiveCodeBench, Vals AI run value
RankModelProviderPass@1 accuracy
1Claude Fable 5.1Anthropic90.5
2Claude Fable 5Anthropic89.8
3Gemini 3.8 FlashGoogle89.5
4Claude Opus 5Anthropic89.0
5Gemini 3.7 FlashGoogle88.7
6Gemini 3.1 ProGoogle88.5
7Grok 4.6xAI88.2
8Gemini 3.6 FlashGoogle88.1
9Qwen3.8 MaxAlibaba87.9
10Claude Opus 4.8Anthropic87.8
11Gemini 3.5 FlashGoogle87.6
12DeepSeek V4 Pro 0813DeepSeek87.5
13Grok 4.5xAI87.4
14DeepSeek V4 Flash 0731DeepSeek87.3
15Kimi K3Moonshot AI87.2
16Qwen3.7 MaxAlibaba87.1
17Kimi K2.6Moonshot AI86.8
18Nemotron 3 UltraNVIDIA86.0
18Qwen3.6 PlusAlibaba86.0
20GPT-5.6 TerraOpenAI85.9
20TMInkling-SmallThinking Machines Lab85.9
20Muse Spark 1.1Meta85.9
23Gemini 3 FlashGoogle85.6
24TMInklingThinking Machines Lab85.5
25GPT-5.5OpenAI85.3
26Claude Opus 4.7Anthropic85.1
27Grok 4.3xAI84.5
28Grok 4.20xAI84.3
29GPT-5.4 nanoOpenAI84.0
29Ling 3.0 FlashInclusionAI84.0
29Qwen3.8-27BAlibaba84.0
32GPT-5.6 SolOpenAI82.6
33Claude Sonnet 5Anthropic82.4
34MiniMax M3MiniMax82.2
35Claude Sonnet 4.6Anthropic82.1
36GPT-5.4 miniOpenAI81.5
36MiMo-V2.5Xiaomi81.5
38GLM-5.1Z.AI81.4
38MiMo-V2.5-ProXiaomi81.4
40GLM-5.3Z.AI80.5
40GLM-5.3-FlashZ.AI80.5
42Gemini 3.1 Flash-LiteGoogle80.1
43MiniMax M2.7MiniMax79.9
44Gemini 3.5 Flash-LiteGoogle79.0
45GLM-5.2Z.AI69.5
46Laguna M.1Poolside68.1
47Laguna XS.2Poolside67.8
48Claude Haiku 4.5Anthropic41.2

Evidence key: Observed

Rows are ordered by the value Vals AI LiveCodeBench, Vals AI run leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard