Skip to main content
ModelScale

Coding benchmark

Vibe Code Bench leaderboard

Vibe Code Bench v1.1. Every model the catalog carries a published Vibe Code Bench value for, ranked by that value.

CategoryCoding
MeasureFull-stack app implementation benchmark
TasksEnd-to-end web application builds
DifficultyEnd-to-end software delivery

Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.

Vibe Code Bench ranking

41 models with a published Vibe Code Bench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Vibe Code Bench v1.1 value
RankModelProviderFull-stack app implementation benchmark
1Claude Opus 4.7Anthropic71.00
2GPT-5.5OpenAI69.85
3GPT-5.4OpenAI67.42
4GPT-5.3 CodexOpenAI61.77
5Claude Opus 4.6Anthropic57.57
6GPT-5.2OpenAI53.50
7Claude Opus 4.6 (Adaptive)Anthropic53.50
8Claude Sonnet 4.6Anthropic51.48
9DeepSeek V4 Pro 0813DeepSeek49.93
10Gemini 3.5 FlashGoogle48.68
11GPT-5.4 miniOpenAI47.97
12GPT-5.2-CodexOpenAI37.91
13Kimi K2.6Moonshot AI37.89
14Gemini 3.1 ProGoogle32.03
15GLM-5.1Z.AI31.46
16MiniMax M2.7MiniMax27.04
17GPT-5.4 nanoOpenAI26.10
18Qwen3.6 PlusAlibaba25.56
19GPT-5.1OpenAI24.61
20GLM-5 (Reasoning)Z.AI23.36
21Claude Sonnet 4.5 ThinkingAnthropic22.62
22GPT-5.1-Codex-MaxOpenAI22.17
23Claude Opus 4.5 ThinkingAnthropic20.63
24Gemini 3 FlashGoogle20.20
25GPT-5 (high)OpenAI20.09
26Muse SparkMeta19.67
27Kimi K2.5 (Reasoning)Moonshot AI17.54
28Qwen3.5 PlusAlibaba15.74
29MiniMax M2.5MiniMax14.85
30Gemini 3 ProGoogle14.30
31GPT-5 miniOpenAI14.17
32GPT-5.1-CodexOpenAI13.12
33Claude Haiku 4.5 ThinkingAnthropic11.39
34DeepSeek V3.2 (Thinking)DeepSeek5.11
35Grok 4.20xAI4.06
36Qwen3 MaxAlibaba3.51
37GLM-4.6Z.AI3.09
38Grok 4.1 Fast (Reasoning)xAI1.20
39Gemini 2.5 ProGoogle0.40
40Gemini 3.1 Flash-LiteGoogle0.00
40Grok 4 Fast (Reasoning)xAI0.00

Evidence key: Observed

Rows are ordered by the value Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard