Coding benchmark
Vibe Code Bench leaderboard
Vibe Code Bench v1.1. Every model the catalog carries a published Vibe Code Bench value for, ranked by that value.
Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.
Vibe Code Bench ranking
41 models with a published Vibe Code Bench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Full-stack app implementation benchmark |
|---|---|---|---|
| 1 | Anthropic | 71.00 | |
| 2 | OpenAI | 69.85 | |
| 3 | OpenAI | 67.42 | |
| 4 | OpenAI | 61.77 | |
| 5 | Anthropic | 57.57 | |
| 6 | OpenAI | 53.50 | |
| 7 | Anthropic | 53.50 | |
| 8 | Anthropic | 51.48 | |
| 9 | DeepSeek | 49.93 | |
| 10 | 48.68 | ||
| 11 | OpenAI | 47.97 | |
| 12 | OpenAI | 37.91 | |
| 13 | Moonshot AI | 37.89 | |
| 14 | 32.03 | ||
| 15 | Z.AI | 31.46 | |
| 16 | MiniMax | 27.04 | |
| 17 | OpenAI | 26.10 | |
| 18 | Alibaba | 25.56 | |
| 19 | OpenAI | 24.61 | |
| 20 | Z.AI | 23.36 | |
| 21 | Anthropic | 22.62 | |
| 22 | OpenAI | 22.17 | |
| 23 | Anthropic | 20.63 | |
| 24 | 20.20 | ||
| 25 | OpenAI | 20.09 | |
| 26 | Meta | 19.67 | |
| 27 | Moonshot AI | 17.54 | |
| 28 | Alibaba | 15.74 | |
| 29 | MiniMax | 14.85 | |
| 30 | 14.30 | ||
| 31 | OpenAI | 14.17 | |
| 32 | OpenAI | 13.12 | |
| 33 | Anthropic | 11.39 | |
| 34 | DeepSeek | 5.11 | |
| 35 | xAI | 4.06 | |
| 36 | Alibaba | 3.51 | |
| 37 | Z.AI | 3.09 | |
| 38 | xAI | 1.20 | |
| 39 | 0.40 | ||
| 40 | 0.00 | ||
| 40 | xAI | 0.00 |
Evidence key: Observed
Rows are ordered by the value Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.