Coding benchmark
App-Bench leaderboard
Every model the catalog carries a published App-Bench value for, ranked by that value.
CategoryCoding
MeasureBest-of-three one-shot feature completion
Tasks6 full-stack app-building tasks
DifficultyProduction-style full-stack application generation
Published byApp-Bench
A six-task full-stack web-app benchmark that measures how much required functionality an AI builder or coding assistant delivers from one prompt without human code edits.
No model in the catalog has a published App-Bench score.
The benchmark is defined by App-Bench, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.