Skip to main content
ModelScale

Coding benchmark

VulcanBench v3 leaderboard

Every model the catalog carries a published VulcanBench v3 value for, ranked by that value.

CategoryCoding
MeasurePass@1 with low, medium, and high effort
Tasks23 post-cutoff repository tasks in the v3 report
DifficultyProfessional multi-file software engineering
Published byVulcanBench

An open software-engineering benchmark built from real merged post-cutoff pull requests across Python, Rust, TypeScript, JavaScript, and Go repositories.

VulcanBench v3 ranking

14 models with a published VulcanBench v3 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published VulcanBench v3 value
RankModelProviderPass@1 with low, medium, and high effort
1Grok 4.5xAI89.9
2Claude Fable 5Anthropic89.5
3DeepSeek V4 Flash 0731DeepSeek88.4
4Claude Opus 5Anthropic87.0
4GPT-5.6 SolOpenAI87.0
4GPT-5.6 TerraOpenAI87.0
4Grok 4.6xAI87.0
4Muse Spark 1.2Meta87.0
9GPT-5.6 LunaOpenAI85.5
10Qwen3.8-27BAlibaba82.6
11Qwen3.8 MaxAlibaba81.2
12GLM-5.3Z.AI78.3
13Claude Haiku 4.5Anthropic76.2
14Kimi K3Moonshot AI73.7

Evidence key: Observed

Rows are ordered by the value VulcanBench published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard