Skip to main content
ModelScale

Reasoning benchmark

WildBench leaderboard

Every model the catalog carries a published WildBench value for, ranked by that value.

CategoryReasoning
MeasureReal-world task evaluation
Tasks1,024 real-world tasks
DifficultyDiverse real-world scenarios

An automated evaluation framework using 1,000+ real-world user tasks covering reasoning, planning, coding, and creative writing. Highly correlated with Chatbot Arena human preference rankings.

No model in the catalog has a published WildBench score.

The benchmark is defined by WildBench: Benchmarking Language Models with Challenging Tasks from Real Users in the Wild, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Reasoning capability leaderboard