Reasoning benchmark
WildBench leaderboard
Every model the catalog carries a published WildBench value for, ranked by that value.
CategoryReasoning
MeasureReal-world task evaluation
Tasks1,024 real-world tasks
DifficultyDiverse real-world scenarios
An automated evaluation framework using 1,000+ real-world user tasks covering reasoning, planning, coding, and creative writing. Highly correlated with Chatbot Arena human preference rankings.
No model in the catalog has a published WildBench score.
The benchmark is defined by WildBench: Benchmarking Language Models with Challenging Tasks from Real Users in the Wild, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.