external benchmark
LiveBench leaderboard
Every model the catalog carries a published LiveBench value for, ranked by that value.
Categoryexternal
MeasureMean of category averages
Tasks23 objective tasks across 7 categories
DifficultyBroad frontier-model evaluation
A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.
No model in the catalog has a published LiveBench score.
The benchmark is defined by LiveBench: A Challenging, Contamination-Free LLM Benchmark, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.