Skip to main content
ModelScale

external benchmark

LiveBench leaderboard

Every model the catalog carries a published LiveBench value for, ranked by that value.

Categoryexternal
MeasureMean of category averages
Tasks23 objective tasks across 7 categories
DifficultyBroad frontier-model evaluation

A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.

No model in the catalog has a published LiveBench score.

The benchmark is defined by LiveBench: A Challenging, Contamination-Free LLM Benchmark, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards