Skip to main content
ModelScale

Reasoning benchmark

Conceptual Reasoning leaderboard

Conceptual Reasoning Benchmark. Every model the catalog carries a published Conceptual Reasoning value for, ranked by that value.

CategoryReasoning
MeasureAverage pairwise-ranking loss against expert ratings
Tasks224 texts and 608 within-text critique pairs
DifficultyFuzzy, expert-rated argumentative reasoning

Tests whether model judgments rank argumentative critiques in the same order as expert human ratings across philosophy, AI alignment, and other concept-heavy texts.

No model in the catalog has a published Conceptual Reasoning score.

The benchmark is defined by Conceptual Reasoning Benchmark Results, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Reasoning capability leaderboard