Reasoning benchmark
Conceptual Reasoning leaderboard
Conceptual Reasoning Benchmark. Every model the catalog carries a published Conceptual Reasoning value for, ranked by that value.
CategoryReasoning
MeasureAverage pairwise-ranking loss against expert ratings
Tasks224 texts and 608 within-text critique pairs
DifficultyFuzzy, expert-rated argumentative reasoning
Published byConceptual Reasoning Benchmark Results
Tests whether model judgments rank argumentative critiques in the same order as expert human ratings across philosophy, AI alignment, and other concept-heavy texts.
No model in the catalog has a published Conceptual Reasoning score.
The benchmark is defined by Conceptual Reasoning Benchmark Results, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.