Reasoning benchmark
MuSR leaderboard
Testing the Limits of Chain-of-thought with Multistep Soft Reasoning. Every model the catalog carries a published MuSR value for, ranked by that value.
CategoryReasoning
MeasureNarrative-based reasoning
TasksMulti-step reasoning
DifficultyComplex reasoning tasks
A dataset for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Tests the ability to perform complex, structured reasoning.
No model in the catalog has a published MuSR score.
The benchmark is defined by MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.