Skip to main content
ModelScale

Reasoning benchmark

MuSR leaderboard

Testing the Limits of Chain-of-thought with Multistep Soft Reasoning. Every model the catalog carries a published MuSR value for, ranked by that value.

CategoryReasoning
MeasureNarrative-based reasoning
TasksMulti-step reasoning
DifficultyComplex reasoning tasks

A dataset for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Tests the ability to perform complex, structured reasoning.

No model in the catalog has a published MuSR score.

The benchmark is defined by MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Reasoning capability leaderboard