Skip to main content
ModelScale

Reasoning benchmark

LisanBench leaderboard

Every model the catalog carries a published LisanBench value for, ranked by that value.

CategoryReasoning
MeasureDifficulty-weighted word-chain reasoning
Tasks50 starting words × 3 trials
DifficultyOpen-ended lexical planning

A word-chain reasoning benchmark that tests planning, recall, constraint following, and vocabulary depth by asking models to extend non-repeating edit-distance-1 chains.

No model in the catalog has a published LisanBench score.

The benchmark is defined by LisanBench methodology, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Reasoning capability leaderboard