Skip to main content
ModelScale

Reasoning benchmark

WinoGrande leaderboard

Every model the catalog carries a published WinoGrande value for, ranked by that value.

CategoryReasoning
MeasureExact match
TasksCoreference resolution questions
DifficultyCommonsense reasoning

A commonsense coreference benchmark reported in DeepSeek-V4 base-model evaluations.

No model in the catalog has a published WinoGrande score.

The benchmark is defined by DeepSeek-V4 Technical Report, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Reasoning capability leaderboard