Reasoning benchmark
CLUEWSC leaderboard
Every model the catalog carries a published CLUEWSC value for, ranked by that value.
CategoryReasoning
MeasureExact match
TasksChinese coreference questions
DifficultyChinese commonsense reasoning
Published byDeepSeek-V4 Technical Report
A Chinese Winograd Schema Challenge benchmark reported in DeepSeek-V4 base-model evaluations.
No model in the catalog has a published CLUEWSC score.
The benchmark is defined by DeepSeek-V4 Technical Report, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.