Skip to main content
ModelScale

Reasoning benchmark

BBH leaderboard

BIG-Bench Hard. Every model the catalog carries a published BBH value for, ranked by that value.

CategoryReasoning
MeasureMixed reasoning tasks
Tasks23 tasks
DifficultyAdvanced reasoning

A suite of 23 challenging tasks from the BIG-Bench collaborative benchmark where prior language models failed to exceed average human performance, even with chain-of-thought prompting.

BBH ranking

3 models with a published BBH value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published BIG-Bench Hard value
RankModelProviderMixed reasoning tasks
1SPSoofi S 30B-A3BSoofi Project78.8
2OPMiniCPM5-1BOpenBMB71.89
3Gemma 4 12BGoogle53

Evidence key: Observed

Rows are ordered by the value Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Reasoning capability leaderboard