Reasoning benchmark
BBH leaderboard
BIG-Bench Hard. Every model the catalog carries a published BBH value for, ranked by that value.
A suite of 23 challenging tasks from the BIG-Bench collaborative benchmark where prior language models failed to exceed average human performance, even with chain-of-thought prompting.
BBH ranking
3 models with a published BBH value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Mixed reasoning tasks |
|---|---|---|---|
| 1 | SPSoofi S 30B-A3B | Soofi Project | 78.8 |
| 2 | OPMiniCPM5-1B | OpenBMB | 71.89 |
| 3 | 53 |
Evidence key: Observed
Rows are ordered by the value Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.