Skip to main content
ModelScale

Reasoning benchmark

BullshitBench v2 leaderboard

Every model the catalog carries a published BullshitBench v2 value for, ranked by that value.

CategoryReasoning
MeasurePrompt challenge and refusal evaluation
TasksNonsensical and flawed prompts across multiple domains
DifficultyRobustness and critical reasoning

A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.

No model in the catalog has a published BullshitBench v2 score.

The benchmark is defined by BullshitBench: Measuring whether AI models challenge nonsensical prompts, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Reasoning capability leaderboard