Reasoning benchmark
BullshitBench v2 leaderboard
Every model the catalog carries a published BullshitBench v2 value for, ranked by that value.
CategoryReasoning
MeasurePrompt challenge and refusal evaluation
TasksNonsensical and flawed prompts across multiple domains
DifficultyRobustness and critical reasoning
A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.
No model in the catalog has a published BullshitBench v2 score.
The benchmark is defined by BullshitBench: Measuring whether AI models challenge nonsensical prompts, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.