Skip to main content
ModelScale

Reasoning benchmark

Pencil Puzzle Bench leaderboard

Every model the catalog carries a published Pencil Puzzle Bench value for, ranked by that value.

CategoryReasoning
MeasureDirect and agentic puzzle solve rate
Tasks300 evaluation puzzles
DifficultyMulti-step verifiable reasoning

A multi-step verifiable reasoning benchmark that evaluates whether models can solve pencil puzzles with unique solutions.

No model in the catalog has a published Pencil Puzzle Bench score.

The benchmark is defined by Pencil Puzzle Bench, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Reasoning capability leaderboard