Skip to main content
ModelScale

Agentic benchmark

PinchBench leaderboard

Every model the catalog carries a published PinchBench value for, ranked by that value.

CategoryAgentic
MeasureAverage success rate from official runs
Tasks23 OpenClaw agent tasks
DifficultyLong-horizon agent workflows
Published byAbout PinchBench

An OpenClaw agent benchmark from Kilo that measures successful task completion across standardized real-world agent workflows.

PinchBench ranking

6 models with a published PinchBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published PinchBench value
RankModelProviderAverage success rate from official runs
1PAPokee-Isaac 28BPokee AI95.7
2Nemotron 3 UltraNVIDIA90.0
3Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA83.4
4LLaDA2.2-flashInclusionAI81.7
5LFM2.5-2.6BLiquidAI68.2
6LLaDA2.2-miniInclusionAI62.3

Evidence key: Observed

Rows are ordered by the value About PinchBench published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard