Skip to main content
ModelScale

Agentic benchmark

DeepPlanning leaderboard

Every model the catalog carries a published DeepPlanning value for, ranked by that value.

CategoryAgentic
MeasureLong-horizon planning benchmark
TasksTravel planning and constrained shopping
DifficultyConstrained agent planning

A long-horizon planning benchmark that tests whether agents can optimize under explicit time, budget, and feasibility constraints.

DeepPlanning ranking

7 models with a published DeepPlanning value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published DeepPlanning value
RankModelProviderLong-horizon planning benchmark
1Qwen3.7 PlusAlibaba62.3
2Qwen3.6 PlusAlibaba41.5
3Qwen3.5 397BAlibaba37.6
4Claude Opus 4.5Anthropic26.4
5Qwen3.6-35B-A3BAlibaba25.9
6GLM-5Z.AI14.6
7Kimi K2.5Moonshot AI14.4

Evidence key: Observed

Rows are ordered by the value DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard