Skip to main content
ModelScale

external benchmark

ALE-Bench leaderboard

Agents Last Exam. Every model the catalog carries a published ALE-Bench value for, ranked by that value.

Categoryexternal
MeasurePass rate, partial-credit score, cost, token, and duration metadata
Tasks152 ALE-V1 professional workflow tasks across 13 top-level domains
DifficultyReal-world agentic workflows
Published byAgents Last Exam

A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.

No model in the catalog has a published ALE-Bench score.

The benchmark is defined by Agents Last Exam, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards