Skip to main content
ModelScale

Agentic benchmark

TAU-bench leaderboard

Tool-Agent-User Benchmark. Every model the catalog carries a published TAU-bench value for, ranked by that value.

CategoryAgentic
MeasureDomain-specific pass^1 through pass^4 task success
TasksAirline and retail task sets in the archived 2024 release
DifficultyPolicy-constrained, multi-turn customer service

Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.

No model in the catalog has a published TAU-bench score.

The benchmark is defined by τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Agentic capability leaderboard