Agentic benchmark
TAU-bench leaderboard
Tool-Agent-User Benchmark. Every model the catalog carries a published TAU-bench value for, ranked by that value.
CategoryAgentic
MeasureDomain-specific pass^1 through pass^4 task success
TasksAirline and retail task sets in the archived 2024 release
DifficultyPolicy-constrained, multi-turn customer service
Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.
No model in the catalog has a published TAU-bench score.
The benchmark is defined by τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.