Agentic benchmark
WebArena leaderboard
WebArena Web Agent Benchmark. Every model the catalog carries a published WebArena value for, ranked by that value.
WebArena tests whether a browser-agent system can complete 812 long-horizon tasks inside self-hosted replicas of functional websites. It checks the requested end state, so a result reflects the model, agent scaffold, browser interface, action budget, and evaluator together—not the base model alone.
The benchmark is defined by WebArena: A Realistic Web Environment for Building Autonomous Agents, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.