Skip to main content
ModelScale

Agentic benchmark

EdgeBench leaderboard

Every model the catalog carries a published EdgeBench value for, ranked by that value.

CategoryAgentic
MeasureLong-horizon interactive agent evaluation
Tasks134 tasks (51 public) across 6 domains
DifficultyDay-scale expert tasks

A ByteDance Seed benchmark of 134 real-world, day-scale tasks that measures how autonomous agents learn from environment feedback over 12+ hour interaction horizons, spanning scientific and ML, systems and software engineering, optimization, knowledge, formal, and game domains.

No model in the catalog has a published EdgeBench score.

The benchmark is defined by EdgeBench technical report, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Agentic capability leaderboard