Agentic benchmark
InferenceBench leaderboard
Every model the catalog carries a published InferenceBench value for, ranked by that value.
CategoryAgentic
MeasureTwo-hour autonomous CLI agent run
Tasks4 inference-serving optimization scenarios
DifficultyOpen-ended ML systems engineering
Published byInferenceBench
A benchmark for open-ended LLM inference optimization by AI agents. Agents receive a base model, one H100, and a fixed time budget to build a valid OpenAI-compatible inference server that improves serving speed.
No model in the catalog has a published InferenceBench score.
The benchmark is defined by InferenceBench, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.