Agentic benchmark
AI4AI-Bench leaderboard
AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement. Every model the catalog carries a published AI4AI-Bench value for, ranked by that value.
CategoryAgentic
MeasureFour-hour code rewrite followed by sealed training and held-out evaluation
Tasks10 AI training-algorithm design tasks
DifficultyEnd-to-end AI research and algorithm design
Tests whether coding agents can improve the training algorithm inside an existing AI research codebase, then survive a sealed training run and held-out evaluation.
No model in the catalog has a published AI4AI-Bench score.
The benchmark is defined by AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.