Skip to main content
ModelScale

Agentic benchmark

MLE-Bench Lite leaderboard

Every model the catalog carries a published MLE-Bench Lite value for, ranked by that value.

CategoryAgentic
MeasureAutonomous iterative ML optimization
TasksLow-resource ML competitions
DifficultyAgentic machine learning

A lightweight machine-learning competition benchmark that measures whether models can iteratively train, evaluate, and improve ML systems in low-resource settings.

MLE-Bench Lite ranking

3 models with a published MLE-Bench Lite value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published MLE-Bench Lite value
RankModelProviderAutonomous iterative ML optimization
1Atria Dawn PreviewShanghai Artificial Intelligence Laboratory86.2
2MiniMax M2.7MiniMax66.6
3Agents-A1-4BInternScience22.7

Evidence key: Observed

Rows are ordered by the value MiniMax M2.7: Early Echoes of Self-Evolution published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard