Skip to main content
ModelScale

external benchmark

WeirdML leaderboard

WeirdML v2. Every model the catalog carries a published WeirdML value for, ranked by that value.

Categoryexternal
MeasureAverage accuracy across tasks
Tasks17 novel ML engineering tasks
DifficultyNovel dataset modeling and iterative debugging
Published byWeirdML

A machine-learning engineering benchmark that tests whether LLMs can train models on novel datasets, write PyTorch code, and improve through iterative feedback.

No model in the catalog has a published WeirdML score.

The benchmark is defined by WeirdML, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards