Skip to main content
ModelScale

Mathematics benchmark

FrontierMath (legacy) leaderboard

FrontierMath legacy aggregate. Every model the catalog carries a published FrontierMath (legacy) value for, ranked by that value.

CategoryMathematics
MeasureOpen-ended mathematical reasoning with tool access
TasksHistorical aggregate
DifficultyResearch-level mathematics

Legacy FrontierMath values retained for historical model pages. This field is not used in current rankings because it can mix prior benchmark versions and slices.

FrontierMath (legacy) ranking

7 models with a published FrontierMath (legacy) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published FrontierMath legacy aggregate value
RankModelProviderOpen-ended mathematical reasoning with tool access
1GPT-5.6 SolOpenAI89
2GPT-5.6 TerraOpenAI84.9
3GPT-5.6 LunaOpenAI78.6
4GPT-5.5 ProOpenAI52.4
5GPT-5.5OpenAI51.7
6GPT-5.4 ProOpenAI50
7Claude Opus 4.7 (Adaptive)Anthropic43.8

Evidence key: Observed

Rows are ordered by the value FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards