Mathematics benchmark
FrontierMath (legacy) leaderboard
FrontierMath legacy aggregate. Every model the catalog carries a published FrontierMath (legacy) value for, ranked by that value.
Legacy FrontierMath values retained for historical model pages. This field is not used in current rankings because it can mix prior benchmark versions and slices.
FrontierMath (legacy) ranking
7 models with a published FrontierMath (legacy) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Open-ended mathematical reasoning with tool access |
|---|---|---|---|
| 1 | OpenAI | 89 | |
| 2 | OpenAI | 84.9 | |
| 3 | OpenAI | 78.6 | |
| 4 | OpenAI | 52.4 | |
| 5 | OpenAI | 51.7 | |
| 6 | OpenAI | 50 | |
| 7 | Anthropic | 43.8 |
Evidence key: Observed
Rows are ordered by the value FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.