Mathematics benchmark
FrontierMath v2 (Tiers 1-3) leaderboard
FrontierMath v2 Tiers 1-3. Every model the catalog carries a published FrontierMath v2 (Tiers 1-3) value for, ranked by that value.
Epoch AI's corrected v2 core FrontierMath suite of private advanced mathematics problems. Models can reason iteratively and use Python; scores are pass rates on the private set.
FrontierMath v2 (Tiers 1-3) ranking
51 models with a published FrontierMath v2 (Tiers 1-3) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Python-enabled iterative mathematical problem solving |
|---|---|---|---|
| 1 | OpenAI | 89.000 | |
| 2 | OpenAI | 84.900 | |
| 3 | OpenAI | 78.600 | |
| 4 | OpenAI | 51.700 | |
| 5 | OpenAI | 51.000 | |
| 6 | OpenAI | 50.000 | |
| 7 | OpenAI | 47.600 | |
| 8 | Anthropic | 47.241 | |
| 9 | Anthropic | 43.793 | |
| 10 | Anthropic | 40.700 | |
| 10 | OpenAI | 40.700 | |
| 12 | Meta | 39.000 | |
| 13 | 38.966 | ||
| 13 | Moonshot AI | 38.966 | |
| 15 | 37.600 | ||
| 16 | 36.900 | ||
| 17 | 35.640 | ||
| 18 | Z.AI | 33.448 | |
| 19 | Anthropic | 32.400 | |
| 20 | OpenAI | 31.034 | |
| 21 | OpenAI | 28.280 | |
| 22 | Moonshot AI | 27.900 | |
| 23 | Alibaba | 26.207 | |
| 24 | OpenAI | 25.860 | |
| 25 | OpenAI | 24.828 | |
| 26 | Alibaba | 23.103 | |
| 27 | DeepSeek | 22.100 | |
| 28 | Moonshot AI | 21.404 | |
| 29 | Alibaba | 21.034 | |
| 30 | Anthropic | 20.690 | |
| 31 | xAI | 19.655 | |
| 32 | OpenAI | 18.685 | |
| 33 | Z.AI | 16.434 | |
| 34 | 14.138 | ||
| 35 | Anthropic | 13.495 | |
| 36 | OpenAI | 9.310 | |
| 37 | Alibaba | 8.481 | |
| 38 | Alibaba | 6.207 | |
| 39 | Anthropic | 5.903 | |
| 40 | OpenAI | 5.517 | |
| 41 | 4.844 | ||
| 42 | OpenAI | 4.483 | |
| 43 | Z.AI | 3.819 | |
| 44 | xAI | 3.793 | |
| 45 | Z.AI | 2.439 | |
| 46 | Anthropic | 2.069 | |
| 47 | DeepSeek | 1.724 | |
| 48 | OpenAI | 1.034 | |
| 49 | Meta | 0.690 | |
| 50 | OpenAI | 0.345 | |
| 51 | Meta | 0.000 |
Evidence key: Observed
Rows are ordered by the value FrontierMath v2 benchmark hub published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.