Mathematics benchmark
FrontierMath v2 (Tier 4) leaderboard
FrontierMath v2 Tier 4. Every model the catalog carries a published FrontierMath v2 (Tier 4) value for, ranked by that value.
Epoch AI's corrected v2 Tier 4 expansion, a separate set of exceptionally difficult research-level mathematics problems evaluated with Python-enabled iterative reasoning.
FrontierMath v2 (Tier 4) ranking
45 models with a published FrontierMath v2 (Tier 4) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Python-enabled iterative mathematical problem solving |
|---|---|---|---|
| 1 | OpenAI | 97.600 | |
| 2 | OpenAI | 83.000 | |
| 3 | OpenAI | 68.300 | |
| 4 | OpenAI | 58.500 | |
| 5 | OpenAI | 39.600 | |
| 6 | OpenAI | 37.500 | |
| 7 | OpenAI | 35.400 | |
| 8 | Anthropic | 31.250 | |
| 9 | OpenAI | 27.100 | |
| 10 | Anthropic | 22.917 | |
| 11 | Anthropic | 22.900 | |
| 12 | OpenAI | 18.800 | |
| 13 | 18.750 | ||
| 14 | 16.700 | ||
| 15 | Meta | 14.600 | |
| 16 | 14.583 | ||
| 17 | Moonshot AI | 14.580 | |
| 18 | Z.AI | 12.500 | |
| 18 | OpenAI | 12.500 | |
| 20 | Alibaba | 8.333 | |
| 21 | Anthropic | 8.300 | |
| 22 | OpenAI | 6.250 | |
| 22 | OpenAI | 6.250 | |
| 24 | Moonshot AI | 4.200 | |
| 25 | Anthropic | 4.167 | |
| 25 | Anthropic | 4.167 | |
| 25 | 4.167 | ||
| 25 | 4.167 | ||
| 25 | 4.167 | ||
| 25 | Alibaba | 4.167 | |
| 31 | Z.AI | 2.128 | |
| 32 | DeepSeek | 2.100 | |
| 32 | Z.AI | 2.100 | |
| 34 | Anthropic | 2.083 | |
| 34 | xAI | 2.083 | |
| 34 | OpenAI | 2.083 | |
| 34 | Alibaba | 2.083 | |
| 38 | OpenAI | 2.080 | |
| 39 | Anthropic | 0.000 | |
| 39 | Z.AI | 0.000 | |
| 39 | OpenAI | 0.000 | |
| 39 | xAI | 0.000 | |
| 39 | Moonshot AI | 0.000 | |
| 39 | Alibaba | 0.000 | |
| 39 | Alibaba | 0.000 |
Evidence key: Observed
Rows are ordered by the value FrontierMath Tier 4 v2 leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.