Skip to main content
ModelScale

Mathematics benchmark

FrontierMath v2 (Tier 4) leaderboard

FrontierMath v2 Tier 4. Every model the catalog carries a published FrontierMath v2 (Tier 4) value for, ranked by that value.

CategoryMathematics
MeasurePython-enabled iterative mathematical problem solving
Tasks43 private extreme-difficulty mathematics problems
DifficultyResearch-level mathematics requiring hours or days of expert work

Epoch AI's corrected v2 Tier 4 expansion, a separate set of exceptionally difficult research-level mathematics problems evaluated with Python-enabled iterative reasoning.

FrontierMath v2 (Tier 4) ranking

45 models with a published FrontierMath v2 (Tier 4) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published FrontierMath v2 Tier 4 value
RankModelProviderPython-enabled iterative mathematical problem solving
1GPT-6 AstraOpenAI97.600
2GPT-5.6 SolOpenAI83.000
3GPT-5.6 TerraOpenAI68.300
4GPT-5.6 LunaOpenAI58.500
5GPT-5.5 ProOpenAI39.600
6GPT-5.4 ProOpenAI37.500
7GPT-5.5OpenAI35.400
8Claude Opus 4.8Anthropic31.250
9GPT-5.4OpenAI27.100
10Claude Opus 4.7Anthropic22.917
11Claude Opus 4.6Anthropic22.900
12GPT-5.2OpenAI18.800
13Gemini 3 ProGoogle18.750
14Gemini 3.1 ProGoogle16.700
15Muse SparkMeta14.600
16Gemini 3.5 FlashGoogle14.583
17Kimi K2.6Moonshot AI14.580
18GLM-5.1Z.AI12.500
18GPT-5.1OpenAI12.500
20Qwen3.6 PlusAlibaba8.333
21Claude Sonnet 4.6Anthropic8.300
22GPT-5.4 nanoOpenAI6.250
22o4-mini (high)OpenAI6.250
24Kimi K2.5Moonshot AI4.200
25Claude Opus 4.5Anthropic4.167
25Claude Sonnet 4.5Anthropic4.167
25Gemini 2.5 FlashGoogle4.167
25Gemini 2.5 ProGoogle4.167
25Gemini 3 FlashGoogle4.167
25Qwen 3.6 Max (preview)Alibaba4.167
31GLM-4.6Z.AI2.128
32DeepSeek V3.2DeepSeek2.100
32GLM-5Z.AI2.100
34Claude Haiku 4.5Anthropic2.083
34Grok 4xAI2.083
34o3OpenAI2.083
34Qwen3.5 PlusAlibaba2.083
38GPT-5.4 miniOpenAI2.080
39Claude 3.5 SonnetAnthropic0.000
39GLM-4.7Z.AI0.000
39GPT-4.1OpenAI0.000
39Grok 3 [Beta]xAI0.000
39Kimi K2Moonshot AI0.000
39Qwen3 235B 2507 (Reasoning)Alibaba0.000
39Qwen3.5 FlashAlibaba0.000

Evidence key: Observed

Rows are ordered by the value FrontierMath Tier 4 v2 leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards