Skip to main content
ModelScale

Mathematics benchmark

FrontierMath v2 (Tiers 1-3) leaderboard

FrontierMath v2 Tiers 1-3. Every model the catalog carries a published FrontierMath v2 (Tiers 1-3) value for, ranked by that value.

CategoryMathematics
MeasurePython-enabled iterative mathematical problem solving
Tasks295 private advanced mathematics problems
DifficultyFrom olympiad-plus to early research mathematics

Epoch AI's corrected v2 core FrontierMath suite of private advanced mathematics problems. Models can reason iteratively and use Python; scores are pass rates on the private set.

FrontierMath v2 (Tiers 1-3) ranking

51 models with a published FrontierMath v2 (Tiers 1-3) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published FrontierMath v2 Tiers 1-3 value
RankModelProviderPython-enabled iterative mathematical problem solving
1GPT-5.6 SolOpenAI89.000
2GPT-5.6 TerraOpenAI84.900
3GPT-5.6 LunaOpenAI78.600
4GPT-5.5OpenAI51.700
5GPT-5.5 ProOpenAI51.000
6GPT-5.4 ProOpenAI50.000
7GPT-5.4OpenAI47.600
8Claude Opus 4.8Anthropic47.241
9Claude Opus 4.7Anthropic43.793
10Claude Opus 4.6Anthropic40.700
10GPT-5.2OpenAI40.700
12Muse SparkMeta39.000
13Gemini 3.5 FlashGoogle38.966
13Kimi K2.6Moonshot AI38.966
15Gemini 3 ProGoogle37.600
16Gemini 3.1 ProGoogle36.900
17Gemini 3 FlashGoogle35.640
18GLM-5.1Z.AI33.448
19Claude Sonnet 4.6Anthropic32.400
20GPT-5.1OpenAI31.034
21GPT-5.4 miniOpenAI28.280
22Kimi K2.5Moonshot AI27.900
23Qwen3.6 PlusAlibaba26.207
24GPT-5.4 nanoOpenAI25.860
25o4-mini (high)OpenAI24.828
26Qwen 3.6 Max (preview)Alibaba23.103
27DeepSeek V3.2DeepSeek22.100
28Kimi K2Moonshot AI21.404
29Qwen3.5 PlusAlibaba21.034
30Claude Opus 4.5Anthropic20.690
31Grok 4xAI19.655
32o3OpenAI18.685
33GLM-5Z.AI16.434
34Gemini 2.5 ProGoogle14.138
35Claude Sonnet 4.5Anthropic13.495
36o1OpenAI9.310
37Qwen3 235B 2507 (Reasoning)Alibaba8.481
38Qwen3.5 FlashAlibaba6.207
39Claude Haiku 4.5Anthropic5.903
40GPT-4.1OpenAI5.517
41Gemini 2.5 FlashGoogle4.844
42GPT-4.1 miniOpenAI4.483
43GLM-4.6Z.AI3.819
44Grok 3 [Beta]xAI3.793
45GLM-4.7Z.AI2.439
46Claude 3.5 SonnetAnthropic2.069
47DeepSeek V3DeepSeek1.724
48GPT-4.1 nanoOpenAI1.034
49Llama 4 MaverickMeta0.690
50GPT-4oOpenAI0.345
51Llama 4 ScoutMeta0.000

Evidence key: Observed

Rows are ordered by the value FrontierMath v2 benchmark hub published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards