Skip to main content
ModelScale

Mathematics benchmark

AIME26 leaderboard

AIME 2026. Every model the catalog carries a published AIME26 value for, ranked by that value.

CategoryMathematics
MeasureShort-answer mathematics
TasksCompetition math problems
DifficultyOlympiad-style mathematics

A 2026 American Invitational Mathematics Examination snapshot used in frontier-model comparison tables for mathematical reasoning.

AIME26 ranking

28 models with a published AIME26 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published AIME 2026 value
RankModelProviderShort-answer mathematics
1GLM-5.2Z.AI99.2
2STA.X K2SK Telecom97.1
2TMInklingThinking Machines Lab97.1
4Kimi K2.6Moonshot AI96.4
5PMTernary Bonsai 2 27BPrism ML95.8
6GLM-5Z.AI95.8
6Kimi K2.5Moonshot AI95.8
8Solar Open 2Upstage95.7
9TMInkling-SmallThinking Machines Lab95.5
10GLM-5.1Z.AI95.3
10Qwen3.6 PlusAlibaba95.3
10Solar Pro 4Upstage95.3
13Claude Opus 4.5Anthropic95.1
14Muse Glimmer 30BMeta94.7
15MAI-Thinking-1Microsoft94.5
16Qwen3.6-27BAlibaba94.1
17Qwen3.5 397BAlibaba93.3
18Ling 3.0 FlashInclusionAI93.2
19Qwen3.6-35B-A3BAlibaba92.7
20K-EXAONE 2.0LG AI Research92.3
21ZYZAYA1-8BZyphra89.1
22OPMiniCPM5-2BOpenBMB86.5
23Gemma 4 12BGoogle77.5
24ZYZAYA1-74B-PreviewZyphra76.4
25LongCat-Flash-Lite-SparseMeituan65.7
26LFM2.5-8B-A1BLiquidAI50.0
27OPMiniCPM5-1BOpenBMB40.4
28LLaDA2.2-miniInclusionAI35.0

Evidence key: Observed

Rows are ordered by the value Qwen3.6 launch benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards