Mathematics benchmark
AIME26 leaderboard
AIME 2026. Every model the catalog carries a published AIME26 value for, ranked by that value.
A 2026 American Invitational Mathematics Examination snapshot used in frontier-model comparison tables for mathematical reasoning.
AIME26 ranking
28 models with a published AIME26 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Short-answer mathematics |
|---|---|---|---|
| 1 | Z.AI | 99.2 | |
| 2 | STA.X K2 | SK Telecom | 97.1 |
| 2 | TMInkling | Thinking Machines Lab | 97.1 |
| 4 | Moonshot AI | 96.4 | |
| 5 | PMTernary Bonsai 2 27B | Prism ML | 95.8 |
| 6 | Z.AI | 95.8 | |
| 6 | Moonshot AI | 95.8 | |
| 8 | Upstage | 95.7 | |
| 9 | TMInkling-Small | Thinking Machines Lab | 95.5 |
| 10 | Z.AI | 95.3 | |
| 10 | Alibaba | 95.3 | |
| 10 | Upstage | 95.3 | |
| 13 | Anthropic | 95.1 | |
| 14 | Meta | 94.7 | |
| 15 | Microsoft | 94.5 | |
| 16 | Alibaba | 94.1 | |
| 17 | Alibaba | 93.3 | |
| 18 | InclusionAI | 93.2 | |
| 19 | Alibaba | 92.7 | |
| 20 | LG AI Research | 92.3 | |
| 21 | ZYZAYA1-8B | Zyphra | 89.1 |
| 22 | OPMiniCPM5-2B | OpenBMB | 86.5 |
| 23 | 77.5 | ||
| 24 | ZYZAYA1-74B-Preview | Zyphra | 76.4 |
| 25 | Meituan | 65.7 | |
| 26 | LiquidAI | 50.0 | |
| 27 | OPMiniCPM5-1B | OpenBMB | 40.4 |
| 28 | InclusionAI | 35.0 |
Evidence key: Observed
Rows are ordered by the value Qwen3.6 launch benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.