Mathematics benchmark
AIME 2025 leaderboard
American Invitational Mathematics Examination 2025. Every model the catalog carries a published AIME 2025 value for, ranked by that value.
The most recent AIME examination, featuring 15 challenging mathematics problems testing olympiad-level mathematical reasoning with integer answers from 000-999.
AIME 2025 ranking
16 models with a published AIME 2025 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Integer answers 000-999 |
|---|---|---|---|
| 1 | Microsoft | 97 | |
| 2 | Moonshot AI | 96.1 | |
| 2 | Moonshot AI | 96.1 | |
| 4 | Z.AI | 95.7 | |
| 5 | PMTernary Bonsai 2 27B | Prism ML | 95 |
| 6 | Xiaomi | 94.1 | |
| 7 | IBM | 89.17 | |
| 8 | Anthropic | 87 | |
| 9 | IBM | 86.67 | |
| 10 | OPMiniCPM5-2B | OpenBMB | 86.5 |
| 11 | LG AI Research | 85.3 | |
| 12 | NVIDIA | 82.1 | |
| 13 | IBM | 78.33 | |
| 14 | LiquidAI | 51.87 | |
| 15 | LiquidAI | 42.53 | |
| 16 | OPMiniCPM5-1B | OpenBMB | 40.42 |
Evidence key: Observed
Rows are ordered by the value American Invitational Mathematics Examination published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.