Mathematics benchmark
MATH-500 leaderboard
MATH-500 Problem Set. Every model the catalog carries a published MATH-500 value for, ranked by that value.
A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.
MATH-500 ranking
7 models with a published MATH-500 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Free-form mathematical answers |
|---|---|---|---|
| 1 | PMTernary Bonsai 2 27B | Prism ML | 98.8 |
| 2 | Meituan | 95.8 | |
| 3 | OPMiniCPM5-2B | OpenBMB | 94.6 |
| 4 | OPMiniCPM5-1B | OpenBMB | 91.6 |
| 5 | LiquidAI | 88.76 | |
| 6 | KAKanana-2 1.3B Instruct | Kakao | 61.4 |
| 7 | KAKanana-2 3B Instruct | Kakao | 61.2 |
Evidence key: Observed
Rows are ordered by the value Measuring Mathematical Problem Solving With the MATH Dataset published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.