Reasoning benchmark
MRCRv2 leaderboard
Every model the catalog carries a published MRCRv2 value for, ranked by that value.
A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.
MRCRv2 ranking
9 models with a published MRCRv2 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Multi-round long-context evaluation |
|---|---|---|---|
| 1 | SASakana Fugu-Ultra | Sakana AI | 93.6 |
| 2 | Alibaba | 92.9 | |
| 3 | Alibaba | 91.7 | |
| 4 | Alibaba | 90.4 | |
| 5 | SASakana Fugu | Sakana AI | 86.6 |
| 6 | 77.3 | ||
| 7 | 72.2 | ||
| 8 | PAPokee-Isaac 28B | Pokee AI | 60.7 |
| 9 | 43.4 |
Evidence key: Observed
Rows are ordered by the value Introducing GPT-5.2 and GPT-5.2 Pro published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.