Skip to main content
ModelScale

Reasoning benchmark

MRCRv2 leaderboard

Every model the catalog carries a published MRCRv2 value for, ranked by that value.

CategoryReasoning
MeasureMulti-round long-context evaluation
TasksLong-context retrieval
DifficultyHard long-context

A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.

MRCRv2 ranking

9 models with a published MRCRv2 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published MRCRv2 value
RankModelProviderMulti-round long-context evaluation
1SASakana Fugu-UltraSakana AI93.6
2Qwen3.8 MaxAlibaba92.9
3Qwen3.7 PlusAlibaba91.7
4Qwen3.7 MaxAlibaba90.4
5SASakana FuguSakana AI86.6
6Gemini 3.5 FlashGoogle77.3
7Gemini 3.5 Flash-LiteGoogle72.2
8PAPokee-Isaac 28BPokee AI60.7
9Gemma 4 12BGoogle43.4

Evidence key: Observed

Rows are ordered by the value Introducing GPT-5.2 and GPT-5.2 Pro published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Reasoning capability leaderboard