Skip to main content
ModelScale

Reasoning benchmark

MLCR-AA leaderboard

Medical Long Context Reasoning (MLCR-AA). Every model the catalog carries a published MLCR-AA value for, ranked by that value.

CategoryReasoning
MeasureAccuracy
TasksLong, fragmented medical-record reasoning
DifficultyLong-context medical reasoning

An open benchmark from Wisedocs, run by Artificial Analysis, measuring how well models reason over long, fragmented medical records with multi-document synthesis.

MLCR-AA ranking

16 models with a published MLCR-AA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Medical Long Context Reasoning (MLCR-AA) value
RankModelProviderAccuracy
1Claude Fable 5.1Anthropic71.1
2Claude Fable 5Anthropic64.4
3Claude Opus 5Anthropic55.6
4Claude Sonnet 5Anthropic55.0
5GLM-5.3-FlashZ.AI51.1
6GLM-5.3Z.AI48.3
7Muse Spark 1.3Meta41.1
8Kimi K3Moonshot AI38.3
9GPT-6 AstraOpenAI35.0
10GPT-5.6 TerraOpenAI31.7
11GPT-5.6 SolOpenAI26.1
12Gemini 3.8 FlashGoogle21.7
12Qwen3.8-27BAlibaba21.7
14Muse Glimmer 30BMeta20.0
15GPT-5.6 LunaOpenAI19.4
16DeepSeek V4 Pro 0813DeepSeek17.8

Evidence key: Observed

Rows are ordered by the value Medical Long Context Reasoning (MLCR-AA) published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Reasoning capability leaderboard