Reasoning benchmark
MLCR-AA leaderboard
Medical Long Context Reasoning (MLCR-AA). Every model the catalog carries a published MLCR-AA value for, ranked by that value.
An open benchmark from Wisedocs, run by Artificial Analysis, measuring how well models reason over long, fragmented medical records with multi-document synthesis.
MLCR-AA ranking
16 models with a published MLCR-AA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Accuracy |
|---|---|---|---|
| 1 | Anthropic | 71.1 | |
| 2 | Anthropic | 64.4 | |
| 3 | Anthropic | 55.6 | |
| 4 | Anthropic | 55.0 | |
| 5 | Z.AI | 51.1 | |
| 6 | Z.AI | 48.3 | |
| 7 | Meta | 41.1 | |
| 8 | Moonshot AI | 38.3 | |
| 9 | OpenAI | 35.0 | |
| 10 | OpenAI | 31.7 | |
| 11 | OpenAI | 26.1 | |
| 12 | 21.7 | ||
| 12 | Alibaba | 21.7 | |
| 14 | Meta | 20.0 | |
| 15 | OpenAI | 19.4 | |
| 16 | DeepSeek | 17.8 |
Evidence key: Observed
Rows are ordered by the value Medical Long Context Reasoning (MLCR-AA) published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.