Knowledge benchmark
MedXpertQA (Text) leaderboard
MedXpertQA Text. Every model the catalog carries a published MedXpertQA (Text) value for, ranked by that value.
A medical multiple-choice benchmark spanning many specialties with 10 answer options per question.
MedXpertQA (Text) ranking
5 models with a published MedXpertQA (Text) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Medical MCQ |
|---|---|---|---|
| 1 | 71.5 | ||
| 2 | OpenAI | 59.6 | |
| 3 | Meta | 52.6 | |
| 4 | Anthropic | 52.1 | |
| 5 | xAI | 50.2 |
Evidence key: Observed
Rows are ordered by the value Muse Spark Eval Methodology published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.