Skip to main content
ModelScale

Knowledge benchmark

MedXpertQA (Text) leaderboard

MedXpertQA Text. Every model the catalog carries a published MedXpertQA (Text) value for, ranked by that value.

CategoryKnowledge
MeasureMedical MCQ
Tasks2,450 medical multiple-choice questions
DifficultyProfessional medical knowledge

A medical multiple-choice benchmark spanning many specialties with 10 answer options per question.

MedXpertQA (Text) ranking

5 models with a published MedXpertQA (Text) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published MedXpertQA Text value
RankModelProviderMedical MCQ
1Gemini 3.1 ProGoogle71.5
2GPT-5.4OpenAI59.6
3Muse SparkMeta52.6
4Claude Opus 4.6Anthropic52.1
5Grok 4.20xAI50.2

Evidence key: Observed

Rows are ordered by the value Muse Spark Eval Methodology published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard