Skip to main content
ModelScale

Multimodal & Grounded benchmark

MedXpertQA (MM) leaderboard

MedXpertQA Multimodal. Every model the catalog carries a published MedXpertQA (MM) value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureMedical visual MCQ
Tasks2,000 multimodal medical questions
DifficultyClinical multimodal reasoning

A multimodal medical multiple-choice benchmark covering clinical images such as X-rays, histology, and dermatology.

MedXpertQA (MM) ranking

8 models with a published MedXpertQA (MM) value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published MedXpertQA Multimodal value
RankModelProviderMedical visual MCQ
1Gemini 3.1 ProGoogle81.3
2Qwen3.8 MaxAlibaba80.4
3Muse SparkMeta78.4
4GPT-5.4OpenAI77.1
5Qwen3.7 PlusAlibaba71.0
6Grok 4.20xAI65.8
7Claude Opus 4.6Anthropic64.8
8Gemma 4 12BGoogle48.7

Evidence key: Observed

Rows are ordered by the value Muse Spark Eval Methodology published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards