Knowledge benchmark
HealthBench Hard leaderboard
Every model the catalog carries a published HealthBench Hard value for, ranked by that value.
A harder subset of OpenAI's HealthBench for evaluating open-ended medical and health reasoning with rubric-based grading.
HealthBench Hard ranking
9 models with a published HealthBench Hard value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Open-ended health evaluation |
|---|---|---|---|
| 1 | Meta | 42.8 | |
| 2 | OpenAI | 40.1 | |
| 3 | OpenAI | 36.3 | |
| 4 | OpenAI | 33.1 | |
| 5 | OpenAI | 32.7 | |
| 6 | OpenAI | 32.0 | |
| 7 | 20.6 | ||
| 8 | xAI | 20.3 | |
| 9 | Anthropic | 14.8 |
Evidence key: Observed
Rows are ordered by the value Muse Spark Eval Methodology published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.