HealthBench Hard Leaderboard
Last updated August 16, 2026 · 9 models evaluated
Muse Spark leads the HealthBench Hard leaderboard at 0.428, ahead of GPT-5.6 Sol (0.331) and GPT-5.6 Terra (0.327). HealthBench Hard is the hardest slice of OpenAI's HealthBench benchmark: the 1,000 health conversations frontier models handled worst, graded against physician-written rubrics on a 0 to 1 scale. The best score at the benchmark's May 2025 release was o3's 0.320.
Leaderboard
1,000-conversation set, one response per conversation. Updated August 16, 2026.
Full ranking
| # | model | score | size | context | cost in / out per 1M | license | |
|---|---|---|---|---|---|---|---|
| 1 | Muse Spark Meta | 0.428 | — | 1.0M | $1.25 / $4.25 | proprietary | |
| 2 | GPT-5.6 Sol OpenAI | 0.331 | — | 1.1M | $5.00 / $30.00 | proprietary | |
| 3 | GPT-5.6 Terra OpenAI | 0.327 | — | 1.1M | $2.00 / $12.00 | proprietary | |
| 4 | GPT-5.6 Luna OpenAI | 0.320 | — | 1.1M | $0.20 / $1.20 | proprietary | |
| 5 | GPT OSS 120B OpenAI | 0.300 | 117B | 131K | — | open | |
| 6 | GPT-5.3 Chat OpenAI | 0.259 | — | 128K | $1.75 / $14.00 | proprietary | |
| 7 | GPT-5.5 Instant OpenAI | 0.229 | — | 400K | $5.00 / $30.00 | proprietary | |
| 8 | GPT OSS 20B OpenAI | 0.108 | 21B | 131K | — | open | |
| 9 | GPT-5 OpenAI | 0.016 | — | 400K | $1.25 / $10.00 | proprietary | |
Scores are rubric credit across the 1,000-conversation set, one response per conversation, graded by the benchmark's default grader (GPT-4.1). Prices are vendor list rates per 1M tokens, checked August 16, 2026. GPT-5.6 models charge higher rates above 272K input tokens. GPT OSS models are open weights without vendor list pricing. Every model name links through to its full result.
What the benchmark measures
HealthBench Hard keeps only the failures. OpenAI's HealthBench grades models on 5,000 realistic health conversations, from emergency triage to global health, against 48,562 rubric criteria written by 262 physicians across 60 countries. The Hard subset is the 1,000 conversations on which frontier models scored worst at release, which makes it the part of HealthBench that still separates current models. Two open-weights models, GPT OSS 120B and GPT OSS 20B, sit on the board alongside the closed frontier. The full construction is on the benchmark page.
How to read the scores
A score is rubric points earned over maximum positive points, clipped to 0 to 1. Low numbers are the design, not a flaw: the subset exists because models fail on it, the field's best result is 0.428, and 1.0 is nowhere in sight. Because the grader is itself a model, a number graded under any other configuration does not belong in this table; everything here shares one run protocol, laid out on the methodology page.
Head to head
The pairings worth deciding between get full pages; the compare page holds the complete score-difference matrix.
- 0.428 vs 0.331 · Muse Spark by 0.097
- 0.331 vs 0.327 · GPT-5.6 Sol by 0.004
- 0.320 vs 0.300 · GPT-5.6 Luna by 0.020
- 0.300 vs 0.108 · GPT OSS 120B by 0.192
- 0.331 vs 0.016 · GPT-5.6 Sol by 0.315
- 0.320 vs 0.259 · GPT-5.6 Luna by 0.061
Data
Everything on this page ships as JSON and CSV under CC BY 4.0, at URLs that never move. The data page has the citation format and version history.