GPT-5 on HealthBench Hard
rank 9 of 9 · updated August 16, 2026
On the 9-model HealthBench Hard board, GPT-5 holds rank 9 with a score of 0.016. OpenAI's August 2025 flagship, still served at unchanged prices but superseded by the GPT-5.6 series. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.
Score and API facts
| rank | 9 of 9 |
|---|---|
| score | 0.016 |
| lab | OpenAI |
| context window | 400K tokens |
| API price per 1M tokens | $1.25 in / $10.00 out |
| license | proprietary |
| released | 2025-08-07 |
Where it sits
Muse Spark tops the board at 0.428, which puts GPT-5 0.412 off the lead. One place up is GPT OSS 20B at 0.108. Every row on the board comes out of one run, so a gap between two models is measured on the same conversations under the same grader.
What does GPT-5 score on HealthBench Hard?
As of August 16, 2026, GPT-5 scores 0.016 on HealthBench Hard, 9 of 9 models on the board.
What does GPT-5 cost per million tokens?
OpenAI lists GPT-5 at $1.25 per million input tokens and $10.00 per million output tokens.
Head to head
The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.
- GPT-5 vs Muse Spark0.016 vs 0.428 · Muse Spark by 0.412
- 0.016 vs 0.331 · GPT-5.6 Sol by 0.315
- GPT-5 vs GPT-5.6 Terra0.016 vs 0.327 · GPT-5.6 Terra by 0.311
- GPT-5 vs GPT-5.6 Luna0.016 vs 0.320 · GPT-5.6 Luna by 0.304
- GPT-5 vs GPT OSS 120B0.016 vs 0.300 · GPT OSS 120B by 0.284
- GPT-5 vs GPT-5.3 Chat0.016 vs 0.259 · GPT-5.3 Chat by 0.243
- GPT-5 vs GPT-5.5 Instant0.016 vs 0.229 · GPT-5.5 Instant by 0.213
- GPT-5 vs GPT OSS 20B0.016 vs 0.108 · GPT OSS 20B by 0.092
How the conversations are graded is on the methodology page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.