GPT OSS 20B on HealthBench Hard
rank 8 of 9 · updated August 16, 2026
On the 9-model HealthBench Hard board, GPT OSS 20B holds rank 8 with a score of 0.108. The smaller Apache 2.0 open-weights release, sized to run locally in 16GB of memory. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.
Score and API facts
| rank | 8 of 9 |
|---|---|
| score | 0.108 |
| lab | OpenAI |
| context window | 131K tokens |
| API price per 1M tokens | no published list pricing |
| license | open |
| parameters | 21B |
| released | 2025-08-05 |
Where it sits
Muse Spark tops the board at 0.428, which puts GPT OSS 20B 0.320 off the lead. One place up is GPT-5.5 Instant at 0.229. One place down is GPT-5 at 0.016. Every row on the board comes out of one run, so a gap between two models is measured on the same conversations under the same grader.
What does GPT OSS 20B score on HealthBench Hard?
As of August 16, 2026, GPT OSS 20B scores 0.108 on HealthBench Hard, 8 of 9 models on the board.
What does GPT OSS 20B cost per million tokens?
OpenAI publishes no list pricing for GPT OSS 20B.
Head to head
The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.
- GPT OSS 20B vs Muse Spark0.108 vs 0.428 · Muse Spark by 0.320
- GPT OSS 20B vs GPT-5.6 Sol0.108 vs 0.331 · GPT-5.6 Sol by 0.223
- GPT OSS 20B vs GPT-5.6 Terra0.108 vs 0.327 · GPT-5.6 Terra by 0.219
- GPT OSS 20B vs GPT-5.6 Luna0.108 vs 0.320 · GPT-5.6 Luna by 0.212
- 0.108 vs 0.300 · GPT OSS 120B by 0.192
- GPT OSS 20B vs GPT-5.3 Chat0.108 vs 0.259 · GPT-5.3 Chat by 0.151
- GPT OSS 20B vs GPT-5.5 Instant0.108 vs 0.229 · GPT-5.5 Instant by 0.121
- GPT OSS 20B vs GPT-50.108 vs 0.016 · GPT OSS 20B by 0.092
How the conversations are graded is on the methodology page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.