HealthBench Hard

OpenAI logoGPT OSS 120B on HealthBench Hard

rank 5 of 9 · updated August 16, 2026

On the 9-model HealthBench Hard board, GPT OSS 120B holds rank 5 with a score of 0.300. OpenAI's most capable open-weights model, an Apache 2.0 mixture-of-experts release that fits on a single 80GB GPU. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.

Score and API facts

rank5 of 9
score0.300
labOpenAI
context window131K tokens
API price per 1M tokensno published list pricing
licenseopen
parameters117B
released2025-08-05

Where it sits

Muse Spark tops the board at 0.428, which puts GPT OSS 120B 0.128 off the lead. One place up is GPT-5.6 Luna at 0.320. One place down is GPT-5.3 Chat at 0.259. Every row on the board comes out of one run, so a gap between two models is measured on the same conversations under the same grader.

What does GPT OSS 120B score on HealthBench Hard?

As of August 16, 2026, GPT OSS 120B scores 0.300 on HealthBench Hard, 5 of 9 models on the board.

What does GPT OSS 120B cost per million tokens?

OpenAI publishes no list pricing for GPT OSS 120B.

Head to head

The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.

How the conversations are graded is on the methodology page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.