HealthBench Hard

The best AI for hard medical questions, August 2026

based on the HealthBench Hard leaderboard, updated August 16, 2026

Measured on the 1,000 health conversations frontier models handle worst, Muse Spark is the strongest model right now at 0.428, and the gap behind it is wide. The right pick still changes with the constraint you are under, so the table below answers the question four ways. The numbers come from the HealthBench Hard leaderboard, which this site maintains.

The short version

Highest scoreMuse Spark0.428, 0.097 ahead of everything else, at $1.25 / $4.25 per 1M tokens.
Best OpenAI modelGPT-5.6 Sol0.331, with GPT-5.6 Terra (0.327) close behind at less than half the price.
High volume on a budgetGPT-5.6 Luna0.320 at $0.20 / $1.20 per 1M tokens, within 0.011 of the best GPT-5.6 tier.
Local weightsGPT OSS 120B0.300 at rank 5, Apache 2.0, 117B parameters on a single 80GB GPU.

Why these picks

Two facts organize this board. First, Muse Spark's lead is real: 0.097 over the best GPT-5.6 model, a chasm next to the near-ties inside that tier. Second, the GPT-5.6 tier itself is close: 0.331, 0.327, and 0.320 span just 0.011, so inside OpenAI's lineup the honest recommendation is priced, not scored: GPT-5.6 Luna at a twenty-fifth of GPT-5.6 Sol's input rate gives up almost nothing on this benchmark.

The open-weights picture

GPT OSS 120B scores 0.300 at rank 5, above three of OpenAI's own closed models on this board, and its Apache 2.0 license makes it the default for health systems that keep records on-premises. GPT OSS 20B (0.108) trades most of that capability for a footprint that runs in 16GB of memory. Meta has said Muse Spark's weights will open later in 2026, which would move the top of this board into the open column.

What this page is not

Rubric adherence on hard, written conversations is all a benchmark score measures. Bedside judgment, latency, integration cost, and whether a vendor signs what your compliance team needs are outside it, and nothing here is medical advice. Before a single number drives a procurement call, read the methodology, and read the benchmark page for what the conversations contain.