HealthBench Hard

Compare models

updated August 16, 2026

Matchups that decide deployments get a full page each, with the score gap, list prices, and a costed usage workload. The matrix below covers every remaining pairing.

Head-to-head pages

Score-difference matrix

Each cell is the row model's score minus the column model's, so positive means the row wins. Column numbers follow rank.

model123456789
1. Muse Spark·+0.097+0.101+0.108+0.128+0.169+0.199+0.320+0.412
2. GPT-5.6 Sol-0.097·+0.004+0.011+0.031+0.072+0.102+0.223+0.315
3. GPT-5.6 Terra-0.101-0.004·+0.007+0.027+0.068+0.098+0.219+0.311
4. GPT-5.6 Luna-0.108-0.011-0.007·+0.020+0.061+0.091+0.212+0.304
5. GPT OSS 120B-0.128-0.031-0.027-0.020·+0.041+0.071+0.192+0.284
6. GPT-5.3 Chat-0.169-0.072-0.068-0.061-0.041·+0.030+0.151+0.243
7. GPT-5.5 Instant-0.199-0.102-0.098-0.091-0.071-0.030·+0.121+0.213
8. GPT OSS 20B-0.320-0.223-0.219-0.212-0.192-0.151-0.121·+0.092
9. GPT-5-0.412-0.315-0.311-0.304-0.284-0.243-0.213-0.092·

On a 1,000-conversation set, anything within 0.01 sits inside plausible re-run variation; read those cells as ties.

Pricing and context windows for every model are on the models page.