HealthBench Hard

GPT-5.6 Luna vs GPT-5.3 Chat on HealthBench Hard

updated August 16, 2026

A gap of 0.061 separates these two on the 1,000 hardest HealthBench conversations: GPT-5.6 Luna at 0.320, GPT-5.3 Chat at 0.259. The table adds what each one costs at list rates.

Side by side

OpenAI logoGPT-5.6 LunaOpenAI logoGPT-5.3 Chat
score0.3200.259
rank4 of 96 of 9
context window1.1M128K
price per 1M tokens, in / out$0.20 / $1.20$1.75 / $14.00
1,000-exchange workload$1.24$13.30
released2026-07-092026-03-05
licenseproprietaryproprietary

The workload row prices 1,000 exchanges of 2,000 input and 700 output tokens each, at the list rates current on August 16, 2026. GPT-5.6 models charge higher rates above 272K input tokens. GPT OSS models are open weights without vendor list pricing.

Reading the matchup

A generation cleanup: Luna scores 0.061 higher, costs about a ninth as much per input token, and GPT-5.3 Chat is already deprecated, with OpenAI pointing API users to the 5.6 series. The only reason to run 5.3 Chat on this workload is an integration that has not migrated yet.

Which scores higher on HealthBench Hard, GPT-5.6 Luna or GPT-5.3 Chat?

GPT-5.6 Luna. On the 1,000-conversation set it scores 0.320 to GPT-5.3 Chat's 0.259, a margin of 0.061 as of August 16, 2026.

Which is cheaper to run, GPT-5.6 Luna or GPT-5.3 Chat?

GPT-5.6 Luna. The same workload of 1,000 exchanges (2,000 input and 700 output tokens each) comes to $1.24 on it and $13.30 on the other.

Related comparisons

Each model's full page: GPT-5.6 Luna and GPT-5.3 Chat. Every other pairing lives in the score-difference matrix.