Muse Spark on HealthBench Hard
rank 1 of 9 · updated August 16, 2026
On the 9-model HealthBench Hard board, Muse Spark holds rank 1 with a score of 0.428. Meta's frontier multimodal reasoning model, sold through the Meta Model API, with weights announced to open later in 2026. HealthBench Hard tests models on the 1,000 health conversations the frontier found hardest, with each response judged against its conversation's physician-written rubric and scored between 0 and 1.
Score and API facts
| rank | 1 of 9 |
|---|---|
| score | 0.428 |
| lab | Meta |
| context window | 1.0M tokens |
| API price per 1M tokens | $1.25 in / $4.25 out |
| license | proprietary |
| released | 2026-08-05 |
Where it sits
Nothing ranks above it: Muse Spark holds first place with 0.097 of clear air over GPT-5.6 Sol. Every row on the board comes out of one run, so a gap between two models is measured on the same conversations under the same grader.
What does Muse Spark score on HealthBench Hard?
As of August 16, 2026, Muse Spark scores 0.428 on HealthBench Hard, 1 of 9 models on the board.
What does Muse Spark cost per million tokens?
Meta lists Muse Spark at $1.25 per million input tokens and $4.25 per million output tokens.
Head to head
The pairings that earned a full page are linked below; the rest of the differences live in the score-difference matrix.
- 0.428 vs 0.331 · Muse Spark by 0.097
- Muse Spark vs GPT-5.6 Terra0.428 vs 0.327 · Muse Spark by 0.101
- Muse Spark vs GPT-5.6 Luna0.428 vs 0.320 · Muse Spark by 0.108
- Muse Spark vs GPT OSS 120B0.428 vs 0.300 · Muse Spark by 0.128
- Muse Spark vs GPT-5.3 Chat0.428 vs 0.259 · Muse Spark by 0.169
- Muse Spark vs GPT-5.5 Instant0.428 vs 0.229 · Muse Spark by 0.199
- Muse Spark vs GPT OSS 20B0.428 vs 0.108 · Muse Spark by 0.320
- Muse Spark vs GPT-50.428 vs 0.016 · Muse Spark by 0.412
How the conversations are graded is on the methodology page, and how the subset was selected is on the benchmark page. The full ranking is on the leaderboard.