Muse Spark 1.1 on HealthBench Professional
rank 6 of 22 · updated September 8, 2026
Muse Spark 1.1 scores 0.593 on HealthBench Professional, rank 6 of 22 models on the board. Meta's July 2026 update to Muse Spark, sold through the Meta Model API at $1.25 per million input tokens. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.
Result and API facts
| rank | 6 of 22 |
|---|---|
| score | 0.593 |
| lab | Meta |
| context window | 1.0M tokens |
| API price per 1M tokens | $1.25 in / $4.25 out |
| license | proprietary |
| source | Muse Spark 1.1 Evaluation Report (model card) |
| released | 2026-07-09 |
Position in the field
The gap to the leader, Claude Fable 5 at 0.660, is 0.067. Directly above sits Claude Opus 5 at 0.598. Directly below sits Claude Sonnet 5 at 0.578. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.
Source of this score
Read from Muse Spark 1.1 Evaluation Report (model card, Meta, 2026-07-09). Vendor-reported. Confidence: verified. Configuration: length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44).
Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
p. 101, Figure 44 'General capability benchmark results' (image), row HealthBench Professional, column Muse Spark 1.1; protocol p. 104 (printed 103): HealthBench Pro comprises 525 evaluation data points graded by rubrics. We use GPT-5.4 with low reasoning effort as the grader and report the length-normalized rubric score as done in their paper. · full entry on the sources page
What does Muse Spark 1.1 score on HealthBench Professional?
Muse Spark 1.1 scores 0.593 on HealthBench Professional, which places it at rank 6 of 22 models on the board as of September 8, 2026. The number was read from Muse Spark 1.1 Evaluation Report, listed on the sources page.
How much does Muse Spark 1.1 cost to run?
Muse Spark 1.1 is priced at $1.25 per million input tokens and $4.25 per million output tokens through Meta's API.
Head to head
Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.
- Muse Spark 1.1 vs Claude Fable 50.593 vs 0.660 · Claude Fable 5 by 0.067
- Muse Spark 1.1 vs GPT-6 Astra0.593 vs 0.634 · GPT-6 Astra by 0.041
- Muse Spark 1.1 vs Claude Fable 5.10.593 vs 0.621 · Claude Fable 5.1 by 0.028
- Muse Spark 1.1 vs GPT-5.6 Sol0.593 vs 0.605 · GPT-5.6 Sol by 0.012
- Muse Spark 1.1 vs Claude Opus 50.593 vs 0.598 · Claude Opus 5 by 0.005
- Muse Spark 1.1 vs Claude Sonnet 50.593 vs 0.578 · Muse Spark 1.1 by 0.015
- Muse Spark 1.1 vs GPT-5.6 Terra0.593 vs 0.577 · Muse Spark 1.1 by 0.016
- Muse Spark 1.1 vs Claude Opus 4.80.593 vs 0.558 · Muse Spark 1.1 by 0.035
- Muse Spark 1.1 vs GPT-5.6 Luna0.593 vs 0.557 · Muse Spark 1.1 by 0.036
- Muse Spark 1.1 vs Muse Spark0.593 vs 0.541 · Muse Spark 1.1 by 0.052
- Muse Spark 1.1 vs GPT-5.6 Sol (August)0.593 vs 0.540 · Muse Spark 1.1 by 0.053
- Muse Spark 1.1 vs Claude Opus 4.70.593 vs 0.519 · Muse Spark 1.1 by 0.074
- Muse Spark 1.1 vs GPT-5.50.593 vs 0.518 · Muse Spark 1.1 by 0.075
- Muse Spark 1.1 vs GPT-5.40.593 vs 0.481 · Muse Spark 1.1 by 0.112
- Muse Spark 1.1 vs GPT-50.593 vs 0.462 · Muse Spark 1.1 by 0.131
- Muse Spark 1.1 vs GPT-5.20.593 vs 0.459 · Muse Spark 1.1 by 0.134
- Muse Spark 1.1 vs Claude Sonnet 4.60.593 vs 0.442 · Muse Spark 1.1 by 0.151
- Muse Spark 1.1 vs GPT-5.6 Luna (August)0.593 vs 0.441 · Muse Spark 1.1 by 0.152
- Muse Spark 1.1 vs GPT-5.10.593 vs 0.396 · Muse Spark 1.1 by 0.197
- Muse Spark 1.1 vs GPT-5.5 Instant0.593 vs 0.384 · Muse Spark 1.1 by 0.209
- Muse Spark 1.1 vs MAI-Thinking-10.593 vs 0.350 · Muse Spark 1.1 by 0.243
Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.