Grok 4.6 on HealthBench Professional
rank 23 of 31 · updated September 30, 2026
Grok 4.6 scores 0.485 on HealthBench Professional, rank 23 of 31 models on the board. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.
Result and API facts
| rank | 23 of 31 |
|---|---|
| score | 0.485 |
| lab | SpaceX AI |
| context window | 500K tokens |
| API price per 1M tokens | $2.00 in / $6.00 out |
| license | proprietary |
| source | Introducing Grok 4.7 (launch post) |
| released | 2026-08-12 |
Position in the field
The gap to the leader, GPT-6 Astra (Anthropic run) at 0.703, is 0.218. Directly above sits GPT-5.5 at 0.518. Directly below sits GPT-5.4 at 0.481. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.
Source of this score
Read from Introducing Grok 4.7 (launch post, SpaceXAI, 2026-09-21). Vendor-reported. Confidence: verified. Configuration: SpaceXAI evaluation; High effort; HealthBench Professional.
Grok 4.7 xHigh Grok 4.6 High GPT-5.6 Sol Max Fable 5.1 Max Clinical reasoningHealthBench Professional 56.7% 48.5% 60.5% 62.1%
Model Improvements comparison table, Clinical reasoning / HealthBench Professional row; Grok 4.6 column; September 21, 2026. · full entry on the sources page
What does Grok 4.6 score on HealthBench Professional?
Grok 4.6 scores 0.485 on HealthBench Professional, which places it at rank 23 of 31 models on the board as of September 30, 2026. The number was read from Introducing Grok 4.7, listed on the sources page.
How much does Grok 4.6 cost to run?
Grok 4.6 is priced at $2.00 per million input tokens and $6.00 per million output tokens through SpaceX AI's API.
Head to head
Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.
- Grok 4.6 vs GPT-6 Astra (Anthropic run)0.485 vs 0.703 · GPT-6 Astra (Anthropic run) by 0.218
- Grok 4.6 vs Claude Sonnet 5.50.485 vs 0.692 · Claude Sonnet 5.5 by 0.207
- Grok 4.6 vs Claude Fable 50.485 vs 0.660 · Claude Fable 5 by 0.175
- Grok 4.6 vs Claude Opus 5.50.485 vs 0.656 · Claude Opus 5.5 by 0.171
- Grok 4.6 vs GPT-6 Astra0.485 vs 0.647 · GPT-6 Astra by 0.162
- Grok 4.6 vs Claude Fable 5 (September card)0.485 vs 0.633 · Claude Fable 5 (September card) by 0.148
- Grok 4.6 vs Claude Fable 5.10.485 vs 0.621 · Claude Fable 5.1 by 0.136
- Grok 4.6 vs GPT-6 Luna0.485 vs 0.608 · GPT-6 Luna by 0.123
- Grok 4.6 vs GPT-6 Sol0.485 vs 0.608 · GPT-6 Sol by 0.123
- Grok 4.6 vs GPT-5.6 Sol0.485 vs 0.605 · GPT-5.6 Sol by 0.120
- Grok 4.6 vs Claude Opus 50.485 vs 0.598 · Claude Opus 5 by 0.113
- Grok 4.6 vs Muse Spark 1.10.485 vs 0.593 · Muse Spark 1.1 by 0.108
- Grok 4.6 vs Claude Sonnet 50.485 vs 0.578 · Claude Sonnet 5 by 0.093
- Grok 4.6 vs GPT-5.6 Terra0.485 vs 0.577 · GPT-5.6 Terra by 0.092
- Grok 4.6 vs Claude Opus 4.8 (Opus 4.8 grader)0.485 vs 0.574 · Claude Opus 4.8 (Opus 4.8 grader) by 0.089
- Grok 4.6 vs Grok 4.70.485 vs 0.567 · Grok 4.7 by 0.082
- Grok 4.6 vs Claude Opus 4.80.485 vs 0.558 · Claude Opus 4.8 by 0.073
- Grok 4.6 vs GPT-5.6 Luna0.485 vs 0.557 · GPT-5.6 Luna by 0.072
- Grok 4.6 vs Muse Spark0.485 vs 0.541 · Muse Spark by 0.056
- Grok 4.6 vs GPT-5.6 Sol (August)0.485 vs 0.540 · GPT-5.6 Sol (August) by 0.055
- Grok 4.6 vs Claude Opus 4.70.485 vs 0.519 · Claude Opus 4.7 by 0.034
- Grok 4.6 vs GPT-5.50.485 vs 0.518 · GPT-5.5 by 0.033
- Grok 4.6 vs GPT-5.40.485 vs 0.481 · Grok 4.6 by 0.004
- Grok 4.6 vs GPT-50.485 vs 0.462 · Grok 4.6 by 0.023
- Grok 4.6 vs GPT-5.20.485 vs 0.459 · Grok 4.6 by 0.026
- Grok 4.6 vs Claude Sonnet 4.60.485 vs 0.442 · Grok 4.6 by 0.043
- Grok 4.6 vs GPT-5.6 Luna (August)0.485 vs 0.441 · Grok 4.6 by 0.044
- Grok 4.6 vs GPT-5.10.485 vs 0.396 · Grok 4.6 by 0.089
- Grok 4.6 vs GPT-5.5 Instant0.485 vs 0.384 · Grok 4.6 by 0.101
- Grok 4.6 vs MAI-Thinking-10.485 vs 0.350 · Grok 4.6 by 0.135
Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.