Grok 4.7 on HealthBench Professional
rank 16 of 31 · updated September 30, 2026
Grok 4.7 scores 0.567 on HealthBench Professional, rank 16 of 31 models on the board. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.
Result and API facts
| rank | 16 of 31 |
|---|---|
| score | 0.567 |
| lab | SpaceX AI |
| context window | 500K tokens |
| API price per 1M tokens | $2.00 in / $6.00 out |
| license | proprietary |
| source | Introducing Grok 4.7 (launch post) |
| released | 2026-09-21 |
Position in the field
The gap to the leader, GPT-6 Astra (Anthropic run) at 0.703, is 0.136. Directly above sits Claude Opus 4.8 (Opus 4.8 grader) at 0.574. Directly below sits Claude Opus 4.8 at 0.558. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.
Source of this score
Read from Introducing Grok 4.7 (launch post, SpaceXAI, 2026-09-21). Vendor-reported. Confidence: verified. Configuration: SpaceXAI evaluation; xHigh effort; HealthBench Professional.
Grok 4.7 xHigh Grok 4.6 High GPT-5.6 Sol Max Fable 5.1 Max Clinical reasoningHealthBench Professional 56.7% 48.5% 60.5% 62.1%
Model Improvements comparison table, Clinical reasoning / HealthBench Professional row; Grok 4.7 column; September 21, 2026. · full entry on the sources page
What does Grok 4.7 score on HealthBench Professional?
Grok 4.7 scores 0.567 on HealthBench Professional, which places it at rank 16 of 31 models on the board as of September 30, 2026. The number was read from Introducing Grok 4.7, listed on the sources page.
How much does Grok 4.7 cost to run?
Grok 4.7 is priced at $2.00 per million input tokens and $6.00 per million output tokens through SpaceX AI's API.
Head to head
Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.
- Grok 4.7 vs GPT-6 Astra (Anthropic run)0.567 vs 0.703 · GPT-6 Astra (Anthropic run) by 0.136
- Grok 4.7 vs Claude Sonnet 5.50.567 vs 0.692 · Claude Sonnet 5.5 by 0.125
- Grok 4.7 vs Claude Fable 50.567 vs 0.660 · Claude Fable 5 by 0.093
- Grok 4.7 vs Claude Opus 5.50.567 vs 0.656 · Claude Opus 5.5 by 0.089
- Grok 4.7 vs GPT-6 Astra0.567 vs 0.647 · GPT-6 Astra by 0.080
- Grok 4.7 vs Claude Fable 5 (September card)0.567 vs 0.633 · Claude Fable 5 (September card) by 0.066
- Grok 4.7 vs Claude Fable 5.10.567 vs 0.621 · Claude Fable 5.1 by 0.054
- Grok 4.7 vs GPT-6 Luna0.567 vs 0.608 · GPT-6 Luna by 0.041
- Grok 4.7 vs GPT-6 Sol0.567 vs 0.608 · GPT-6 Sol by 0.041
- Grok 4.7 vs GPT-5.6 Sol0.567 vs 0.605 · GPT-5.6 Sol by 0.038
- Grok 4.7 vs Claude Opus 50.567 vs 0.598 · Claude Opus 5 by 0.031
- Grok 4.7 vs Muse Spark 1.10.567 vs 0.593 · Muse Spark 1.1 by 0.026
- Grok 4.7 vs Claude Sonnet 50.567 vs 0.578 · Claude Sonnet 5 by 0.011
- Grok 4.7 vs GPT-5.6 Terra0.567 vs 0.577 · GPT-5.6 Terra by 0.010
- Grok 4.7 vs Claude Opus 4.8 (Opus 4.8 grader)0.567 vs 0.574 · Claude Opus 4.8 (Opus 4.8 grader) by 0.007
- Grok 4.7 vs Claude Opus 4.80.567 vs 0.558 · Grok 4.7 by 0.009
- Grok 4.7 vs GPT-5.6 Luna0.567 vs 0.557 · Grok 4.7 by 0.010
- Grok 4.7 vs Muse Spark0.567 vs 0.541 · Grok 4.7 by 0.026
- Grok 4.7 vs GPT-5.6 Sol (August)0.567 vs 0.540 · Grok 4.7 by 0.027
- Grok 4.7 vs Claude Opus 4.70.567 vs 0.519 · Grok 4.7 by 0.048
- Grok 4.7 vs GPT-5.50.567 vs 0.518 · Grok 4.7 by 0.049
- Grok 4.7 vs Grok 4.60.567 vs 0.485 · Grok 4.7 by 0.082
- Grok 4.7 vs GPT-5.40.567 vs 0.481 · Grok 4.7 by 0.086
- Grok 4.7 vs GPT-50.567 vs 0.462 · Grok 4.7 by 0.105
- Grok 4.7 vs GPT-5.20.567 vs 0.459 · Grok 4.7 by 0.108
- Grok 4.7 vs Claude Sonnet 4.60.567 vs 0.442 · Grok 4.7 by 0.125
- Grok 4.7 vs GPT-5.6 Luna (August)0.567 vs 0.441 · Grok 4.7 by 0.126
- Grok 4.7 vs GPT-5.10.567 vs 0.396 · Grok 4.7 by 0.171
- Grok 4.7 vs GPT-5.5 Instant0.567 vs 0.384 · Grok 4.7 by 0.183
- Grok 4.7 vs MAI-Thinking-10.567 vs 0.350 · Grok 4.7 by 0.217
Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.