HealthBench Professional

SpaceX AI logoGrok 4.6 on HealthBench Professional

rank 23 of 31 · updated September 30, 2026

Grok 4.6 scores 0.485 on HealthBench Professional, rank 23 of 31 models on the board. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.

Result and API facts

rank23 of 31
score0.485
labSpaceX AI
context window500K tokens
API price per 1M tokens$2.00 in / $6.00 out
licenseproprietary
sourceIntroducing Grok 4.7 (launch post)
released2026-08-12

Position in the field

The gap to the leader, GPT-6 Astra (Anthropic run) at 0.703, is 0.218. Directly above sits GPT-5.5 at 0.518. Directly below sits GPT-5.4 at 0.481. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.

Source of this score

Read from Introducing Grok 4.7 (launch post, SpaceXAI, 2026-09-21). Vendor-reported. Confidence: verified. Configuration: SpaceXAI evaluation; High effort; HealthBench Professional.

Grok 4.7 xHigh Grok 4.6 High GPT-5.6 Sol Max Fable 5.1 Max Clinical reasoningHealthBench Professional 56.7% 48.5% 60.5% 62.1%

Model Improvements comparison table, Clinical reasoning / HealthBench Professional row; Grok 4.6 column; September 21, 2026. · full entry on the sources page

What does Grok 4.6 score on HealthBench Professional?

Grok 4.6 scores 0.485 on HealthBench Professional, which places it at rank 23 of 31 models on the board as of September 30, 2026. The number was read from Introducing Grok 4.7, listed on the sources page.

How much does Grok 4.6 cost to run?

Grok 4.6 is priced at $2.00 per million input tokens and $6.00 per million output tokens through SpaceX AI's API.

Head to head

Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.

Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.