HealthBench Professional

Sources

14 documents · 22 of 22 rows documented · updated September 8, 2026

This page lists the document behind each score on the leaderboard. Arcophos does not run HealthBench Professional; the numbers are copied from vendor system cards, official update reports, model cards, launch announcements, and named leaderboards, exactly as those documents print them. Entries are grouped by document. Each one gives the model, the score as printed, the excerpt the number was read from, where it sits in the document, who reported it, and a confidence label. Rows with no located document are listed at the end rather than left out. How the labels are assigned and why 2 documents can disagree is on the methodology page.

Claude Fable 5 and Claude Mythos 5 System Card

system card, first party · Anthropic · anthropic.com · published June 9, 2026 · retrieved September 7, 2026

  • Claude Fable 566.0primary source · vendor-reported · verified · length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 70.3%). Measured as Claude Mythos 5; Anthropic itself prints 66.0 in the Fable 5 column of the Opus 5 card with footnote Mythos 5.
    HealthBench Professional 66.0 - 64.7 56.9 51.8 -
    p. 252, Table 8.1.A, row HealthBench Professional, column Mythos 5 (Fable 5 column '-'); Figure 8.18.2.A p. 298 bar label 66.0% on Claude Mythos 5
  • Claude Opus 4.856.9conflicting
    HealthBench Professional 66.0 - 64.7 56.9 51.8 -
    p. 252, Table 8.1.A, column Opus 4.8

Claude Fable 5.1 and Claude Mythos 5.1 System Card

system card, first party · Anthropic · anthropic.com

  • Claude Fable 563.3%conflicting
    HealthBench Professional 62.1% 63.3% 59.8% –
    p. 167, Table 8.1.A, column 'Claude Fable 5/Mythos 5'
  • Claude Fable 5.162.1%primary source · vendor-reported · verified · length-adjusted (method published in the HealthBench Professional paper); Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over five trials, no tools or customized system prompt; Fable 5.1 run with safety classifiers active and a refusal-fallback to Claude Opus 5 (raw 74.2%).
    On HealthBench Professional, Claude Fable 5.1 achieved a raw score of 74.2%, ahead of Claude Opus 5 at 73.4%, Fable 5 at 68.9%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Fable 5.1 achieved a score of 62.1%.
    p. 199, sec. 8.17.2 (same figure printed as 62.1% in Table 8.1.A, p. 167)
  • Claude Opus 559.8%corroborating
    HealthBench Professional 62.1% 63.3% 59.8% –
    p. 167, Table 8.1.A, column 'Claude Opus 5'; raw 73.4% also reprinted on p. 199, sec. 8.17.2

GPT-5.5 Instant System Card

system card, first party · OpenAI · deploymentsafety.openai.com · published May 5, 2026 · retrieved September 7, 2026

  • GPT-5.5 Instant38.4primary source · vendor-reported · verified · length-adjusted (40.7 unadjusted, 2,775 mean response chars); GPT-5.5 Instant system card Table 5, column GPT-5.5 INSTANT.
    HealthBench Professional 37.6 (40.4, 2,973) 35.7 (38.3, 2,872) 32.9 (33.8, 2,285) 38.4 (40.7, 2,775)
    Section 4.1 HealthBench, Table 5 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5 INSTANT

GPT-5.6 - August Updates (system card addendum)

system card, first party · OpenAI · cdn.openai.com · published August 6, 2026 · retrieved September 7, 2026

  • GPT-5.6 Sol (August)54.0primary source · vendor-reported · verified · ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Sol (August) (56.6 unadjusted, 2,894 chars)
    HealthBench Professional 32.9 (33.8, 2,285) 38.4 (40.7, 2,775) 54.0 (56.6, 2,894) 44.1 (46.8, 2,920)
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Sol (August)
  • GPT-5.6 Luna (August)44.1primary source · vendor-reported · verified · ChatGPT production (Instant) deployment setting, length-adjusted, GPT-5.6 August Updates PDF column GPT-5.6 Luna (August) (46.8 unadjusted, 2,920 chars)
    HealthBench Professional 32.9 (33.8, 2,285) 38.4 (40.7, 2,775) 54.0 (56.6, 2,894) 44.1 (46.8, 2,920)
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.6 Luna (August)
  • GPT-5.5 Instant38.4corroborating
    HealthBench Professional 32.9 (33.8, 2,285) 38.4 (40.7, 2,775) 54.0 (56.6, 2,894) 44.1 (46.8, 2,920)
    p. 11, section 5.1 HealthBench, table "Reported as length-adjusted score (unadjusted, mean response length in characters)", column GPT-5.5 Instant

GPT-5.6 Preview System Card

system card, first party · OpenAI · deploymentsafety.openai.com

  • GPT-5.6 Sol60.5corroborating
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-SOL
  • GPT-5.6 Terra57.7corroborating
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-TERRA
  • GPT-5.6 Luna55.7corroborating
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.6-LUNA
  • GPT-5.551.8corroborating
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.5
  • GPT-5.448.1corroborating
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.4
  • GPT-546.2corroborating
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5
  • GPT-5.245.9corroborating
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.2
  • GPT-5.139.6corroborating
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)) (preview card, published 2026-06-26), column GPT-5.1

GPT-5.6 System Card

system card, first party · OpenAI · deploymentsafety.openai.com · published July 9, 2026 · retrieved September 7, 2026

  • GPT-5.6 Sol60.5primary source · vendor-reported · verified · length-adjusted, max reasoning effort (64.1 unadjusted, 3,228 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-SOL. API max-effort setting, not the August ChatGPT production setting measured in row 492.
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-SOL
  • GPT-5.6 Terra57.7primary source · vendor-reported · verified · length-adjusted, max reasoning effort (62.4 unadjusted, 3,618 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-TERRA.
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-TERRA
  • GPT-5.6 Luna55.7primary source · vendor-reported · verified · length-adjusted, max reasoning effort (59.8 unadjusted, 3,389 mean response chars); GPT-5.6 system card Table 6, column GPT-5.6-LUNA. API max-effort setting, not the August ChatGPT production setting measured in row 493.
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.6-LUNA
  • GPT-5.551.8primary source · vendor-reported · verified · length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.5 (57.2 unadjusted, 3818 chars)
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.5
  • GPT-5.448.1primary source · vendor-reported · verified · length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.4 (51.9 unadjusted, 3308 chars)
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.4
  • GPT-546.2primary source · vendor-reported · verified · length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5 (51.0 unadjusted, 3616 chars)
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5
  • GPT-5.245.9primary source · vendor-reported · verified · length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.2 (50.0 unadjusted, 3400 chars)
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.2
  • GPT-5.139.6primary source · vendor-reported · verified · length-adjusted, max reasoning effort, GPT-5.6 system card Table 6 column GPT-5.1 (48.0 unadjusted, 4863 chars)
    HealthBench Professional length-adjusted 46.2 (51.0, 3616) 39.6 (48.0, 4863) 45.9 (50.0, 3400) 48.1 (51.9, 3308) 51.8 (57.2, 3818) 60.5 (64.1, 3228) 57.7 (62.4, 3618) 55.7 (59.8, 3389)
    Section 5.1 HealthBench, Table 6 (reported as length-adjusted score (unadjusted, mean response length in characters)), column GPT-5.1

GPT-6 Astra System Card

system card, first party · OpenAI · deploymentsafety.openai.com · published September 3, 2026 · retrieved September 8, 2026

  • GPT-6 Astra63.4primary source · vendor-reported · verified · length-adjusted, max reasoning effort (69.5 unadjusted, 4,097 mean response chars); GPT-6 Astra system card Table 6, column 'gpt-6 Astra'.
    Astra has a length-adjusted HealthBench Professional score of 63.4 (+2.9 relative to GPT-5.6 Sol), HealthBench score of 58.1 (+1.1), HealthBench Hard score of 36.3 (+3.2), and HealthBench Consensus score of 95.8 (+0.3).
    p. 19, sec. 6.1 (Table 6 cell, column 'gpt-6 Astra': 63.4 (69.5, 4097))

GPT-6 Astra System Card - HealthBench (Deployment Safety Hub)

system card, first party · OpenAI · deploymentsafety.openai.com

  • GPT-6 Astra63.4mirror
    Astra has a length-adjusted HealthBench Professional score of 63.4 (+2.9 relative to GPT-5.6 Sol)
    section 6.1, HTML rendering of the same card

GPT-6 Astra: A new generation of intelligence

launch post, first party · OpenAI · openai.com

  • GPT-6 Astra63.4%corroborating
    | HealthBench Professional (length-adjusted) | 63.4% | 60.5% | 58.1% | 60.9% | 56.4% | 52.1% |
    Science and Health table

MAI-Thinking-1: Building a Hill-Climbing Machine

model card, first party · Microsoft AI · microsoft.ai · published August 12, 2026 · retrieved September 7, 2026

  • MAI-Thinking-135primary source · vendor-reported · verified · length-adjusted (HealthBench Professional length penalty), standard GPT-5.4 grader and OpenAI rubrics, Microsoft AI run; printed at integer precision.
    Model AIR-Bench CyberSec Instruct CyberSec Auto Long Fact Truthful QA HealthBench Prof. MedXpert QA MAI-Thinking-1 88 63 63 98 88 35 43 Sonnet 4.6 88 62 56 98 88 38 49
    p. 54, Table 12 'Post-trained model evaluation results on various public benchmarks', Health group, column HealthBench Prof.; protocol Appendix K.6 p. 106: HealthBench Professional introduces a length penalty for the primary metric, to correct for a well-observed correlation between lengthy responses and artificially increased LLM-grader scores. For all reported scores, we use the standard GPT-5.4 grader and rubrics provided by OpenAI.

Muse Spark 1.1 Evaluation Report

model card, first party · Meta · research.meta.ai · published July 9, 2026 · retrieved September 7, 2026

  • Muse Spark 1.159.3primary source · vendor-reported · verified · length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44)
    Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
    p. 101, Figure 44 'General capability benchmark results' (image), row HealthBench Professional, column Muse Spark 1.1; protocol p. 104 (printed 103): HealthBench Pro comprises 525 evaluation data points graded by rubrics. We use GPT-5.4 with low reasoning effort as the grader and report the length-normalized rubric score as done in their paper.
  • Muse Spark54.1primary source · vendor-reported · verified · length-normalized, GPT-5.4 low-reasoning grader; Muse Spark (1.0) column in the Muse Spark 1.1 Evaluation Report Figure 44
    Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
    p. 101, Figure 44 (image), row HealthBench Professional, column Muse Spark; protocol p. 104 (printed 103)

System Card: Claude Opus 4.8

system card, first party · Anthropic · anthropic.com · published May 28, 2026 · retrieved September 7, 2026

  • Claude Opus 4.855.8%primary source · vendor-reported · verified · length-adjusted, adaptive thinking at max effort, Claude Sonnet 4.6 grader (the grader used by the Opus 4.8 card itself); the other Claude rows on this board use the Claude Opus 4.8 grader, under which Anthropic later prints 57.4 for this model.
    Claude Opus 4.8 scores 55.8%, a meaningful improvement over Claude Opus 4.7 at 51.9% and Claude Sonnet 4.6 at 41.7%.
    p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229
  • Claude Opus 4.751.9%primary source · vendor-reported · verified · length-adjusted; adaptive thinking at max effort; Claude Sonnet 4.6 grader; 5 trials; comparison model in the Opus 4.8 card
    Claude Opus 4.8 scores 55.8%, a meaningful improvement over Claude Opus 4.7 at 51.9% and Claude Sonnet 4.6 at 41.7%.
    p. 228, section 8.14.1 HealthBench Professional; Figure 8.14.A p. 229
  • Claude Sonnet 4.641.7%conflicting
    Claude Opus 4.8 scores 55.8%, a meaningful improvement over Claude Opus 4.7 at 51.9% and Claude Sonnet 4.6 at 41.7%.
    p. 228, section 8.14.1

System Card: Claude Opus 5

system card, first party · Anthropic · anthropic.com

  • Claude Fable 566.0corroborating
    HealthBench Professional 59.8 57.4 66.011 60.5
    p. 152, Table 8.1.A, column Fable 5 (footnote 11: Mythos 5)
  • Claude Opus 559.8%primary source · vendor-reported · verified · length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 73.4%).
    Claude Opus 5 achieved a raw score of 73.4%, which is the highest amongst all Claude models, ahead of Claude Mythos 5 at 70.3%, Claude Opus 4.8 at 60.3%, and Claude Sonnet 5 at 62.4%. After length adjustment, which penalizes verbose model responses, Claude Opus 5 achieved a score of 59.8%.
    p. 189, section 8.15.2 HealthBench Professional results; also Table 8.1.A p. 152 ('HealthBench Professional 59.8 ...')
  • Claude Sonnet 557.8%corroborating
    Figure 8.15.2.A bar labels: Claude Sonnet 5 62.4% raw / 57.8% length-adjusted (prose: 'Claude Sonnet 5 at 62.4%')
    p. 189, section 8.15.2, Figure 8.15.2.A
  • Claude Opus 4.857.4conflicting
    HealthBench Professional 59.8 57.4 66.011 60.5
    p. 152, Table 8.1.A, column Opus 4.8

System Card: Claude Sonnet 5

system card, first party · Anthropic · anthropic.com · published June 30, 2026 · retrieved September 7, 2026

  • Claude Sonnet 557.8primary source · vendor-reported · verified · length-adjusted, Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt (raw 62.4%).
    HealthBench Professional 57.8 44.2 51.8 -
    p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 5; Figure 8.12.2.A p. 139
  • Claude Sonnet 4.644.2primary source · vendor-reported · verified · length-adjusted; Anthropic protocol: adaptive thinking at max effort, Claude Opus 4.8 grader, averaged over 5 trials, no tools or custom system prompt; comparison column in the Sonnet 5 card. Other Anthropic prints: 44.4% (Fable card Figure 8.18.2.A), 41.7% (Opus 4.8 card, Sonnet 4.6 grader)
    HealthBench Professional 57.8 44.2 51.8 -
    p. 115, Table 8.1.A, row HealthBench Professional, column Claude Sonnet 4.6

The same source fields are included in the JSON and CSV downloads on the data page.