Independent benchmark analysis / Original Nature study, corrected version of record

MultiMedQA

Seven question sources do not produce one clinical score.

MultiMedQA combines professional examinations, biomedical research questions and consumer health questions. We use the original Nature study to examine why those sources demand different evaluation endpoints. Multiple-choice answer accuracy is straightforward to count; a long answer can contain correct information, omit an important qualification and still receive mixed judgments. Our analysis concentrates on the consumer-question evaluation and the relationship between source inventories and the smaller human-rated sample. Published Med-PaLM and Flan-PaLM measurements remain historical study results. The coverage explorer is our interpretation of what each task format can and cannot support.

01 / What is being tested?

The task, before the score.

input
An examination question, a question plus research abstract, or a consumer health query.
output
A selected answer or a generated long-form response.
unit
One question; human evaluation uses a selected 140-question subset.
setting
Original 2023 paper; separate automated multiple-choice and human long-form evaluations.

Data origin. Seven components: MedQA, MedMCQA, PubMedQA, medical MMLU subjects, LiveQA, MedicationQA and HealthSearchQA. HealthSearchQA contains search-engine-surfaced consumer queries, not clinical encounters. [1][2][3]

Source datasets
7

Six existing components plus HealthSearchQA.

Abstract; Methods [1][2]
HealthSearchQA
3,173 questions

Version-of-record source inventory, not all human-rated.

Methods: HealthSearchQA [1]
Human-rated subset
140 questions

100 HealthSearchQA, 20 LiveQA and 20 MedicationQA.

Human evaluation results [1]
MedQA test set
1,273 questions

The original study’s exam test denominator.

Methods: MedQA [1]
Consumer evaluation
3 answer sources

Flan-PaLM 540B, Med-PaLM 540B and clinician answers.

Human evaluation results [1]
  1. 01

    Identify the question family

    Professional, research and consumer questions have different input and output contracts.

    [1]
  2. 02

    Generate answers

    The paper evaluates prompted model outputs; the consumer subset also receives clinician-authored answers.

    [1]
  3. 03

    Use the matching evaluation

    Answer keys score MCQs; clinicians and lay raters assess different dimensions of long-form answers.

    [1]
  4. 04

    Preserve the axis

    Report consensus, omissions and possible harm separately rather than collapsing them into a generic quality percentage.

    [1]

Dataset anatomy

The original human evaluation sample

HealthSearchQA sample

Selected for human evaluation.

100 questions[1]
LiveQA sample

Selected for human evaluation.

20 questions[1]
MedicationQA sample

Selected for human evaluation.

20 questions[1]

These are 140 selected questions, not the full sizes of the three source collections. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Task-specific accuracy and human ratings

Direction depends on the axis: consensus higher, judged potential harm lower.

Multiple-choice correctness and human judgments of long-form answers are separate endpoints.

Scoring definition

MCQ accuracy = correct answers / evaluated questions; human rating rate = responses assigned a category / responses rated for that axis

Do not average MCQ accuracy with harm or consensus percentages. Keep question subset, rater design and model configuration attached. [1]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Original study: agreement with scientific consensus

Nature 2023 original study; 140-question long-form evaluation; one clinician rating per answer in this pilot.

Responses judged aligned with consensus · %
050100
Reported
Clinician answersExpert-authored reference responses.
92.9%
Med-PaLM 540BInstruction prompt-tuned model.
92.6%
Flan-PaLM 540BInstruction-tuned baseline.
61.9%

Paper-reported percentages; bootstrap uncertainty is described in the paper. Similar point estimates do not establish clinical equivalence.

Source: Human evaluation results: Scientific consensus; Figure 4a [1]

Paper-reported results / selected rows

Original study: judged potential for harm

Same original consumer-question study; subjective hypothetical-harm ratings, not observed adverse events.

Responses judged potentially harmful · %
050100
Reported
Clinician answersLower is better for this axis.
5.7%
Med-PaLM 540BLower is better for this axis.
5.9%
Flan-PaLM 540BLower is better for this axis.
29.7%

Keep separate from consensus. These are evaluator judgments about answers, not measured patient harm rates.

Source: Possible extent and likelihood of harm; Figure 4d [1]

04 / Our original analysis

What follows from the design?

01

The inventory is larger than the rated evidence

Published evidence

HealthSearchQA contains 3,173 questions; 100 enter the original human-rated subset. [1]

Our interpretation

A large source collection does not make every reported human judgment a large-sample estimate. Track source size and evaluated sample separately.

02

The same answer can have mixed qualities

Published evidence

The human framework separately assesses incorrect content, missing content and other axes. [1]

Our interpretation

A single overall accuracy number erases error types that may matter differently to readers. Preserve a vector of results.

03

Similar averages are not equivalence

Published evidence

Clinician and Med-PaLM consensus point estimates are close in the original pilot. [1]

Our interpretation

Without an equivalence design and its uncertainty, this is not evidence that they can replace one another in practice.

04

Follow-up evaluations change the unit of comparison

Published evidence

The later Med-PaLM 2 study distinguishes 140-question and 1,066-question sets. [3]

Our interpretation

Model generations and sample expansions must remain separate in comparisons; a shared suite name does not fix the protocol.

05 / Scope of the evidence

Where this benchmark stops.

Question answering is a bounded task

These tasks do not directly evaluate actions in an EHR, actual patient dialogue or clinical outcomes. [1]

Pilot human ratings

The original long-form study uses a small selected question set and subjective rating axes. [1]

Historical public sources

Source overlap and changing scientific knowledge complicate interpretation across later model generations. [1]

Component licensing

No umbrella dataset license is established by the paper’s open-access article license. [1]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Published component datasets and HealthSearchQA supplement
License
Article CC BY 4.0; dataset licenses vary and must be checked separately.
Conditions
Original model code/weights were not released with the paper. We reproduce aggregate measurements, not question or patient content.
[1]

Evidence trail

Read the originals.

  1. Large language models encode clinical knowledge ↗

    Singhal, Azizi, Tu et al.; Nature. Version-of-record MultiMedQA definition and historical human evaluation.

  2. Large Language Models Encode Clinical Knowledge: publication record ↗

    Google Research. Official creator record explaining the seven-source suite and Med-PaLM study.

  3. Toward expert-level medical question answering with large language models ↗

    Singhal et al.; Nature Medicine. Follow-up study distinguishes MultiMedQA 140 from expanded 1,066-question evaluation.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗