Independent benchmark analysis / Original Nature study, corrected version of record
MultiMedQA
Seven question sources do not produce one clinical score.
MultiMedQA combines professional examinations, biomedical research questions and consumer health questions. We use the original Nature study to examine why those sources demand different evaluation endpoints. Multiple-choice answer accuracy is straightforward to count; a long answer can contain correct information, omit an important qualification and still receive mixed judgments. Our analysis concentrates on the consumer-question evaluation and the relationship between source inventories and the smaller human-rated sample. Published Med-PaLM and Flan-PaLM measurements remain historical study results. The coverage explorer is our interpretation of what each task format can and cannot support.
These are 140 selected questions, not the full sizes of the three source collections. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Task-specific accuracy and human ratings
Direction depends on the axis: consensus higher, judged potential harm lower.
Multiple-choice correctness and human judgments of long-form answers are separate endpoints.
Scoring definition
MCQ accuracy = correct answers / evaluated questions; human rating rate = responses assigned a category / responses rated for that axis
Do not average MCQ accuracy with harm or consensus percentages. Keep question subset, rater design and model configuration attached. [1]
03 / Measured evidence
Results, with their conditions attached.
Paper-reported results / selected rows
Original study: agreement with scientific consensus
Nature 2023 original study; 140-question long-form evaluation; one clinician rating per answer in this pilot.
Singhal et al.; Nature Medicine. Follow-up study distinguishes MultiMedQA 140 from expanded 1,066-question evaluation.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.