{"publication":"Healthcare Benchmarks AI","url":"https://healthcarebenchmarks.ai","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"multimedqa","name":"MultiMedQA","shortName":"MultiMedQA","version":"Original Nature study, corrected version of record","creators":"Singhal, Azizi, Tu et al.; Google Research and collaborators","paperDate":"2023-07-12","headline":"Seven question sources do not produce one clinical score.","summary":"MultiMedQA combines professional examinations, biomedical research questions and consumer health questions. We use the original Nature study to examine why those sources demand different evaluation endpoints. Multiple-choice answer accuracy is straightforward to count; a long answer can contain correct information, omit an important qualification and still receive mixed judgments. Our analysis concentrates on the consumer-question evaluation and the relationship between source inventories and the smaller human-rated sample. Published Med-PaLM and Flan-PaLM measurements remain historical study results. The coverage explorer is our interpretation of what each task format can and cannot support.","task":{"input":"An examination question, a question plus research abstract, or a consumer health query.","output":"A selected answer or a generated long-form response.","unit":"One question; human evaluation uses a selected 140-question subset.","setting":"Original 2023 paper; separate automated multiple-choice and human long-form evaluations."},"dataOrigin":"Seven components: MedQA, MedMCQA, PubMedQA, medical MMLU subjects, LiveQA, MedicationQA and HealthSearchQA. HealthSearchQA contains search-engine-surfaced consumer queries, not clinical encounters.","facts":[{"label":"Source datasets","value":"7","detail":"Six existing components plus HealthSearchQA.","sourceIds":["mmqa","mmqa-google"],"locator":"Abstract; Methods"},{"label":"HealthSearchQA","value":"3,173 questions","detail":"Version-of-record source inventory, not all human-rated.","sourceIds":["mmqa"],"locator":"Methods: HealthSearchQA"},{"label":"Human-rated subset","value":"140 questions","detail":"100 HealthSearchQA, 20 LiveQA and 20 MedicationQA.","sourceIds":["mmqa"],"locator":"Human evaluation results"},{"label":"MedQA test set","value":"1,273 questions","detail":"The original study’s exam test denominator.","sourceIds":["mmqa"],"locator":"Methods: MedQA"},{"label":"Consumer evaluation","value":"3 answer sources","detail":"Flan-PaLM 540B, Med-PaLM 540B and clinician answers.","sourceIds":["mmqa"],"locator":"Human evaluation results"}],"metric":{"name":"Task-specific accuracy and human ratings","description":"Multiple-choice correctness and human judgments of long-form answers are separate endpoints.","formula":"MCQ accuracy = correct answers / evaluated questions; human rating rate = responses assigned a category / responses rated for that axis","direction":"Direction depends on the axis: consensus higher, judged potential harm lower.","comparability":"Do not average MCQ accuracy with harm or consensus percentages. Keep question subset, rater design and model configuration attached.","sourceIds":["mmqa"]},"workflow":[{"label":"Identify the question family","detail":"Professional, research and consumer questions have different input and output contracts.","sourceIds":["mmqa"]},{"label":"Generate answers","detail":"The paper evaluates prompted model outputs; the consumer subset also receives clinician-authored answers.","sourceIds":["mmqa"]},{"label":"Use the matching evaluation","detail":"Answer keys score MCQs; clinicians and lay raters assess different dimensions of long-form answers.","sourceIds":["mmqa"]},{"label":"Preserve the axis","detail":"Report consensus, omissions and possible harm separately rather than collapsing them into a generic quality percentage.","sourceIds":["mmqa"]}],"slices":[{"label":"HealthSearchQA sample","value":100,"unit":"questions","detail":"Selected for human evaluation.","sourceIds":["mmqa"]},{"label":"LiveQA sample","value":20,"unit":"questions","detail":"Selected for human evaluation.","sourceIds":["mmqa"]},{"label":"MedicationQA sample","value":20,"unit":"questions","detail":"Selected for human evaluation.","sourceIds":["mmqa"]}],"sliceTitle":"The original human evaluation sample","sliceNote":"These are 140 selected questions, not the full sizes of the three source collections.","results":[{"id":"mm-consensus","title":"Original study: agreement with scientific consensus","metric":"Responses judged aligned with consensus","unit":"%","lower":0,"upper":100,"scope":"Nature 2023 original study; 140-question long-form evaluation; one clinician rating per answer in this pilot.","sourceIds":["mmqa"],"locator":"Human evaluation results: Scientific consensus; Figure 4a","rows":[{"label":"Clinician answers","value":92.9,"display":"92.9%","detail":"Expert-authored reference responses."},{"label":"Med-PaLM 540B","value":92.6,"display":"92.6%","detail":"Instruction prompt-tuned model."},{"label":"Flan-PaLM 540B","value":61.9,"display":"61.9%","detail":"Instruction-tuned baseline."}],"note":"Paper-reported percentages; bootstrap uncertainty is described in the paper. Similar point estimates do not establish clinical equivalence."},{"id":"mm-harm","title":"Original study: judged potential for harm","metric":"Responses judged potentially harmful","unit":"%","lower":0,"upper":100,"scope":"Same original consumer-question study; subjective hypothetical-harm ratings, not observed adverse events.","sourceIds":["mmqa"],"locator":"Possible extent and likelihood of harm; Figure 4d","rows":[{"label":"Clinician answers","value":5.7,"display":"5.7%","detail":"Lower is better for this axis."},{"label":"Med-PaLM 540B","value":5.9,"display":"5.9%","detail":"Lower is better for this axis."},{"label":"Flan-PaLM 540B","value":29.7,"display":"29.7%","detail":"Lower is better for this axis."}],"note":"Keep separate from consensus. These are evaluator judgments about answers, not measured patient harm rates."}],"analysis":[{"heading":"The inventory is larger than the rated evidence","evidence":"HealthSearchQA contains 3,173 questions; 100 enter the original human-rated subset.","interpretation":"A large source collection does not make every reported human judgment a large-sample estimate. Track source size and evaluated sample separately.","sourceIds":["mmqa"]},{"heading":"The same answer can have mixed qualities","evidence":"The human framework separately assesses incorrect content, missing content and other axes.","interpretation":"A single overall accuracy number erases error types that may matter differently to readers. Preserve a vector of results.","sourceIds":["mmqa"]},{"heading":"Similar averages are not equivalence","evidence":"Clinician and Med-PaLM consensus point estimates are close in the original pilot.","interpretation":"Without an equivalence design and its uncertainty, this is not evidence that they can replace one another in practice.","sourceIds":["mmqa"]},{"heading":"Follow-up evaluations change the unit of comparison","evidence":"The later Med-PaLM 2 study distinguishes 140-question and 1,066-question sets.","interpretation":"Model generations and sample expansions must remain separate in comparisons; a shared suite name does not fix the protocol.","sourceIds":["mmqa-followup"]}],"limitations":[{"title":"Question answering is a bounded task","detail":"These tasks do not directly evaluate actions in an EHR, actual patient dialogue or clinical outcomes.","sourceIds":["mmqa"]},{"title":"Pilot human ratings","detail":"The original long-form study uses a small selected question set and subjective rating axes.","sourceIds":["mmqa"]},{"title":"Historical public sources","detail":"Source overlap and changing scientific knowledge complicate interpretation across later model generations.","sourceIds":["mmqa"]},{"title":"Component licensing","detail":"No umbrella dataset license is established by the paper’s open-access article license.","sourceIds":["mmqa"]}],"access":{"status":"Published component datasets and HealthSearchQA supplement","license":"Article CC BY 4.0; dataset licenses vary and must be checked separately.","restrictions":"Original model code/weights were not released with the paper. We reproduce aggregate measurements, not question or patient content.","url":"https://www.nature.com/articles/s41586-023-06291-2","sourceIds":["mmqa"]},"sourceIds":["mmqa","mmqa-google","mmqa-followup"]}],"explorer":{"kind":"coverage","title":"What kind of question is being answered?","intro":"A seven-component suite contains different evidence contracts. Compare the source, input, output and scoring boundary before comparing a percentage.","caution":"Rows describe the original MultiMedQA study. They do not merge its MCQ and human-rating endpoints.","sourceIds":["mmqa"],"parameters":[],"rows":[{"label":"MedQA","category":"Professional exams","input":"USMLE-style vignette and options","output":"Selected answer","metric":"MCQ accuracy","constraint":"A complete vignette supplies evidence the model did not gather.","benchmarkSlug":"multimedqa","sourceIds":["mmqa"]},{"label":"MedMCQA","category":"Professional exams","input":"Indian medical exam question and options","output":"Selected answer","metric":"MCQ accuracy","constraint":"Source jurisdiction and exam mix differ from MedQA.","benchmarkSlug":"multimedqa","sourceIds":["mmqa"]},{"label":"MMLU medical subjects","category":"Professional exams","input":"Subject-specific medical question and options","output":"Selected answer","metric":"Per-subject accuracy","constraint":"One suite component contains multiple subject subsets.","benchmarkSlug":"multimedqa","sourceIds":["mmqa"]},{"label":"PubMedQA","category":"Research evidence","input":"Research question and abstract context","output":"Yes, no or maybe","metric":"Answer accuracy","constraint":"Supporting abstract is supplied; retrieval is not evaluated.","benchmarkSlug":"multimedqa","sourceIds":["mmqa"]},{"label":"LiveQA","category":"Consumer questions","input":"Consumer query","output":"Long-form answer","metric":"Human evaluation on selected subset","constraint":"A source question count is not a rated-answer count.","benchmarkSlug":"multimedqa","sourceIds":["mmqa"]},{"label":"MedicationQA","category":"Consumer questions","input":"Consumer medication query","output":"Long-form answer","metric":"Human evaluation on selected subset","constraint":"Potential harm is a rater judgment, not an adverse-event outcome.","benchmarkSlug":"multimedqa","sourceIds":["mmqa"]},{"label":"HealthSearchQA","category":"Consumer questions","input":"Commonly searched health question","output":"Long-form answer","metric":"Human evaluation on selected subset","constraint":"Search queries do not establish a clinical diagnosis or patient cohort.","benchmarkSlug":"multimedqa","sourceIds":["mmqa"]}]},"references":[{"id":"mmqa","title":"Large language models encode clinical knowledge","organization":"Singhal, Azizi, Tu et al.; Nature","url":"https://www.nature.com/articles/s41586-023-06291-2","note":"Version-of-record MultiMedQA definition and historical human evaluation.","locator":"Human evaluation results; Figure 4; Methods: Datasets","version":"12 July 2023, corrected 27 July 2023"},{"id":"mmqa-google","title":"Large Language Models Encode Clinical Knowledge: publication record","organization":"Google Research","url":"https://research.google/pubs/large-language-models-encode-clinical-knowledge/","note":"Official creator record explaining the seven-source suite and Med-PaLM study.","locator":"Abstract and publication links","version":"2023 publication"},{"id":"mmqa-followup","title":"Toward expert-level medical question answering with large language models","organization":"Singhal et al.; Nature Medicine","url":"https://www.nature.com/articles/s41591-024-03423-7","note":"Follow-up study distinguishes MultiMedQA 140 from expanded 1,066-question evaluation.","locator":"Methods: Datasets; Independent evaluation","version":"2025 journal publication"}]}