# Healthcare Benchmarks AI > Independent MultiMedQA analysis covering seven datasets, human-rated consumer answers, sample denominators and the limits of medical question-answering scores. MultiMedQA is a suite of question sources, with different answer formats and different ways to judge success. We analyze the original Nature study’s exam, research and consumer tasks, with particular attention to its human-rated answer sample. The original contribution here is an explicit map between each source, its input and its scoring boundary. Historical Med-PaLM results remain labeled as published observations. Our guides explain why the suite name alone is insufficient to compare later evaluations. ## Provenance Independent analysis published by Arcophos. Benchmark creation belongs to the credited authors. Result rows are selected paper-reported measurements with their source versions and evaluation conditions, not new Arcophos runs or a live leaderboard. ## Benchmark dossiers - [MultiMedQA](https://healthcarebenchmarks.ai/benchmarks/multimedqa/): Seven question sources do not produce one clinical score. Source version: Original Nature study, corrected version of record. ## Original analyses - [What are the seven datasets in MultiMedQA?](https://healthcarebenchmarks.ai/guides/multimedqa-seven-datasets/): Map MultiMedQA’s examination, research and consumer questions to their answer formats and evaluation boundaries. - [What does MultiMedQA’s 140-question human evaluation show?](https://healthcarebenchmarks.ai/guides/multimedqa-human-evaluation-140/): Interpret the original consumer-answer study without confusing rater judgments, dataset size and patient outcomes. - [How can MultiMedQA results be compared across studies?](https://healthcarebenchmarks.ai/guides/multimedqa-compare-studies-and-versions/): Check component membership, question counts, answer generation and rater design before treating MultiMedQA scores as comparable. ## Inspect the evidence - [Evidence JSON](https://healthcarebenchmarks.ai/evidence.json): Task definitions, dataset facts, scoring rules, source-version results, our interpretations, and reference IDs. - [Sources](https://healthcarebenchmarks.ai/sources/): Original papers and repositories with evidence locators. - [Editorial method](https://healthcarebenchmarks.ai/methodology/): Source reconciliation and interpretation boundaries. - [About](https://healthcarebenchmarks.ai/about/): Ownership and corrections. Analysis updated: 2026-09-28