# Clinical Benchmark > Independent analysis of MedQA and MedMCQA: verified dataset splits, historical retrieval experiments, answer-selection metrics and an interactive accuracy interval calculator. MedQA and MedMCQA evaluate medical examination answer selection. Their scores depend on which questions, options and retrieval resources a system receives. This publication makes those conditions visible: exact denominators, source-version discrepancies, historical experiments and the limits of an accuracy claim. Start with the benchmark dossiers, inspect the illustrative interval calculator, then use the comparison worksheet to document a run. Arcophos provides the analysis and tools; the original research teams created the benchmarks. ## Provenance Independent analysis published by Arcophos. Benchmark creation belongs to the credited authors. Result rows are selected paper-reported measurements with their source versions and evaluation conditions, not new Arcophos runs or a live leaderboard. ## Benchmark dossiers - [MedQA · USMLE four-option](https://clinicalbenchmark.com/benchmarks/medqa/): A medical answer-selection score with a specific denominator. Source version: Original 2020 paper; USMLE four-option evaluation. - [MedMCQA](https://clinicalbenchmark.com/benchmarks/medmcqa/): The split and the retrieval source belong beside the score. Source version: CHIL 2022 paper; exact split counts from §3/Table 2. ## Original analyses - [MedQA vs MedMCQA: write the comparison contract first](https://clinicalbenchmark.com/guides/medqa-vs-medmcqa-comparison-contract/): Match split, options, retrieval and output parsing before interpreting medical examination accuracy. - [Why the MedMCQA split needs more than a label](https://clinicalbenchmark.com/guides/medmcqa-validation-test-split-discrepancy/): Resolve the original source disagreement using exact statistics, file identity and a transparent evaluation receipt. - [Read accuracy intervals without losing the system comparison](https://clinicalbenchmark.com/guides/accuracy-intervals-and-retrieval-effects/): Use the MedQA denominator and MedMCQA context experiment to separate score resolution, sampling uncertainty and pipeline changes. ## Inspect the evidence - [Evidence JSON](https://clinicalbenchmark.com/evidence.json): Task definitions, dataset facts, scoring rules, source-version results, our interpretations, and reference IDs. - [Sources](https://clinicalbenchmark.com/sources/): Original papers and repositories with evidence locators. - [Editorial method](https://clinicalbenchmark.com/methodology/): Source reconciliation and interpretation boundaries. - [About](https://clinicalbenchmark.com/about/): Ownership and corrections. Analysis updated: 2026-09-28