{"publication":"Clinical Benchmark","url":"https://clinicalbenchmark.com","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"medqa","name":"MedQA · USMLE four-option","shortName":"MedQA","version":"Original 2020 paper; USMLE four-option evaluation","creators":"Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang and Peter Szolovits","paperDate":"2020-09-28","headline":"A medical answer-selection score with a specific denominator.","summary":"MedQA turns medical examination questions into a reproducible answer-selection task. This analysis centers the English USMLE four-option test set, because a benchmark name alone does not identify the answer space or evaluation split. The useful comparison is between systems given equivalent questions, options and retrieval resources. Our analysis separates those conditions, shows the resolution implied by the test denominator, and explains the gap between selecting an answer and executing a clinical workflow. Historical baseline results below are reported by the original paper; they are neither new Arcophos runs nor a current model ranking.","task":{"input":"A medical examination question and candidate answer options; retrieved textbook evidence in the original OpenQA pipeline.","output":"One selected answer option.","unit":"Question","setting":"English USMLE subset, four options, held-out test partition."},"dataOrigin":"Medical examination and preparation material collected from public websites; not longitudinal patient records.","facts":[{"label":"English questions","value":"12,723","detail":"Full USMLE partition, across its three splits.","sourceIds":["medqa-paper"],"locator":"Table 3"},{"label":"Training questions","value":"10,178","detail":"Official random training partition.","sourceIds":["medqa-paper","medqa-repo"],"locator":"Table 3; README Data"},{"label":"Development questions","value":"1,272","detail":"Distinct from the 1,273-question test partition.","sourceIds":["medqa-paper"],"locator":"Table 3"},{"label":"Test questions","value":"1,273","detail":"Denominator for the four-option USMLE results shown here.","sourceIds":["medqa-paper"],"locator":"Tables 3 and 8"},{"label":"Reported answer options","value":"4","detail":"The author repository also distinguishes the four-option variant.","sourceIds":["medqa-paper","medqa-repo"],"locator":"Table 3; README Data"}],"metric":{"name":"Answer-selection accuracy","description":"Correctly selected answers divided by evaluated questions, multiplied by 100.","formula":"accuracy = 100 × correct answers / evaluated questions","direction":"Higher is better for this answer-selection task.","comparability":"Match language, split, four/five-option variant, retrieval, prompting and output parsing before comparing scores.","sourceIds":["medqa-paper","medqa-repo"]},"workflow":[{"label":"Fix the partition","detail":"Select the USMLE language subset and the four-option test partition.","sourceIds":["medqa-paper","medqa-repo"]},{"label":"Retrieve evidence","detail":"Original retrieval baselines search a textbook collection before answer scoring.","sourceIds":["medqa-paper"]},{"label":"Select the answer","detail":"The reader or retrieval system assigns a choice from the supplied options.","sourceIds":["medqa-paper"]},{"label":"Count agreement","detail":"Compare the predicted option with the reference and report test accuracy.","sourceIds":["medqa-paper"]}],"slices":[{"label":"Training","value":10178,"unit":"questions","detail":"English USMLE","sourceIds":["medqa-paper"]},{"label":"Development","value":1272,"unit":"questions","detail":"English USMLE","sourceIds":["medqa-paper"]},{"label":"Test","value":1273,"unit":"questions","detail":"English USMLE","sourceIds":["medqa-paper"]}],"sliceTitle":"The English USMLE partition","sliceNote":"Published split counts; these are questions, not patients. They sum to 12,723.","results":[{"id":"medqa-original","title":"Original paper: selected USMLE baselines","metric":"Answer-selection accuracy","unit":"%","lower":0,"upper":100,"scope":"September 2020 paper, four-option USMLE test; original retrieval/reader protocols.","sourceIds":["medqa-paper"],"locator":"Table 8","rows":[{"label":"IR-Custom","value":36.1,"display":"36.1","detail":"Retrieval baseline"},{"label":"BERT-Base-En","value":34.3,"display":"34.3","detail":"Original paper reader pipeline"},{"label":"BioBERT-Base","value":34.1,"display":"34.1","detail":"Original paper reader pipeline"},{"label":"BioBERT-Large","value":36.7,"display":"36.7","detail":"Original paper reader pipeline"}],"note":"Paper-reported historical results. No uncertainty intervals are provided in this table; do not present these as current frontier performance."}],"analysis":[{"heading":"One question has a visible arithmetic effect.","evidence":"The USMLE test partition contains 1,273 questions.","interpretation":"One additional correct answer changes accuracy by 100/1,273 ≈ 0.0786 percentage points. This is score resolution, not uncertainty.","sourceIds":["medqa-paper"]},{"heading":"A variant label is part of the measurement.","evidence":"The repository separately provides four-option data.","interpretation":"A score without its option format leaves the task partly unspecified. A chance correction cannot reconstruct removed distractors.","sourceIds":["medqa-repo"]},{"heading":"A correct letter leaves other behavior unmeasured.","evidence":"The scored output is the selected answer.","interpretation":"The score alone cannot determine whether accompanying prose contradicts that choice or whether an actionable clinical workflow was completed.","sourceIds":["medqa-paper"]}],"limitations":[{"title":"Examination setting","detail":"Results concern prepared questions and reference options; they do not observe prospective care outcomes.","sourceIds":["medqa-paper"]},{"title":"Exposure remains unknown","detail":"A public random split does not establish that a later pretrained model never encountered the material.","sourceIds":["medqa-repo"]},{"title":"Pipeline dependence","detail":"Retrieval and answer extraction can change the measured system even when the base model name remains constant.","sourceIds":["medqa-paper"]}],"access":{"status":"Public author repository with data download links.","license":"Author repository: MIT. Underlying source-material rights are not independently assessed here.","restrictions":"Use the exact variant and split; consult source-specific terms before redistribution.","url":"https://github.com/jind11/MedQA","sourceIds":["medqa-repo"]},"sourceIds":["medqa-paper","medqa-repo"]},{"slug":"medmcqa","name":"MedMCQA","shortName":"MedMCQA","version":"CHIL 2022 paper; exact split counts from §3/Table 2","creators":"Ankit Pal, Logesh Kumar Umapathi and Malaikannan Sankarasubbu","paperDate":"2022-04-07","headline":"The split and the retrieval source belong beside the score.","summary":"MedMCQA is a medical examination question benchmark whose original evaluation deliberately separated examination sources. That makes split identity a central part of the result. The source documents also contain inconsistent development/test headings, so this page anchors exact counts to the paper statistics and makes the disagreement visible. The original baseline comparison provides a useful controlled story: the same named reader receives different context sources, and its answer-selection accuracy changes. We use that historical experiment to explain system comparability without implying that an examination score establishes clinical effectiveness.","task":{"input":"An English medical examination question with four options; optional retrieved context in the baseline study.","output":"The selected option index.","unit":"Question","setting":"Mock/test-series training material; NEET PG development and AIIMS PG test questions."},"dataOrigin":"Examination websites, books, mock tests and test series; image/table-dependent questions were filtered out.","facts":[{"label":"Training questions","value":"182,822","detail":"Exact paper statistics.","sourceIds":["medmcqa-paper"],"locator":"Section 3; Table 2"},{"label":"Validation questions","value":"4,183","detail":"Exact statistics; README headings and rounded prose conflict.","sourceIds":["medmcqa-paper","medmcqa-repo"],"locator":"Section 3; README Data Splits"},{"label":"Test questions","value":"6,150","detail":"The paper exact statistics distinguish this from validation.","sourceIds":["medmcqa-paper","medmcqa-repo"],"locator":"Section 3; Table 2"},{"label":"Questions across splits","value":"193,155","detail":"Exact total, rather than the abstract’s approximate 194k claim.","sourceIds":["medmcqa-paper"],"locator":"Section 2.3; Table 1"},{"label":"Subject categories","value":"21","detail":"Includes the Unknown category in the paper subject table.","sourceIds":["medmcqa-paper"],"locator":"Table 3"},{"label":"Option indices","value":"1–4","detail":"Author test-submission format; do not assume every derivative loader uses the same encoding.","sourceIds":["medmcqa-repo"],"locator":"Submission checklist"}],"metric":{"name":"Answer-selection accuracy","description":"The proportion of questions assigned their reference option. The paper table uses decimals; its discussion identifies 0.47 as 47% accuracy.","formula":"accuracy = 100 × correct option predictions / evaluated questions","direction":"Higher is better.","comparability":"Name validation versus test and the retrieval source. Author README headings disagree with the exact paper split statistics; record the actual file/version used.","sourceIds":["medmcqa-paper","medmcqa-repo"]},"workflow":[{"label":"Separate examination sources","detail":"Training uses mock/test-series questions; development and test use separate examination families.","sourceIds":["medmcqa-paper"]},{"label":"Select context","detail":"The paper compares no context, Wikipedia context and PubMed context.","sourceIds":["medmcqa-paper"]},{"label":"Fine-tune the reader","detail":"Each question/option pair is scored; the highest-validation checkpoint is selected.","sourceIds":["medmcqa-paper"]},{"label":"Evaluate held-out choices","detail":"Report option accuracy with its split and context setting.","sourceIds":["medmcqa-paper","medmcqa-repo"]}],"slices":[{"label":"Training","value":182822,"unit":"questions","detail":"Exact §3 count","sourceIds":["medmcqa-paper"]},{"label":"Validation","value":4183,"unit":"questions","detail":"Exact §3 count","sourceIds":["medmcqa-paper"]},{"label":"Test","value":6150,"unit":"questions","detail":"Exact §3 count","sourceIds":["medmcqa-paper"]}],"sliceTitle":"Exact split statistics","sliceNote":"The original README swaps validation/test table headings; use exact paper statistics and record the actual downloaded files.","results":[{"id":"medmcqa-context","title":"One reader, three context conditions","metric":"Test accuracy","unit":"%","lower":0,"upper":100,"scope":"CHIL 2022, PubMedBERT reader, original fine-tuning and test evaluation.","sourceIds":["medmcqa-paper"],"locator":"Table 4","rows":[{"label":"PubMedBERT · no context","value":41,"display":"41","detail":"Paper value 0.41"},{"label":"PubMedBERT · Wikipedia","value":42,"display":"42","detail":"Paper value 0.42"},{"label":"PubMedBERT · PubMed","value":47,"display":"47","detail":"Paper value 0.47"}],"note":"Paper-reported historical measurements. Table proportions are multiplied by 100 to display percentage accuracy; no intervals are supplied."}],"analysis":[{"heading":"A result needs a split receipt.","evidence":"Exact paper statistics and author README headings disagree.","interpretation":"Use split name, count, filename and revision together; a label alone can conceal a different denominator.","sourceIds":["medmcqa-paper","medmcqa-repo"]},{"heading":"Retrieval changes the system.","evidence":"PubMedBERT test accuracy is 41, 42 and 47 across the three context settings.","interpretation":"The six-point no-context-to-PubMed difference belongs to this pipeline comparison; it cannot be assigned to a new base model.","sourceIds":["medmcqa-paper"]},{"heading":"Breadth is not an outcome measure.","evidence":"The benchmark covers medical subjects but scores option selection.","interpretation":"Subject coverage and clinical action quality are different dimensions; a broad question bank does not observe decisions made in care.","sourceIds":["medmcqa-paper"]}],"limitations":[{"title":"Source-version ambiguity","detail":"The paper contains inconsistent rounded prose and the README reverses split-table labels. Exact statistics and file provenance must accompany a reproduction.","sourceIds":["medmcqa-paper","medmcqa-repo"]},{"title":"Historical results","detail":"The baseline study predates modern frontier models and is included for its experimental comparison, not its rank.","sourceIds":["medmcqa-paper"]},{"title":"Test-label access","detail":"The author repository says test ground truth is withheld and describes a prediction-submission process. A local validation score should be labeled accordingly.","sourceIds":["medmcqa-repo"]}],"access":{"status":"Public author data links; test submission instructions are documented.","license":"Author repository LICENSE.md: MIT. Third-party package labels may differ.","restrictions":"Verify downloaded-file terms and answer-index encoding; do not infer source-book rights from a repository license.","url":"https://github.com/medmcqa/medmcqa","sourceIds":["medmcqa-repo"]},"sourceIds":["medmcqa-paper","medmcqa-repo"]}],"explorer":{"kind":"accuracy-interval","title":"Inspect the accuracy denominator","intro":"Explore a Wilson 95% interval for a hypothetical answer-selection run. The initial denominator is the 1,273-question MedQA USMLE test set; 1,000 successes is an invented illustration, not a model result.","caution":"This binomial illustration assumes independent question outcomes. It does not account for contamination, question dependence or paired model comparisons, and it does not estimate clinical benefit.","sourceIds":["medqa-paper","wilson"],"rows":[],"parameters":[{"key":"n","value":1273,"label":"MedQA USMLE test questions","sourceIds":["medqa-paper"]},{"key":"successes","value":1000,"label":"Hypothetical correct answers","sourceIds":[]},{"key":"z","value":1.959963984540054,"label":"Two-sided 95% normal critical value","sourceIds":["wilson"]}]},"references":[{"id":"medqa-paper","title":"MedQA: What Disease Does This Patient Have?","organization":"Di Jin and colleagues","url":"https://arxiv.org/html/2009.13081v1","note":"Original MedQA definitions, split statistics and historical retrieval baselines.","locator":"Tables 3 and 8; dataset collection and Appendix A","version":"arXiv v1 · 2020-09-28"},{"id":"medqa-repo","title":"MedQA author repository","organization":"MedQA authors","url":"https://github.com/jind11/MedQA","note":"Official data access, four-option variant and repository license.","locator":"README Data; LICENSE","version":"Commit 27b02f66aac217933c9648a06f82e9f720377925"},{"id":"medmcqa-paper","title":"MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset","organization":"Ankit Pal, Logesh Kumar Umapathi and Malaikannan Sankarasubbu","url":"https://proceedings.mlr.press/v174/pal22a/pal22a.pdf","note":"Original exam-based construction, exact split statistics and retrieval comparison.","locator":"Sections 2–3; Tables 2 and 4","version":"CHIL 2022 · 2022-04-07"},{"id":"medmcqa-repo","title":"MedMCQA author repository","organization":"MedMCQA authors","url":"https://github.com/medmcqa/medmcqa","note":"Original access and submission instructions. README split-table headings disagree with exact paper statistics.","locator":"Data Splits; Model Submission; LICENSE.md","version":"Commit c59ef14ca1990266c4107c7864b45a20fd93e5e0"},{"id":"wilson","title":"Confidence intervals for a proportion","organization":"NIST/SEMATECH","url":"https://www.itl.nist.gov/div898/handbook/prc/section2/prc241.htm","note":"Wilson score interval formula; used only for the illustrative calculator.","locator":"Section 7.2.4.1","version":"Accessed 2026-09-28"}]}