One question has a visible arithmetic effect.
The USMLE test partition contains 1,273 questions. [1]
One additional correct answer changes accuracy by 100/1,273 ≈ 0.0786 percentage points. This is score resolution, not uncertainty.
Independent benchmark analysis / Original 2020 paper; USMLE four-option evaluation
MedQA turns medical examination questions into a reproducible answer-selection task. This analysis centers the English USMLE four-option test set, because a benchmark name alone does not identify the answer space or evaluation split. The useful comparison is between systems given equivalent questions, options and retrieval resources. Our analysis separates those conditions, shows the resolution implied by the test denominator, and explains the gap between selecting an answer and executing a clinical workflow. Historical baseline results below are reported by the original paper; they are neither new Arcophos runs nor a current model ranking.
01 / What is being tested?
Data origin. Medical examination and preparation material collected from public websites; not longitudinal patient records. [1][2]
Full USMLE partition, across its three splits.
Table 3 [1]Distinct from the 1,273-question test partition.
Table 3 [1]Denominator for the four-option USMLE results shown here.
Tables 3 and 8 [1]Select the USMLE language subset and the four-option test partition.
[1][2]Original retrieval baselines search a textbook collection before answer scoring.
[1]The reader or retrieval system assigns a choice from the supplied options.
[1]Compare the predicted option with the reference and report test accuracy.
[1]Dataset anatomy
English USMLE
English USMLE
English USMLE
Published split counts; these are questions, not patients. They sum to 12,723. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Higher is better for this answer-selection task.
03 / Measured evidence
Paper-reported results / selected rows
September 2020 paper, four-option USMLE test; original retrieval/reader protocols.
Paper-reported historical results. No uncertainty intervals are provided in this table; do not present these as current frontier performance.
Source: Table 8 [1]
04 / Our original analysis
The USMLE test partition contains 1,273 questions. [1]
One additional correct answer changes accuracy by 100/1,273 ≈ 0.0786 percentage points. This is score resolution, not uncertainty.
The repository separately provides four-option data. [2]
A score without its option format leaves the task partly unspecified. A chance correction cannot reconstruct removed distractors.
The scored output is the selected answer. [1]
The score alone cannot determine whether accompanying prose contradicts that choice or whether an actionable clinical workflow was completed.
05 / Scope of the evidence
Results concern prepared questions and reference options; they do not observe prospective care outcomes. [1]
A public random split does not establish that a later pretrained model never encountered the material. [2]
Retrieval and answer extraction can change the measured system even when the base model name remains constant. [1]
Evidence trail
Di Jin and colleagues. Original MedQA definitions, split statistics and historical retrieval baselines.
MedQA authors. Official data access, four-option variant and repository license.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.