Independent benchmark analysis / Original 2020 paper; USMLE four-option evaluation

MedQA · USMLE four-option

A medical answer-selection score with a specific denominator.

MedQA turns medical examination questions into a reproducible answer-selection task. This analysis centers the English USMLE four-option test set, because a benchmark name alone does not identify the answer space or evaluation split. The useful comparison is between systems given equivalent questions, options and retrieval resources. Our analysis separates those conditions, shows the resolution implied by the test denominator, and explains the gap between selecting an answer and executing a clinical workflow. Historical baseline results below are reported by the original paper; they are neither new Arcophos runs nor a current model ranking.

01 / What is being tested?

The task, before the score.

input
A medical examination question and candidate answer options; retrieved textbook evidence in the original OpenQA pipeline.
output
One selected answer option.
unit
Question
setting
English USMLE subset, four options, held-out test partition.

Data origin. Medical examination and preparation material collected from public websites; not longitudinal patient records. [1][2]

English questions
12,723

Full USMLE partition, across its three splits.

Table 3 [1]
Training questions
10,178

Official random training partition.

Table 3; README Data [1][2]
Development questions
1,272

Distinct from the 1,273-question test partition.

Table 3 [1]
Test questions
1,273

Denominator for the four-option USMLE results shown here.

Tables 3 and 8 [1]
Reported answer options
4

The author repository also distinguishes the four-option variant.

Table 3; README Data [1][2]
  1. 01

    Fix the partition

    Select the USMLE language subset and the four-option test partition.

    [1][2]
  2. 02

    Retrieve evidence

    Original retrieval baselines search a textbook collection before answer scoring.

    [1]
  3. 03

    Select the answer

    The reader or retrieval system assigns a choice from the supplied options.

    [1]
  4. 04

    Count agreement

    Compare the predicted option with the reference and report test accuracy.

    [1]

Dataset anatomy

The English USMLE partition

Training

English USMLE

10,178 questions[1]
Development

English USMLE

1,272 questions[1]
Test

English USMLE

1,273 questions[1]

Published split counts; these are questions, not patients. They sum to 12,723. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Answer-selection accuracy

Higher is better for this answer-selection task.

Correctly selected answers divided by evaluated questions, multiplied by 100.

Scoring definition

accuracy = 100 × correct answers / evaluated questions

Match language, split, four/five-option variant, retrieval, prompting and output parsing before comparing scores. [1][2]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Original paper: selected USMLE baselines

September 2020 paper, four-option USMLE test; original retrieval/reader protocols.

Answer-selection accuracy · %
050100
Reported
IR-CustomRetrieval baseline
36.1
BERT-Base-EnOriginal paper reader pipeline
34.3
BioBERT-BaseOriginal paper reader pipeline
34.1
BioBERT-LargeOriginal paper reader pipeline
36.7

Paper-reported historical results. No uncertainty intervals are provided in this table; do not present these as current frontier performance.

Source: Table 8 [1]

04 / Our original analysis

What follows from the design?

01

One question has a visible arithmetic effect.

Published evidence

The USMLE test partition contains 1,273 questions. [1]

Our interpretation

One additional correct answer changes accuracy by 100/1,273 ≈ 0.0786 percentage points. This is score resolution, not uncertainty.

02

A variant label is part of the measurement.

Published evidence

The repository separately provides four-option data. [2]

Our interpretation

A score without its option format leaves the task partly unspecified. A chance correction cannot reconstruct removed distractors.

03

A correct letter leaves other behavior unmeasured.

Published evidence

The scored output is the selected answer. [1]

Our interpretation

The score alone cannot determine whether accompanying prose contradicts that choice or whether an actionable clinical workflow was completed.

05 / Scope of the evidence

Where this benchmark stops.

Examination setting

Results concern prepared questions and reference options; they do not observe prospective care outcomes. [1]

Exposure remains unknown

A public random split does not establish that a later pretrained model never encountered the material. [2]

Pipeline dependence

Retrieval and answer extraction can change the measured system even when the base model name remains constant. [1]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Public author repository with data download links.
License
Author repository: MIT. Underlying source-material rights are not independently assessed here.
Conditions
Use the exact variant and split; consult source-specific terms before redistribution.
[2]

Evidence trail

Read the originals.

  1. MedQA: What Disease Does This Patient Have? ↗

    Di Jin and colleagues. Original MedQA definitions, split statistics and historical retrieval baselines.

  2. MedQA author repository ↗

    MedQA authors. Official data access, four-option variant and repository license.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗