Independent benchmark analysis / CHIL 2022 paper; exact split counts from §3/Table 2
MedMCQA
The split and the retrieval source belong beside the score.
MedMCQA is a medical examination question benchmark whose original evaluation deliberately separated examination sources. That makes split identity a central part of the result. The source documents also contain inconsistent development/test headings, so this page anchors exact counts to the paper statistics and makes the disagreement visible. The original baseline comparison provides a useful controlled story: the same named reader receives different context sources, and its answer-selection accuracy changes. We use that historical experiment to explain system comparability without implying that an examination score establishes clinical effectiveness.
01 / What is being tested?
The task, before the score.
- input
- An English medical examination question with four options; optional retrieved context in the baseline study.
- output
- The selected option index.
- unit
- Question
- setting
- Mock/test-series training material; NEET PG development and AIIMS PG test questions.
Data origin. Examination websites, books, mock tests and test series; image/table-dependent questions were filtered out. [3][4]
- Training questions
- 182,822
Exact paper statistics.
Section 3; Table 2 [3] - Validation questions
- 4,183
Exact statistics; README headings and rounded prose conflict.
Section 3; README Data Splits [3][4] - Test questions
- 6,150
The paper exact statistics distinguish this from validation.
Section 3; Table 2 [3][4] - Questions across splits
- 193,155
Exact total, rather than the abstract’s approximate 194k claim.
Section 2.3; Table 1 [3] - Subject categories
- 21
Includes the Unknown category in the paper subject table.
Table 3 [3] - Option indices
- 1–4
Author test-submission format; do not assume every derivative loader uses the same encoding.
Submission checklist [4]
- 01
Separate examination sources
Training uses mock/test-series questions; development and test use separate examination families.
[3] - 02
Select context
The paper compares no context, Wikipedia context and PubMed context.
[3] - 03
Fine-tune the reader
Each question/option pair is scored; the highest-validation checkpoint is selected.
[3] - 04
Evaluate held-out choices
Report option accuracy with its split and context setting.
[3][4]
Dataset anatomy
Exact split statistics
Exact §3 count
Exact §3 count
Exact §3 count
The original README swaps validation/test table headings; use exact paper statistics and record the actual downloaded files. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Answer-selection accuracy
Higher is better.
The proportion of questions assigned their reference option. The paper table uses decimals; its discussion identifies 0.47 as 47% accuracy.
accuracy = 100 × correct option predictions / evaluated questions
Name validation versus test and the retrieval source. Author README headings disagree with the exact paper split statistics; record the actual file/version used. [3][4]
03 / Measured evidence
Results, with their conditions attached.
Paper-reported results / selected rows
One reader, three context conditions
CHIL 2022, PubMedBERT reader, original fine-tuning and test evaluation.
Paper-reported historical measurements. Table proportions are multiplied by 100 to display percentage accuracy; no intervals are supplied.
Source: Table 4 [3]
04 / Our original analysis
What follows from the design?
Retrieval changes the system.
PubMedBERT test accuracy is 41, 42 and 47 across the three context settings. [3]
The six-point no-context-to-PubMed difference belongs to this pipeline comparison; it cannot be assigned to a new base model.
Breadth is not an outcome measure.
The benchmark covers medical subjects but scores option selection. [3]
Subject coverage and clinical action quality are different dimensions; a broad question bank does not observe decisions made in care.
05 / Scope of the evidence
Where this benchmark stops.
Source-version ambiguity
The paper contains inconsistent rounded prose and the README reverses split-table labels. Exact statistics and file provenance must accompany a reproduction. [3][4]
Historical results
The baseline study predates modern frontier models and is included for its experimental comparison, not its rank. [3]
Test-label access
The author repository says test ground truth is withheld and describes a prediction-submission process. A local validation score should be labeled accordingly. [4]
- Availability
- Public author data links; test submission instructions are documented.
- License
- Author repository LICENSE.md: MIT. Third-party package labels may differ.
- Conditions
- Verify downloaded-file terms and answer-index encoding; do not infer source-book rights from a repository license.
Evidence trail
Read the originals.
- MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset ↗
Ankit Pal, Logesh Kumar Umapathi and Malaikannan Sankarasubbu. Original exam-based construction, exact split statistics and retrieval comparison.
- MedMCQA author repository ↗
MedMCQA authors. Original access and submission instructions. README split-table headings disagree with exact paper statistics.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.