Independent benchmark analysis / CHIL 2022 paper; exact split counts from §3/Table 2

MedMCQA

The split and the retrieval source belong beside the score.

MedMCQA is a medical examination question benchmark whose original evaluation deliberately separated examination sources. That makes split identity a central part of the result. The source documents also contain inconsistent development/test headings, so this page anchors exact counts to the paper statistics and makes the disagreement visible. The original baseline comparison provides a useful controlled story: the same named reader receives different context sources, and its answer-selection accuracy changes. We use that historical experiment to explain system comparability without implying that an examination score establishes clinical effectiveness.

01 / What is being tested?

The task, before the score.

input
An English medical examination question with four options; optional retrieved context in the baseline study.
output
The selected option index.
unit
Question
setting
Mock/test-series training material; NEET PG development and AIIMS PG test questions.

Data origin. Examination websites, books, mock tests and test series; image/table-dependent questions were filtered out. [3][4]

Training questions
182,822

Exact paper statistics.

Section 3; Table 2 [3]
Validation questions
4,183

Exact statistics; README headings and rounded prose conflict.

Section 3; README Data Splits [3][4]
Test questions
6,150

The paper exact statistics distinguish this from validation.

Section 3; Table 2 [3][4]
Questions across splits
193,155

Exact total, rather than the abstract’s approximate 194k claim.

Section 2.3; Table 1 [3]
Subject categories
21

Includes the Unknown category in the paper subject table.

Table 3 [3]
Option indices
1–4

Author test-submission format; do not assume every derivative loader uses the same encoding.

Submission checklist [4]
  1. 01

    Separate examination sources

    Training uses mock/test-series questions; development and test use separate examination families.

    [3]
  2. 02

    Select context

    The paper compares no context, Wikipedia context and PubMed context.

    [3]
  3. 03

    Fine-tune the reader

    Each question/option pair is scored; the highest-validation checkpoint is selected.

    [3]
  4. 04

    Evaluate held-out choices

    Report option accuracy with its split and context setting.

    [3][4]

Dataset anatomy

Exact split statistics

Training

Exact §3 count

182,822 questions[3]
Validation

Exact §3 count

4,183 questions[3]
Test

Exact §3 count

6,150 questions[3]

The original README swaps validation/test table headings; use exact paper statistics and record the actual downloaded files. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Answer-selection accuracy

Higher is better.

The proportion of questions assigned their reference option. The paper table uses decimals; its discussion identifies 0.47 as 47% accuracy.

Scoring definition

accuracy = 100 × correct option predictions / evaluated questions

Name validation versus test and the retrieval source. Author README headings disagree with the exact paper split statistics; record the actual file/version used. [3][4]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

One reader, three context conditions

CHIL 2022, PubMedBERT reader, original fine-tuning and test evaluation.

Test accuracy · %
050100
Reported
PubMedBERT · no contextPaper value 0.41
41
PubMedBERT · WikipediaPaper value 0.42
42
PubMedBERT · PubMedPaper value 0.47
47

Paper-reported historical measurements. Table proportions are multiplied by 100 to display percentage accuracy; no intervals are supplied.

Source: Table 4 [3]

04 / Our original analysis

What follows from the design?

01

A result needs a split receipt.

Published evidence

Exact paper statistics and author README headings disagree. [3][4]

Our interpretation

Use split name, count, filename and revision together; a label alone can conceal a different denominator.

02

Retrieval changes the system.

Published evidence

PubMedBERT test accuracy is 41, 42 and 47 across the three context settings. [3]

Our interpretation

The six-point no-context-to-PubMed difference belongs to this pipeline comparison; it cannot be assigned to a new base model.

03

Breadth is not an outcome measure.

Published evidence

The benchmark covers medical subjects but scores option selection. [3]

Our interpretation

Subject coverage and clinical action quality are different dimensions; a broad question bank does not observe decisions made in care.

05 / Scope of the evidence

Where this benchmark stops.

Source-version ambiguity

The paper contains inconsistent rounded prose and the README reverses split-table labels. Exact statistics and file provenance must accompany a reproduction. [3][4]

Historical results

The baseline study predates modern frontier models and is included for its experimental comparison, not its rank. [3]

Test-label access

The author repository says test ground truth is withheld and describes a prediction-submission process. A local validation score should be labeled accordingly. [4]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Public author data links; test submission instructions are documented.
License
Author repository LICENSE.md: MIT. Third-party package labels may differ.
Conditions
Verify downloaded-file terms and answer-index encoding; do not infer source-book rights from a repository license.
[4]

Evidence trail

Read the originals.

  1. MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset ↗

    Ankit Pal, Logesh Kumar Umapathi and Malaikannan Sankarasubbu. Original exam-based construction, exact split statistics and retrieval comparison.

  2. MedMCQA author repository ↗

    MedMCQA authors. Original access and submission instructions. README split-table headings disagree with exact paper statistics.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗