Medical MCQA
Selected option
Accuracy
MedS-Bench · independent analysis
An independent MedS-Bench analysis: clinical task coverage, dataset versions, task-specific metrics, published NER results and an interactive evaluation atlas.
28 source datasets · 52 tasks · 11 high-level categories. A selected label, an extracted entity and a generated explanation need different measurements. [1][2][3]
Selected option
Accuracy
Diagnosis from defined labels
Accuracy
Treatment category
Accuracy
Binary outcome label
Accuracy
One or more labels
Precision / recall / F1
Requested fields
Parsed accuracy
Entity set
F1
Condensed text
BLEU / ROUGE
Explanatory text
BLEU / ROUGE
Label or generated justification
Accuracy or BLEU / ROUGE
Entailment label or explanation
Accuracy or BLEU / ROUGE
Output groupings are Arcophos editorial annotations of the published task families. Their column sizes do not represent dataset or example counts. [1][2]
Selected paper-reported measurements. Dates, model configurations and scoring conditions belong to each panel.
Paper-reported results / selected rows
Final 2025 journal paper, named-entity recognition tasks; few-shot evaluation.
Paper-reported historical results. These averages concern the NER task set only, not the full benchmark. No numeric uncertainty intervals are supplied.
Source: Table 6 [1]
An original analytical tool
Filter the published task families by the kind of output a system must produce. Categories here are Arcophos editorial groupings, not measured performance or an official ranking.
11 of 11 evidence entries shown
A task-family label does not certify a clinical workflow. Counts and scores remain specific to the underlying dataset, selected split and scorer. [1][2]
Final npj Digital Medicine article, 2025
Clinical tasks need different output contracts and different metrics.
MedS-Bench spans answer selection, extraction and generated medical text. Each output needs a different measurement. This publication maps the final journal benchmark to its task definitions, scorers and interpretation limits. Explore the task atlas, inspect the historical NER results, and follow the guides to distinguish a closed-set label from a complete clinical workflow. Arcophos contributes independent analysis; the MedS-Bench authors created the benchmark, data and evaluation methods.
Map clinical language tasks to accuracy, entity F1 and reference-overlap measures before interpreting model performance.
Inspect the answer space behind familiar clinical task names and avoid extending label accuracy into unsupported workflow claims.
Keep training-data counts, benchmark composition, sampled evaluation cases and scorer revisions separate.
A benchmark collection for medical language tasks including answer selection, extraction and generated text. This site uses the final journal version and analyzes each task through its input, output and metric.
No. MedS-Ins is the associated instruction-training collection. Its training-instance totals are not MedS-Bench evaluation size.
Accuracy, entity F1 and text-overlap measures count different things. We keep the reported measurements within their task and metric scope rather than create an unsupported universal ranking.
No. This is independent Arcophos analysis. The benchmark authors’ paper, repository and data cards are linked as the primary sources.
Working tool / saved on this device
Record the task and measurement choices your application needs. This planning worksheet is not an official evaluator or a clinical readiness score.
Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.
Download the evidence ↗