MedS-Bench · independent analysis

Match the clinical task to its metric.

An independent MedS-Bench analysis: clinical task coverage, dataset versions, task-specific metrics, published NER results and an interactive evaluation atlas.

Independent analysis by Arcophos · updated

The clinical task atlas

MedS-Bench ↗

28 source datasets · 52 tasks · 11 high-level categories. A selected label, an extracted entity and a generated explanation need different measurements. [1][2][3]

Select

Medical MCQA

Selected option

Accuracy

What this measures

Correct selection does not score the safety of an explanation. [1][2]

Diagnosis

Diagnosis from defined labels

Accuracy

What this measures

Closed disease vocabulary bounds the task. [1][2]

Treatment planning

Treatment category

Accuracy

What this measures

Eight high-level categories do not constitute a complete plan. [1][2]

Clinical outcome prediction

Binary outcome label

Accuracy

What this measures

The task scores an outcome label, not a care action. [1][2]

Text classification

One or more labels

Precision / recall / F1

What this measures

Multilabel errors need the task-specific aggregation. [1][2]

Extract

Information extraction

Requested fields

Parsed accuracy

What this measures

Current script checks normalized reference containment. [1][2]

Named entity recognition

Entity set

F1

What this measures

The set parser and normalization are part of the metric. [1][2]

Generate

Text summarization

Condensed text

BLEU / ROUGE

What this measures

Reference overlap is not a direct factuality judgment. [1][2]

Explanation and rationale

Explanatory text

BLEU / ROUGE

What this measures

Lexical similarity cannot establish every reasoning step. [1][2]

Mixed

Fact verification

Label or generated justification

Accuracy or BLEU / ROUGE

What this measures

Different subtask outputs use different metrics. [1][2]

Natural language inference

Entailment label or explanation

Accuracy or BLEU / ROUGE

What this measures

Label and generative versions are separate measurements. [1][2]

Output groupings are Arcophos editorial annotations of the published task families. Their column sizes do not represent dataset or example counts. [1][2]

Published results

Source record ↗

Selected paper-reported measurements. Dates, model configurations and scoring conditions belong to each panel.

Paper-reported results / selected rows

Selected results within one metric family

Final 2025 journal paper, named-entity recognition tasks; few-shot evaluation.

Mean task F1 × 100 · points
050100
Reported
GPT-4Average NER F1 reported in the paper
59.52
InternLM 2Average NER F1 reported in the paper
45.69
Llama 3Average NER F1 reported in the paper
23.62
MMedIns-Llama 3Instruction-tuned author model; average NER F1
79.29

Paper-reported historical results. These averages concern the NER task set only, not the full benchmark. No numeric uncertainty intervals are supplied.

Source: Table 6 [1]

An original analytical tool

MedS-Bench task and metric atlas

Evidence explorer

Filter the published task families by the kind of output a system must produce. Categories here are Arcophos editorial groupings, not measured performance or an official ranking.

11 of 11 evidence entries shown

Select

Medical MCQA

Read dossier ↗
Input
Exam question and options
Output
Selected option
Measure
Accuracy
Interpretation boundary

Correct selection does not score the safety of an explanation.

[1][2]
Generate

Text summarization

Read dossier ↗
Input
Medical source text
Output
Condensed text
Measure
BLEU / ROUGE
Interpretation boundary

Reference overlap is not a direct factuality judgment.

[1][2]
Extract

Information extraction

Read dossier ↗
Input
Text plus extraction instruction
Output
Requested fields
Measure
Parsed accuracy
Interpretation boundary

Current script checks normalized reference containment.

[1][2]
Generate

Explanation and rationale

Read dossier ↗
Input
Concept or answer to explain
Output
Explanatory text
Measure
BLEU / ROUGE
Interpretation boundary

Lexical similarity cannot establish every reasoning step.

[1][2]
Extract

Named entity recognition

Read dossier ↗
Input
Biomedical text
Output
Entity set
Measure
F1
Interpretation boundary

The set parser and normalization are part of the metric.

[1][2]
Select

Diagnosis

Read dossier ↗
Input
DDXPlus case information
Output
Diagnosis from defined labels
Measure
Accuracy
Interpretation boundary

Closed disease vocabulary bounds the task.

[1][2]
Select

Treatment planning

Read dossier ↗
Input
SEER case information
Output
Treatment category
Measure
Accuracy
Interpretation boundary

Eight high-level categories do not constitute a complete plan.

[1][2]
Select

Clinical outcome prediction

Read dossier ↗
Input
MIMIC4ED case information
Output
Binary outcome label
Measure
Accuracy
Interpretation boundary

The task scores an outcome label, not a care action.

[1][2]
Select

Text classification

Read dossier ↗
Input
Text plus candidate labels
Output
One or more labels
Measure
Precision / recall / F1
Interpretation boundary

Multilabel errors need the task-specific aggregation.

[1][2]
Mixed

Fact verification

Read dossier ↗
Input
Claim or answer with task context
Output
Label or generated justification
Measure
Accuracy or BLEU / ROUGE
Interpretation boundary

Different subtask outputs use different metrics.

[1][2]
Mixed

Natural language inference

Read dossier ↗
Input
Premise and hypothesis
Output
Entailment label or explanation
Measure
Accuracy or BLEU / ROUGE
Interpretation boundary

Label and generative versions are separate measurements.

[1][2]

A task-family label does not certify a clinical workflow. Counts and scores remain specific to the underlying dataset, selected split and scorer. [1][2]

The benchmark in detail

All dossiers →
Dossier01

Final npj Digital Medicine article, 2025

MedS-Bench ↗

Clinical tasks need different output contracts and different metrics.

UnitTask-specific evaluation exampleMeasureTask-specific accuracy, F1, BLEU and ROUGE

What we examine

MedS-Bench spans answer selection, extraction and generated medical text. Each output needs a different measurement. This publication maps the final journal benchmark to its task definitions, scorers and interpretation limits. Explore the task atlas, inspect the historical NER results, and follow the guides to distinguish a closed-set label from a complete clinical workflow. Arcophos contributes independent analysis; the MedS-Bench authors created the benchmark, data and evaluation methods.

Task
Identify the input, required output and permitted answer space.
Metric
Read accuracy, entity F1 and text overlap as different measurements.
Version
Keep final-paper composition, selected splits and current code distinct.

Analysis & interpretation

All analyses →

Questions, answered

What is MedS-Bench?

A benchmark collection for medical language tasks including answer selection, extraction and generated text. This site uses the final journal version and analyzes each task through its input, output and metric.

Is MedS-Ins the same dataset?

No. MedS-Ins is the associated instruction-training collection. Its training-instance totals are not MedS-Bench evaluation size.

Why does the site avoid a single overall model score?

Accuracy, entity F1 and text-overlap measures count different things. We keep the reported measurements within their task and metric scope rather than create an unsupported universal ranking.

Is this the official MedS-Bench site?

No. This is independent Arcophos analysis. The benchmark authors’ paper, repository and data cards are linked as the primary sources.

Prepare a comparison worksheet

Working tool / saved on this device

Specify a MedS-Bench task profile

Interactive worksheet

Record the task and measurement choices your application needs. This planning worksheet is not an official evaluator or a clinical readiness score.

Define the output

Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.

Download the evidence ↗