Jean-Benoit Delbrouck / Research blog

Research blog · Report-generation evaluation

SPEC-CXR, read for report-generation evaluation

SPEC-CXR scores a generated chest X-ray report finding by finding: an LLM maps it and its reference onto 122 fixed entities and checks presence, location, severity and comparison with priors. What that adds to report-level metrics, how the scorer was validated, and what it says about three leading generators.

Paper Jung Oh Lee, Junwoo Cho, … Donggeun Yoo, Taesoo Kim (Lunit). SPEC-CXR: Advancing Clinical Safety through Entity-Level Performance Evaluation of Chest X-ray Report Generation. MICCAI 2025, LNCS 15966, pp. 594–604 By Jean-Benoit Delbrouck Published 5 Oct 2026 Reading time ~8 min Tags MICCAI 2025 Disclosure I am a co-author of GREEN and RadGraph-F1, two of the metrics this paper positions itself against.
122
entities scored one by one: 89 findings and 33 diagnoses
0.60–0.78
agreement with radiologists’ error counts (tau-b), depending on the LLM judge
50
report pairs behind the entity-level validation
≈0.2
presence F1 of three leading generators, averaged over finding groups
TL;DR
  • A confusion matrix per finding. An LLM maps both reports onto 122 fixed entities, labels each one it finds as matched, missed or invented, and judges the location, severity and comparison of the matches. That gives an F1 per finding, which GREEN, FineRadScore and RadGraph-F1 don’t.
  • At report level it ties the best error counts. Agreement with radiologists is 0.777 (tau-b) with the best of eight LLM judges and 0.597 with the worst. The paper compares that with GREEN’s score (0.64); GREEN’s error count is at 0.79 in the GREEN paper, under an inter-expert bound of 0.81.
  • The new part is the least validated. The entity-level scores rest on a manual check of 50 report pairs, and the released code runs a two-step variant of the validated pipeline.
  • Three leading generators average about 0.2. Presence F1 is 0.19 to 0.22 for CXRReportGen, MAIRA-2 and Med-Gemini, and below 0.2 in 9 of 15 finding groups for the best. These averages are over groups, and not the same groups for each model.

01Why score findings one by one

The CT posts in this series kept hitting the same wall: BLEU, RadGraph-F1 and the like give one number per report, which can’t say which findings a generator misses. For Merlin, the finding-level score had to come from a later paper: a recall of 0.10 over 30 findings. Chest X-ray has long had a per-finding score, the CheXbert F1 over 14 observations. Its newer metrics, GREEN and FineRadScore, are LLM judges that list the errors of each report and still return one score per report. SPEC-CXR, from Lunit, joins the two ideas: an LLM judge with a fixed vocabulary of 122 entities, and three attributes on top of presence.

02How a report pair is scored

Three radiologists drew up the vocabulary, meant to be mutually exclusive: 89 findings in 11 categories, from 23 lung and pleura entries to 2 for the abdomen, and 33 differential diagnoses. Rib fracture and tracheostomy tube have their own entries, and every anatomical region has an “other” entry for the rest.

Figure 1 — One constrained LLM output, two read-outs. The paper’s worked example: three records, four errors. Sources: Section 2; Figure 2.

The scoring is then arithmetic. Our reconstruction from Section 2.2 and Figure 2, not the authors’ code:

# a record: entity, presence, location, severity, comparison

def report_errors(records):       # report level: lower is better
    return sum((r.presence in ("FP", "FN"))
               + (r.location == "not-aligned")
               + (r.severity == "not-aligned")
               + (r.comparison != "other") for r in records)

def entity_f1(pairs, entity):     # entity level, over a test set
    n = Counter(r.presence for records in pairs
                for r in records if r.entity == entity)
    return 2 * n["TP"] / (2 * n["TP"] + n["FP"] + n["FN"])

Attributes are scored on matched findings only, so a miss isn’t penalised twice. Two consequences of the design our observation: a finding that neither report mentions leaves no record, so there is an F1 but no specificity; and matching ignores attributes, so in the example a left pneumothorax matches a right one and costs one location error, not a miss plus an invention.

03Is the scorer right?

The report-level score was tested on ReXVal: 200 MIMIC-CXR report pairs on which six radiologists counted errors in six categories, the same six SPEC-CXR counts. The measure is Kendall’s tau-b between a metric and the radiologists’ mean total error count. The candidates aren’t model outputs: each is another patient’s report, retrieved because it scores well on BLEU, BERTScore, CheXbert or RadGraph, and only impressions were compared.

SPEC-CXR error count, by LLM judgeother metrics in the paper’s Table 1error counts in the GREEN paper, not in that table95% CI, where the source gives one0.40.50.60.70.80.9inter-expert, 0.81GREEN error countGREEN model0.79GREEN error countGPT-40.79SPEC-CXRClaude 3.5 Sonnet0.777G-Rad error countGPT-40.76SPEC-CXRo10.758FineRadScoreClaude 3 Opus0.738SPEC-CXRo3-mini0.734SPEC-CXRLlama 3.1 8B*0.721SPEC-CXRDeepSeek-R1-Distill 32B*0.715FineRadScoreo3-mini0.692SPEC-CXRGPT-4o0.656GREEN scoreGPT-40.640SPEC-CXRClaude 3.5 Haiku0.619RadCliQno LLM0.615SPEC-CXRGPT-4o-mini0.597RaTEScoreno LLM0.527
Figure 2 — With its best judge, SPEC-CXR joins the best existing error counts; with its worst, it does no better than a 2023 metric that uses no LLM. Kendall tau-b; the axis starts at 0.4. Lines are the 95% intervals printed in the source papers; SPEC-CXR gives only the spread of five runs (±0.004 to ±0.035). RadCliQ’s value is on 40 held-out pairs. * fine-tuned by the authors. Sources: Table 1; GREEN paper, Table 6; FineRadScore paper, Tables 10 and 11.
Show as table
MetricLLM judgeTau-bInterval or spread, and source
GREEN error countGREEN model0.7995% CI 0.74–0.83 · GREEN paper, Table 6; not in the paper’s Table 1
GREEN error countGPT-40.7995% CI 0.75–0.83 · GREEN paper, Table 6; not in the paper’s Table 1
SPEC-CXRClaude 3.5 Sonnet0.777± 0.004 (SD of five runs) · the paper, Table 1
G-Rad error countGPT-40.7695% CI 0.70–0.80 · GREEN paper, Table 6; not in the paper’s Table 1
SPEC-CXRo10.758± 0.006 (SD of five runs) · the paper, Table 1
FineRadScoreClaude 3 Opus0.73895% CI 0.680–0.788 · FineRadScore paper, Table 11; quoted in the paper’s Table 1
SPEC-CXRo3-mini0.734± 0.004 (SD of five runs) · the paper, Table 1
SPEC-CXRLlama 3.1 8B*0.721± 0.008 (SD of five runs) · the paper, Table 1
SPEC-CXRDeepSeek-R1-Distill 32B*0.715± 0.009 (SD of five runs) · the paper, Table 1
FineRadScoreo3-mini0.692± 0.028 (SD of five runs) · run by the SPEC-CXR authors, Table 1
SPEC-CXRGPT-4o0.656± 0.014 (SD of five runs) · the paper, Table 1
GREEN scoreGPT-40.64095% CI 0.57–0.70 · GREEN paper, Table 6; quoted in the paper’s Table 1
SPEC-CXRClaude 3.5 Haiku0.619± 0.014 (SD of five runs) · the paper, Table 1
RadCliQno LLM0.61595% CI 0.450–0.749 on 40 held-out pairs · Yu et al. 2023; quoted in the paper’s Table 1
SPEC-CXRGPT-4o-mini0.597± 0.035 (SD of five runs) · the paper, Table 1
RaTEScoreno LLM0.527RaTEScore paper; quoted in the paper’s Table 1
Inter-expert–0.8195% CI 0.78–0.83 · GREEN paper, Table 6

Three things to read off it:

  1. The judge matters more than the vocabulary. The same metric moves from 0.597 with GPT-4o-mini to 0.777 with Claude 3.5 Sonnet. With one judge, o3-mini, going from CheXpert’s 14 labels to the 122 entities adds 0.06 (the paper’s Table 2).
  2. One comparison holds the judge fixed. With o3-mini, SPEC-CXR reaches 0.734 and FineRadScore 0.692.
  3. The ± isn’t a confidence interval. It is the spread of five runs of one judge. The 95% intervals other papers print on these 200 pairs span about 0.1 (0.680 to 0.788 for FineRadScore’s 0.738), enough to hold SPEC-CXR’s best value.
The rows the table leaves out external

The paper’s table lists GREEN at 0.640. That is the GREEN score, which keeps only clinically significant errors and normalises by matched findings. The GREEN paper’s table also gives the model’s plain error count, 0.79, and an inter-expert row, 0.81, which it treats as an upper bound. Against those rows, 0.777 is a tie, not a lead.

It also explains the gap the paper credits to its scoring: ReXVal’s target counts every error once, so an unweighted count fits it best our observation. One reviewer objected to exactly that, a missed pneumothorax costing the same as mild atelectasis; the authors replied that useful weights need the exam’s clinical context.

The entity-level scores, the new part, got a much smaller check: a manual assessment of 50 randomly sampled pairs. The scorer’s labels agreed with it 91.8% of the time on presence, 87.2% on location, 84.5% on severity and 75.3% on comparison. Fifty pairs can’t test 122 entities one by one, and the paper gives no per-entity agreement.

04What it shows about three generators

The paper then scores three generators on the 2,347 studies of ReXrank’s MIMIC-CXR test split: Microsoft’s MAIRA-2 and CXRReportGen, and Google’s Med-Gemini. The results are two radar charts over 15 groups of entities, with no table, so we read the values off the image our measurement.

Med-GeminiMAIRA-2CXRReportGenentities · extractedMed-Gemini00.10.20.30.40.50.60.70.2, the paper’s barForeign body – medical11 · ≈1,8000.63Neck/chest wall soft tissue2 · ≈600.41Heart3 · ≈9100.35Lung and pleura – opacity8 · ≈3,0000.34Diaphragm5 · ≈1500.24Lung and pleura – other10 · ≈8400.22Abdomen2 · ≈300.20Mediastinum5 · ≈1300.19Foreign body – non-medical2 · ≈50.17Quality of exams5 · ≈900.14Lung and pleura – lucency5 · ≈1400.13Differential diagnosis33 · ≈1,2000.09Airway4 · ≈300.08Aorta and other vessel5 · ≈3100.08Bone and joint22 · ≈4100.07
Figure 3 — Tubes, lines and devices are found; most other groups stay below 0.2. Presence F1 per group, sorted by Med-Gemini (its value is in the right-hand column); a missing marker means the paper’s line sits at zero there and its average leaves the group out. Beside each name: entities in the group (our mapping of the released list) and how often the scorer extracted them (the chart’s log scale). Our reading of the radar chart, about ±0.01. Source: Figure 4.
Show as table
Finding groupEntities≈ ExtractedCXRReportGenMAIRA-2Med-Gemini
Foreign body – medical111,8000.530.530.63
Neck/chest wall soft tissue2600.320.280.41
Heart39100.260.310.35
Lung and pleura – opacity83,0000.350.380.34
Diaphragm51500.130.190.24
Lung and pleura – other108400.180.170.22
Abdomen2300.190.080.20
Mediastinum51300.110.130.19
Foreign body – non-medical25––0.17
Quality of exams5900.090.070.14
Lung and pleura – lucency51400.170.210.13
Differential diagnosis331,2000.070.100.09
Airway4300.01–0.08
Aorta and other vessel53100.080.050.08
Bone and joint224100.160.160.07
Mean of the values shown0.1890.2060.222
Average in the paper’s legend0.1890.2050.223

Three things stand out:

  1. Devices are found, most other things aren’t. Tubes, lines and implanted devices score 0.53, 0.53 and 0.63; no other group passes 0.5. Med-Gemini, the best model, is below 0.2 in 9 of the 15 groups, as the paper states; the ninth, the abdomen, sits at 0.196.
  2. The weak groups aren’t only the rare ones. Differential diagnoses were extracted about 1,200 times, third among the groups, and all three models stay near 0.1 or below. The paper doesn’t say why. One candidate: diagnoses are usually written in the impression, and MAIRA-2 generates only findings our observation.
  3. Matched findings are described well, with two cautions. The legends give 0.86 to 0.92 for location, 0.81 to 0.91 for severity and 0.94 to 0.95 for comparison. But on lung opacities, the most frequent group, Med-Gemini’s location and severity accuracies are 0.77 and 0.78 our measurement. And on comparison the scorer disagrees with the manual check 24.7% of the time, while it reports a 5 to 6% error rate for the generators derived.
How the averages were taken our measurement

The caption says the legend averages are means over all 122 entities. All twelve equal, within 0.001, the plain mean of the plotted group values, leaving out the groups where the model’s line drops to zero or breaks off. That keeps 15 for Med-Gemini’s presence F1, 13 for MAIRA-2, 14 for CXRReportGen. So a group of 2 entities weighs as much as one of 33, and the three averages don’t cover the same groups. On the 13 groups all three share, they are 0.24, 0.21 and 0.20 derived: MAIRA-2 and CXRReportGen are level.

05What the paper and the release leave out

  • The validated pipeline. The released code (January 2026) extracts entities from each report separately, derives presence by comparing the two lists, and asks the LLM only for the attributes. The repository’s first README names the single-prompt version as the paper’s method, so the agreement figures above belong to a pipeline the release no longer contains. The default judge is GPT-4o (0.656 in the paper’s Table 1); the fine-tuned open judges, at 0.72, aren’t included; the licence is non-commercial.
  • Entity-level detail. No per-entity support and no confidence intervals; nor who annotated the 50 pairs, which judge produced Figure 4, which report sections were compared, or how the generators were run.
  • Coverage. By the authors’ LLM-based test, the 122 entities fully express 81.9% of test reports, against 56.8% for CheXpert’s labels. About one report in five holds a statement the set can’t represent derived.
  • Scope. One reference dataset and one 200-pair benchmark, both from MIMIC-CXR. Reviewers asked for wider validation and a bias analysis of the LLM; the authors agreed to state the first as a limitation.

06Takeaways for evaluating a report generator

  1. Put a per-finding table next to every report-level score. No report-level number shows that the best of three leading generators is below 0.2 in 9 of 15 finding groups.
  2. Name the LLM judge and keep it fixed. One metric, eight judges: 0.597 to 0.777. The release’s default, GPT-4o, is at 0.656 in the paper.
  3. Extract the reference once. The released code re-extracts it on every call. Cache it, or two models can be scored against different reference entities.
  4. Publish the averaging rule and the support. CXRReportGen’s 0.189 becomes 0.203 on the 13 groups behind MAIRA-2’s 0.205.
  5. Don’t rank LLM error counts on ReXVal alone. The best five sit between 0.76 and 0.79, under an inter-expert bound of 0.81, and 200 pairs don’t separate them.

ASources

  1. Lee, J. O., Cho, J., et al. SPEC-CXR. MICCAI 2025, LNCS 15966, pp. 594–604: paper page with the reviews and the rebuttal · PDF · DOI.
  2. SPEC-CXR code, read at commit 0cb653f (14 Jan 2026); its first README, of 15 Jul 2025, is in the repository’s history.
  3. Ostmeier, S., et al. GREEN: arXiv:2405.03595 (2024), Table 6. Huang, A., et al. FineRadScore: arXiv:2405.20613 (2024), Tables 10 and 11.
  4. Yu, F., et al. RadCliQ: Patterns 4(9), 2023; ReXVal on PhysioNet. Zhao, W., et al. RaTEScore: arXiv:2406.16845 (2024).
  5. Zhang, X., et al. ReXrank: arXiv:2411.15122 (2024). Bannur, S., et al. MAIRA-2: arXiv:2406.04449 (2024) and its model card.
  6. Smit, A., et al. CheXbert: arXiv:2004.09167 (2020). Delbrouck, J.-B., et al. RadGraph-F1: arXiv:2210.12186 (2022).