- A confusion matrix per finding. An LLM maps both reports onto 122 fixed entities, labels each one it finds as matched, missed or invented, and judges the location, severity and comparison of the matches. That gives an F1 per finding, which GREEN, FineRadScore and RadGraph-F1 don’t.
- At report level it ties the best error counts. Agreement with radiologists is 0.777 (tau-b) with the best of eight LLM judges and 0.597 with the worst. The paper compares that with GREEN’s score (0.64); GREEN’s error count is at 0.79 in the GREEN paper, under an inter-expert bound of 0.81.
- The new part is the least validated. The entity-level scores rest on a manual check of 50 report pairs, and the released code runs a two-step variant of the validated pipeline.
- Three leading generators average about 0.2. Presence F1 is 0.19 to 0.22 for CXRReportGen, MAIRA-2 and Med-Gemini, and below 0.2 in 9 of 15 finding groups for the best. These averages are over groups, and not the same groups for each model.
01Why score findings one by one
The CT posts in this series kept hitting the same wall: BLEU, RadGraph-F1 and the like give one number per report, which can’t say which findings a generator misses. For Merlin, the finding-level score had to come from a later paper: a recall of 0.10 over 30 findings. Chest X-ray has long had a per-finding score, the CheXbert F1 over 14 observations. Its newer metrics, GREEN and FineRadScore, are LLM judges that list the errors of each report and still return one score per report. SPEC-CXR, from Lunit, joins the two ideas: an LLM judge with a fixed vocabulary of 122 entities, and three attributes on top of presence.
02How a report pair is scored
Three radiologists drew up the vocabulary, meant to be mutually exclusive: 89 findings in 11 categories, from 23 lung and pleura entries to 2 for the abdomen, and 33 differential diagnoses. Rib fracture and tracheostomy tube have their own entries, and every anatomical region has an “other” entry for the rest.
found in either report
The scoring is then arithmetic. Our reconstruction from Section 2.2 and Figure 2, not the authors’ code:
# a record: entity, presence, location, severity, comparison
def report_errors(records): # report level: lower is better
return sum((r.presence in ("FP", "FN"))
+ (r.location == "not-aligned")
+ (r.severity == "not-aligned")
+ (r.comparison != "other") for r in records)
def entity_f1(pairs, entity): # entity level, over a test set
n = Counter(r.presence for records in pairs
for r in records if r.entity == entity)
return 2 * n["TP"] / (2 * n["TP"] + n["FP"] + n["FN"])
Attributes are scored on matched findings only, so a miss isn’t penalised twice. Two consequences of the design our observation: a finding that neither report mentions leaves no record, so there is an F1 but no specificity; and matching ignores attributes, so in the example a left pneumothorax matches a right one and costs one location error, not a miss plus an invention.
03Is the scorer right?
The report-level score was tested on ReXVal: 200 MIMIC-CXR report pairs on which six radiologists counted errors in six categories, the same six SPEC-CXR counts. The measure is Kendall’s tau-b between a metric and the radiologists’ mean total error count. The candidates aren’t model outputs: each is another patient’s report, retrieved because it scores well on BLEU, BERTScore, CheXbert or RadGraph, and only impressions were compared.
Show as table
| Metric | LLM judge | Tau-b | Interval or spread, and source |
|---|---|---|---|
| GREEN error count | GREEN model | 0.79 | 95% CI 0.74–0.83 · GREEN paper, Table 6; not in the paper’s Table 1 |
| GREEN error count | GPT-4 | 0.79 | 95% CI 0.75–0.83 · GREEN paper, Table 6; not in the paper’s Table 1 |
| SPEC-CXR | Claude 3.5 Sonnet | 0.777 | ± 0.004 (SD of five runs) · the paper, Table 1 |
| G-Rad error count | GPT-4 | 0.76 | 95% CI 0.70–0.80 · GREEN paper, Table 6; not in the paper’s Table 1 |
| SPEC-CXR | o1 | 0.758 | ± 0.006 (SD of five runs) · the paper, Table 1 |
| FineRadScore | Claude 3 Opus | 0.738 | 95% CI 0.680–0.788 · FineRadScore paper, Table 11; quoted in the paper’s Table 1 |
| SPEC-CXR | o3-mini | 0.734 | ± 0.004 (SD of five runs) · the paper, Table 1 |
| SPEC-CXR | Llama 3.1 8B* | 0.721 | ± 0.008 (SD of five runs) · the paper, Table 1 |
| SPEC-CXR | DeepSeek-R1-Distill 32B* | 0.715 | ± 0.009 (SD of five runs) · the paper, Table 1 |
| FineRadScore | o3-mini | 0.692 | ± 0.028 (SD of five runs) · run by the SPEC-CXR authors, Table 1 |
| SPEC-CXR | GPT-4o | 0.656 | ± 0.014 (SD of five runs) · the paper, Table 1 |
| GREEN score | GPT-4 | 0.640 | 95% CI 0.57–0.70 · GREEN paper, Table 6; quoted in the paper’s Table 1 |
| SPEC-CXR | Claude 3.5 Haiku | 0.619 | ± 0.014 (SD of five runs) · the paper, Table 1 |
| RadCliQ | no LLM | 0.615 | 95% CI 0.450–0.749 on 40 held-out pairs · Yu et al. 2023; quoted in the paper’s Table 1 |
| SPEC-CXR | GPT-4o-mini | 0.597 | ± 0.035 (SD of five runs) · the paper, Table 1 |
| RaTEScore | no LLM | 0.527 | RaTEScore paper; quoted in the paper’s Table 1 |
| Inter-expert | – | 0.81 | 95% CI 0.78–0.83 · GREEN paper, Table 6 |
Three things to read off it:
- The judge matters more than the vocabulary. The same metric moves from 0.597 with GPT-4o-mini to 0.777 with Claude 3.5 Sonnet. With one judge, o3-mini, going from CheXpert’s 14 labels to the 122 entities adds 0.06 (the paper’s Table 2).
- One comparison holds the judge fixed. With o3-mini, SPEC-CXR reaches 0.734 and FineRadScore 0.692.
- The ± isn’t a confidence interval. It is the spread of five runs of one judge. The 95% intervals other papers print on these 200 pairs span about 0.1 (0.680 to 0.788 for FineRadScore’s 0.738), enough to hold SPEC-CXR’s best value.
The paper’s table lists GREEN at 0.640. That is the GREEN score, which keeps only clinically significant errors and normalises by matched findings. The GREEN paper’s table also gives the model’s plain error count, 0.79, and an inter-expert row, 0.81, which it treats as an upper bound. Against those rows, 0.777 is a tie, not a lead.
It also explains the gap the paper credits to its scoring: ReXVal’s target counts every error once, so an unweighted count fits it best our observation. One reviewer objected to exactly that, a missed pneumothorax costing the same as mild atelectasis; the authors replied that useful weights need the exam’s clinical context.
The entity-level scores, the new part, got a much smaller check: a manual assessment of 50 randomly sampled pairs. The scorer’s labels agreed with it 91.8% of the time on presence, 87.2% on location, 84.5% on severity and 75.3% on comparison. Fifty pairs can’t test 122 entities one by one, and the paper gives no per-entity agreement.
04What it shows about three generators
The paper then scores three generators on the 2,347 studies of ReXrank’s MIMIC-CXR test split: Microsoft’s MAIRA-2 and CXRReportGen, and Google’s Med-Gemini. The results are two radar charts over 15 groups of entities, with no table, so we read the values off the image our measurement.
Show as table
| Finding group | Entities | ≈ Extracted | CXRReportGen | MAIRA-2 | Med-Gemini |
|---|---|---|---|---|---|
| Foreign body – medical | 11 | 1,800 | 0.53 | 0.53 | 0.63 |
| Neck/chest wall soft tissue | 2 | 60 | 0.32 | 0.28 | 0.41 |
| Heart | 3 | 910 | 0.26 | 0.31 | 0.35 |
| Lung and pleura – opacity | 8 | 3,000 | 0.35 | 0.38 | 0.34 |
| Diaphragm | 5 | 150 | 0.13 | 0.19 | 0.24 |
| Lung and pleura – other | 10 | 840 | 0.18 | 0.17 | 0.22 |
| Abdomen | 2 | 30 | 0.19 | 0.08 | 0.20 |
| Mediastinum | 5 | 130 | 0.11 | 0.13 | 0.19 |
| Foreign body – non-medical | 2 | 5 | – | – | 0.17 |
| Quality of exams | 5 | 90 | 0.09 | 0.07 | 0.14 |
| Lung and pleura – lucency | 5 | 140 | 0.17 | 0.21 | 0.13 |
| Differential diagnosis | 33 | 1,200 | 0.07 | 0.10 | 0.09 |
| Airway | 4 | 30 | 0.01 | – | 0.08 |
| Aorta and other vessel | 5 | 310 | 0.08 | 0.05 | 0.08 |
| Bone and joint | 22 | 410 | 0.16 | 0.16 | 0.07 |
| Mean of the values shown | 0.189 | 0.206 | 0.222 | ||
| Average in the paper’s legend | 0.189 | 0.205 | 0.223 |
Three things stand out:
- Devices are found, most other things aren’t. Tubes, lines and implanted devices score 0.53, 0.53 and 0.63; no other group passes 0.5. Med-Gemini, the best model, is below 0.2 in 9 of the 15 groups, as the paper states; the ninth, the abdomen, sits at 0.196.
- The weak groups aren’t only the rare ones. Differential diagnoses were extracted about 1,200 times, third among the groups, and all three models stay near 0.1 or below. The paper doesn’t say why. One candidate: diagnoses are usually written in the impression, and MAIRA-2 generates only findings our observation.
- Matched findings are described well, with two cautions. The legends give 0.86 to 0.92 for location, 0.81 to 0.91 for severity and 0.94 to 0.95 for comparison. But on lung opacities, the most frequent group, Med-Gemini’s location and severity accuracies are 0.77 and 0.78 our measurement. And on comparison the scorer disagrees with the manual check 24.7% of the time, while it reports a 5 to 6% error rate for the generators derived.
The caption says the legend averages are means over all 122 entities. All twelve equal, within 0.001, the plain mean of the plotted group values, leaving out the groups where the model’s line drops to zero or breaks off. That keeps 15 for Med-Gemini’s presence F1, 13 for MAIRA-2, 14 for CXRReportGen. So a group of 2 entities weighs as much as one of 33, and the three averages don’t cover the same groups. On the 13 groups all three share, they are 0.24, 0.21 and 0.20 derived: MAIRA-2 and CXRReportGen are level.
05What the paper and the release leave out
- The validated pipeline. The released code (January 2026) extracts entities from each report separately, derives presence by comparing the two lists, and asks the LLM only for the attributes. The repository’s first README names the single-prompt version as the paper’s method, so the agreement figures above belong to a pipeline the release no longer contains. The default judge is GPT-4o (0.656 in the paper’s Table 1); the fine-tuned open judges, at 0.72, aren’t included; the licence is non-commercial.
- Entity-level detail. No per-entity support and no confidence intervals; nor who annotated the 50 pairs, which judge produced Figure 4, which report sections were compared, or how the generators were run.
- Coverage. By the authors’ LLM-based test, the 122 entities fully express 81.9% of test reports, against 56.8% for CheXpert’s labels. About one report in five holds a statement the set can’t represent derived.
- Scope. One reference dataset and one 200-pair benchmark, both from MIMIC-CXR. Reviewers asked for wider validation and a bias analysis of the LLM; the authors agreed to state the first as a limitation.
06Takeaways for evaluating a report generator
- Put a per-finding table next to every report-level score. No report-level number shows that the best of three leading generators is below 0.2 in 9 of 15 finding groups.
- Name the LLM judge and keep it fixed. One metric, eight judges: 0.597 to 0.777. The release’s default, GPT-4o, is at 0.656 in the paper.
- Extract the reference once. The released code re-extracts it on every call. Cache it, or two models can be scored against different reference entities.
- Publish the averaging rule and the support. CXRReportGen’s 0.189 becomes 0.203 on the 13 groups behind MAIRA-2’s 0.205.
- Don’t rank LLM error counts on ReXVal alone. The best five sit between 0.76 and 0.79, under an inter-expert bound of 0.81, and 200 pairs don’t separate them.
ASources
- Lee, J. O., Cho, J., et al. SPEC-CXR. MICCAI 2025, LNCS 15966, pp. 594–604: paper page with the reviews and the rebuttal · PDF · DOI.
- SPEC-CXR code, read at commit 0cb653f (14 Jan 2026); its first README, of 15 Jul 2025, is in the repository’s history.
- Ostmeier, S., et al. GREEN: arXiv:2405.03595 (2024), Table 6. Huang, A., et al. FineRadScore: arXiv:2405.20613 (2024), Tables 10 and 11.
- Yu, F., et al. RadCliQ: Patterns 4(9), 2023; ReXVal on PhysioNet. Zhao, W., et al. RaTEScore: arXiv:2406.16845 (2024).
- Zhang, X., et al. ReXrank: arXiv:2411.15122 (2024). Bannur, S., et al. MAIRA-2: arXiv:2406.04449 (2024) and its model card.
- Smit, A., et al. CheXbert: arXiv:2004.09167 (2020). Delbrouck, J.-B., et al. RadGraph-F1: arXiv:2210.12186 (2022).