Jean-Benoit Delbrouck / Research blog

Research blog · Report-generation evaluation

Phrase-grounded fact-checking, read for report-generation evaluation

A MICCAI 2025 model checks a generated chest X-ray report against the image itself: for each finding it answers real or fake and draws a box, with no reference report. Its headline number, a concordance of 0.997 with ground truth, can be reproduced from 28 pairs of averages in one table: what that shows, and what it leaves open.

Paper Razi Mahmood, Diego Machado-Reyes, … Pingkun Yan, Tanveer Syeda-Mahmood. Phrase-grounded Fact-checking for Automatically Generated Chest X-ray Reports. MICCAI 2025, LNCS 15966, pp. 441–452 By Jean-Benoit Delbrouck Published 5 Oct 2026 Reading time ~8 min Tags MICCAI 2025 · Main Disclosure I am a co-author of GREEN and RadGraph-F1. The paper says its score was shown to outperform RadGraph-F1, and a reviewer asked for comparisons with both.
27M
synthetic finding–box pairs made from real reports, about 16 to 17 for each real pair*
0.90–0.94
accuracy at telling real pairs from fake ones, on four test sets
0.49–0.57
mean IoU of its boxes with the reference boxes
0.997
concordance with a reference-based score; 28 pairs of averages give it*

* Our arithmetic on the paper’s Tables 3 and 5.

TL;DR
  • A checker that looks at the image, not at a reference. It takes the image and one finding phrase, such as “yes|edema”, and returns a real-or-fake verdict and a box.
  • It learns from fakes made by rule. Real findings are flipped, moved or swapped, about 16 to 17 fakes for each real pair. In a sample of the released files, seven fakes in ten carry a phrase that is almost never real.
  • 28 pairs of averages reproduce the 0.997. For seven generators on four datasets, the checker’s mean error score stays within 0.022 of a reference-based one. That shows it flags about the right amount of error, not that it flags the right findings.
  • What we can’t tell. How often one verdict on a real generator’s finding is right, and how much of the 0.90 to 0.94 is the mix: by the paper’s own counts, the datasets behind it are 64 to 92% fake.

01Why check a report against the image

The blog’s first post on evaluation read SPEC-CXR, which scores a generated chest X-ray report against a reference report, finding by finding. This paper checks the report against the image instead. That matters at the point of use, where no reference exists.

It comes from Rensselaer Polytechnic Institute, IBM Research, Stanford and Massachusetts General Hospital. Some of the authors built a first checker in 2023: a frozen CLIP encoder with an SVM that labelled whole sentences real or fake. The new one, the FC model, reads normalised finding phrases, trains its encoders and also says where the finding is.

Figure 1 — The model gets the image and a phrase; where the report puts the finding enters only afterwards. The extractor and the region detector come from earlier work; the score is worded as in the first preprint. Sources: Sections 2 to 2.3; Figures 1 and 2.

So the verdict is about presence. The first preprint says the location words are dropped from the model’s input. A misplaced finding shows only as a low overlap between two boxes.

02The training data: fakes made by rule

With no large set of labelled generator errors to learn from, the authors make their own. Each ground-truth report yields real pairs: a finding phrase, and a box around the anatomical regions the finding is attached to. Each real pair then spawns fakes of three kinds. Omissions and wrong severities, the other error types the paper names, are left out.

box of the real pairthe same box, for referencebox of a fake pairOriginalyes|edemareal · E = 1box 0.14, 0.13, 0.72, 0.56Reversalno boxno|edemafake · E = 0box 0, 0, 0, 0Relocationyes|edemafake · E = 0box 0.85, 0.74, 0.1, 0.21Relocationyes|edemafake · E = 0box 0.9, 0.7, 0.1, 0.2Substitutionyes|lung cystfake · E = 0box 0.02, 0.48, 0.1, 0.14
Figure 2 — A fake is the real pair with its polarity flipped, its box moved or its finding swapped. The paper’s own example, drawn to scale; each frame is the image. A relocated box comes from where the finding sits in other images; a swapped finding is one the report doesn’t contain. Source: Table 2.

Look at the two relocated fakes. They keep the image and the phrase of the real pair, yes|edema; only the target changes. Since the model’s input is the image and the phrase, nothing it receives tells the three apart our observation, and the paper doesn’t say what it should learn from them.

DatasetImages and labelsFinding typesReal pairsSynthetic pairsPer real pair
Chest ImaGenome silverMIMIC-CXR; automatic labels491,616,85227,047,05416.7
Chest ImaGenome goldMIMIC-CXR; clinician-verified354,06323,4635.8
MS-CXRMIMIC-CXR; radiologists’ boxes82,24724,33810.8
ChestX-ray8NIH; radiologists’ boxes81,57110,1376.5
VinDr-CXRVietnam; radiologists’ boxes2347,973132,6322.8

Table 1 — Five public datasets, three of them cut from MIMIC-CXR. The silver set supplies the bulk; the paper sets the gold one aside for the generator test. Source: Table 3 (silver counts in full from the first preprint); last column derived.

What the released pairs show our count

The Chest ImaGenome silver pairs are public on Hugging Face as RadCheck: an image id, a phrase such as yes|pleural effusion, a box and a 0/1 label per row, over 22 finding names in our sample (VinDr-CXR’s labels). Fakes outnumber real pairs 17 to 1: 7,465,605 against 429,266 in the validation and test splits.

In a sample of 33,000 validation pairs, seven fakes in ten carry one of eight phrases, among them yes|clavicle fracture and yes|rib fracture, which make up under 3% of the real pairs: the phrase gives the label away before the image is read. The sample also holds 68 image–phrase combinations labelled real with one box and fake with another, as in Figure 2, and 9 labelled both ways with the same empty box.

03Test 1: held-out synthetic pairs

The first test asks the model to do on held-out pairs what it was trained to do: label each pair real or fake, and regress its box.

the FC modelits three ablationsshare of fake pairs in the datasetReal or fake: accuracy2023 sentence classifier0.60.70.80.91.00.9283–85%0.9491–92%0.9285–87%0.9064–73%Where: mean IoU of the boxMAIRA-2Med-RPG0.20.30.40.50.60.540.530.570.49Chest ImaGenome goldMS-CXRChestX-ray8VinDr-CXR
Figure 3 — On three of four sets the verdict sits a few points above the share of fake pairs; the box leads two grounding models by 0.05 or more. The bracket is the share of fake pairs in that dataset by Table 3, read two ways (the synthetic count with or without the real pairs). No intervals are reported. Source: Table 4; brackets derived from Table 3.
Show as table
ImaGenome goldMS-CXRChestX-ray8VinDr-CXR
MethodAcc.mIoUAcc.mIoUAcc.mIoUAcc.mIoU
FC model (FCRegComb)0.920.540.940.530.920.570.900.49
BCE loss (FCRegBCE)0.880.490.920.460.900.530.880.45
Two heads (FCRegDual)0.870.510.890.490.870.510.860.47
Pre-built CLIP (FCRegSep)0.890.380.890.390.920.420.890.37
Med-RPG–0.23–0.32–0.28–0.38
MAIRA-2–0.39–0.48–0.51–0.42
2023 classifier (R/F Model)0.84–0.78–0.81–0.83–
Share of fake pairs (derived)83–85%91–92%85–87%64–73%

Four things to read off it:

  1. An accuracy without its mix. The paper gives no confusion matrix and no balance of the test partitions. By Table 3 the pairs are 83 to 92% fake on three of the four sets derived. If the test partitions keep that mix, answering “fake” every time scores as much: the model is 0.02 to 0.09 above it, the 2023 classifier below it on two or three.
  2. The only spread given is wider than the gaps. Over a ten-fold re-split the paper reports 0.92 ± 0.12, of unstated meaning. The ablations sit within 0.05 of the full model in accuracy.
  3. The box baselines may not be on home ground. MAIRA-2 is 0.05 to 0.15 behind and Med-RPG 0.11 to 0.31, and the paper doesn’t say how either was run. Med-RPG’s own paper reports 0.59 on its MS-CXR split, against 0.32 here, and 0.39 for a version that sees only the phrase.
  4. A mean IoU needs a floor. In our sample of the released pairs, half of the real boxes cover more than 37% of the image, and one fixed box per phrase, set without looking at the image, overlaps them at a mean IoU of 0.50 our count. That is the silver set, not the table’s test sets, but it is the kind of yardstick the 0.49 to 0.57 lacks.

04Test 2: seven real report generators

The second test gives the abstract its only number. Seven generators write reports for the same images, 439 of them on the gold set: RGRG, XrayGPT, R2GenGPT, CvT2DistilGPT2, CheXRepair, MAIRA-2 and an in-house GPT-4 tool. Each report gets the error score of Figure 1 twice: once from the checker’s verdicts and boxes, once with the dataset’s annotations in the checker’s place. The second is a metric from a companion paper by the same group. The two agree with a concordance correlation coefficient of 0.997, which the authors read as support for the checker as a “surrogate for ground truth”.

0.30.30.40.40.50.50.60.60.70.70.80.8equalerror score against the reference annotationserror score from the checkerChest ImaGenome goldVinDr-CXRMS-CXRChestX-ray828 pairs of averages7 generators × 4 datasets0.997their concordance (our arithmetic)0.022largest gap between the two scores82%of the variance is between datasets
Figure 4 — These 28 pairs of averages give the 0.997, and the datasets, more than the generators, spread them. Each point is one generator on one dataset: its mean error score from the checker against the same score from the reference annotations. Dashed line: equal scores. Source: Table 5; the four statistics are derived.
Show as table
ImaGenome goldMS-CXRChestX-ray8VinDr-CXR
GeneratorCheckerRef.CheckerRef.CheckerRef.CheckerRef.
RGRG0.5410.5370.3290.3080.3050.2980.5490.537
XrayGPT0.6220.6260.3880.3910.3770.3550.6180.609
GPT-4, in-house0.6580.6530.4330.4260.3990.4080.6360.630
R2GenGPT0.5870.5850.3770.3740.3460.3330.5810.579
CvT2DistilGPT20.5760.5730.4390.4330.4270.4200.5880.600
CheXRepair0.7440.7330.4660.4610.4390.4320.7090.714
MAIRA-20.6190.6330.4230.4250.4120.4190.5780.569
Concordance (derived)0.9920.9800.9690.986

Lin’s coefficient on the 28 printed pairs is 0.997 derived, the paper’s number to the third decimal. The paper doesn’t say what its coefficient was computed on and gives no per-report value. The rebuttal counts its evidence the same way: six metrics, four datasets and six generators make “144 measurements”.

At that level the agreement holds: the two means never differ by more than 0.022, and within each dataset the seven generators come out in the same order, bar one swap between two that the reference separates by 0.007. Two things temper it. First, 82% of the variance in the reference score lies between datasets derived: both scores are near 0.6 on the gold set and VinDr-CXR and near 0.4 on the other two. Second, the two sides share a front end our observation: the same extractor and region boxes turn each report into findings and locations before either comparison, so their errors move both scores together.

What an agreement of averages shows our observation

Each score averages two shares: the findings judged real, and the box overlap. Over a dataset, two such scores agree when the checker calls about as many findings fake as the reference does and its boxes overlap the stated locations about as much. It can do both while flagging the wrong findings.

Flagging errors in one report needs another number: of the findings the reference contradicts, how many did the checker flag, and how many correct ones did it flag by mistake? On generator output the paper shows five example sentences and no such table.

05What the paper leaves out

The three reviewers started at 2, 3 and 6 out of 6 and all ended at accept.

  • Verdicts on real errors: no per-finding accuracy on generator output, no radiologist, no weighting of errors. Asked whether synthetic errors transfer, the authors pointed to Table 5. And by the first preprint the reference side, too, keeps only the findings a report mentions.
  • The promised comparison: a reviewer asked how the verdicts line up with GREEN, RadGraph-F1, CheXbert-F1 and RadCliQ. The rebuttal quotes concordances of 0.998, 0.999 and 0.998 for the first three and promises the full results in the arXiv version, which has the camera-ready text without them.
  • Uncertainty and bookkeeping: no interval beyond the ±0.12, and no word on which pairs the mean IoU averages over. The gold set has 439 images in the text and 461 in Table 3.
  • Code and model: the paper page lists neither. Without the extractor and the region detector, a report can’t be turned into the phrases and boxes the check needs.

06Takeaways for evaluating without a reference

  1. Validate a checker per finding, on real generator output. Agreement between averages, which gives the 0.997 here, can’t tell a checker that flags the right findings from one that flags the right number of them.
  2. Print the floor next to every score. An accuracy of 0.92 reads differently once 83 to 85% of the dataset’s pairs are fake, and a mean IoU of 0.54 once a fixed box reaches 0.50 on the released silver pairs.
  3. Make rule-made errors as hard as a generator’s. When seven fakes in ten carry a phrase that is almost never real, as in the released sample, accuracy says little about the hard case: a common finding that is or isn’t there.
  4. Use an image-based check beside a reference-based score, not in its place. It needs no second report, but it can’t see what a report leaves out; a score like SPEC-CXR can.

ASources

  1. Mahmood, R., Machado-Reyes, D., Wu, J., Kaviani, P., Wong, K. C. L., D’Souza, N., Kalra, M., Wang, G., Yan, P., Syeda-Mahmood, T. Phrase-grounded Fact-checking for Automatically Generated Chest X-ray Reports. MICCAI 2025, LNCS 15966, pp. 441–452: paper page with the reviews and the rebuttal · PDF · DOI · arXiv:2509.21356 (the camera-ready text).
  2. The first preprint: Mahmood, R., et al. Anatomically-Grounded Fact Checking of Automated Chest X-ray Reports. arXiv:2412.02177 (Dec 2024). The same model and results, with a report-correction step the MICCAI version drops.
  3. The companion metric: Mahmood, R., et al. Evaluating Automated Radiology Report Quality through Fine-Grained Phrasal Grounding of Clinical Findings. arXiv:2412.01031. The 2023 classifier: Mahmood, R., et al. Fact-Checking of AI-Generated Reports. arXiv:2307.14634.
  4. The released pairs: razi-mahmood/RadCheck on Hugging Face, read through its dataset viewer on 5 Oct 2026: exact counts for the validation and test splits, and 330 evenly spaced pages of 100 validation rows as the sample. The fixed box of section 03 is the median box per phrase, fitted on half of the sampled images and scored on the other half.
  5. Chen, Z., et al. Med-RPG: arXiv:2303.07618 (2023), Tables I and II. Bannur, S., et al. MAIRA-2: arXiv:2406.04449 (2024).
  6. Datasets: Chest ImaGenome, MS-CXR and VinDr-CXR on PhysioNet; ChestX-ray8.
  7. Lin, L. I.-K. A concordance correlation coefficient to evaluate reproducibility. Biometrics 45, 255–268 (1989).
  8. Our post on SPEC-CXR.