* Our arithmetic: 15% of each modality’s reports, the paper’s split; plain means of Table 3’s three relation scores.
- A test of the extractor behind RadGraph-F1, on reports it wasn’t trained on. Released RadGraph and RadGraph-XL are scored against two annotators’ consensus labels, next to a RadGraph-style model fine-tuned on each modality.
- Fine-tuning lifts most scores, most on cardiac MRI and head CT. Averaged over Table 3, the fine-tuned models reach 0.80 entity F1 and 0.63 relation F1, against 0.72 and 0.55 for the better released model; on CTPA and ultrasound entities the lead is 0.02 to 0.04.
- Not a like-for-like comparison. Only the fine-tuned model has seen the modality and these annotators’ labels, and by our estimate several scores rest on about ten test items.
- What it doesn’t test. No generated report is scored, so the effect on RadGraph-F1 is open, and the promised annotations and weights are not out.
01Why this paper matters for report-generation evaluation
RadGraph-F1 doesn’t compare the reports word for word. It runs an extractor over the generated report and the reference and compares what comes out: entities labelled as anatomy or as an observation that is present, absent or uncertain, and the relations between them.
RG_ER = the same, with a has-relation flag on each entity
# rewards.py, radgraph package; its README calls RG_ER the widely reported one
So the score is only as good as the extractor on that kind of report. RadGraph’s was trained on chest X-ray reports, RadGraph-XL’s on chest X-ray, chest CT, abdomen/pelvis CT and brain MR. CT and MRI report generators already publish RadGraph-F1, and two papers read in this series didn’t say which RadGraph-F1 they ran (Merlin, Cervical-RG). This paper asks the prior question for four other kinds of report. It sits with the series’ posts on SPEC-CXR and phrase-grounded fact-checking.
02The data: 800 reports, four modalities
The four exams are cardiac MRI, CT pulmonary angiography (CTPA), head CT and abdominal ultrasound. For each, two radiologists or trainees labelled every report independently and then agreed on one version, after training on RadGraph’s guidelines. The schema is RadGraph’s: four entity labels and three relations (modify, located at, suggestive of).
| Cardiac MRI | CTPA | Head CT | Abdominal US | |
|---|---|---|---|---|
| Reports | 100 | 200 | 300 | 200 |
| Train / validation / test reports* | 70 / 15 / 15 | 140 / 30 / 30 | 210 / 45 / 45 | 140 / 30 / 30 |
| Entities | 2,326 | 5,893 | 4,362 | 2,232 |
| Relations | 1,776 | 4,520 | 3,345 | 1,538 |
| Entities per report* | 23 | 29 | 15 | 11 |
| Uncertain observations (≈ in test*) | 57 (≈9) | 247 (≈37) | 74 (≈11) | 152 (≈23) |
| Suggestive-of relations (≈ in test*) | 90 (≈14) | 302 (≈45) | 112 (≈17) | 163 (≈24) |
Table 1 — The new data. Counts from the paper’s Table 1. * Our arithmetic: the 70/15/15 split per modality; entities per report; ≈ in test is 15% of the total, assuming the labels spread evenly over reports. The paper prints 32.9% for ultrasound’s absent observations; 430 of 2,232 is 19.3%. Sources: §2.2, §3.1.
Two things in it matter for the scores. The rare labels are rare: uncertain observations are 1.7% to 6.8% of entities and suggestive-of 3.3% to 10.6% of relations, about 9 and 14 items in the 15 cardiac MRI test reports derived. And the paper’s 60% increase over RadGraph holds for reports, 800 against the 500 of RadGraph’s development set, not for labels: 14,813 entities against 14,579, because these reports carry 11 to 29 entities each against 29 in RadGraph’s derived.
03Results: 28 scores on one scale
Table 3 gives an F1 for each label on each modality’s test reports.
Show as table
| Modality · label | RadGraph | RadGraph-XL | Fine-tuned | Lead* | ≈ in test* |
|---|---|---|---|---|---|
| Cardiac MRI · Anatomy | 0.827 | 0.826 | 0.955 | +0.128 | 147 |
| Cardiac MRI · Obs. present | 0.701 | 0.732 | 0.871 | +0.139 | 176 |
| Cardiac MRI · Obs. absent | 0.659 | 0.720 | 0.757 | +0.037 | 17 |
| Cardiac MRI · Obs. uncertain | 0.207 | 0.343 | 0.424 | +0.081 | 9 |
| Cardiac MRI · Modify | 0.503 | 0.643 | 0.847 | +0.204 | 185 |
| Cardiac MRI · Located at | 0.472 | 0.550 | 0.775 | +0.225 | 68 |
| Cardiac MRI · Suggestive of | 0.261 | 0.440 | 0.444 | +0.004 | 14 |
| CTPA · Anatomy | 0.900 | 0.923 | 0.925 | +0.002 | 401 |
| CTPA · Obs. present | 0.785 | 0.753 | 0.822 | +0.037 | 382 |
| CTPA · Obs. absent | 0.795 | 0.862 | 0.860 | -0.002 | 64 |
| CTPA · Obs. uncertain | 0.595 | 0.549 | 0.656 | +0.061 | 37 |
| CTPA · Modify | 0.580 | 0.626 | 0.664 | +0.038 | 488 |
| CTPA · Located at | 0.464 | 0.466 | 0.468 | +0.002 | 144 |
| CTPA · Suggestive of | 0.590 | 0.526 | 0.608 | +0.018 | 45 |
| Head CT · Anatomy | 0.886 | 0.900 | 0.912 | +0.012 | 265 |
| Head CT · Obs. present | 0.671 | 0.653 | 0.833 | +0.162 | 232 |
| Head CT · Obs. absent | 0.772 | 0.793 | 0.896 | +0.103 | 146 |
| Head CT · Obs. uncertain | 0.444 | 0.400 | 0.667 | +0.223 | 11 |
| Head CT · Modify | 0.555 | 0.675 | 0.797 | +0.122 | 361 |
| Head CT · Located at | 0.415 | 0.488 | 0.602 | +0.114 | 124 |
| Head CT · Suggestive of | 0.370 | 0.514 | 0.400 | -0.114 | 17 |
| Ultrasound · Anatomy | 0.872 | 0.889 | 0.876 | -0.013 | 110 |
| Ultrasound · Obs. present | 0.810 | 0.815 | 0.843 | +0.028 | 137 |
| Ultrasound · Obs. absent | 0.876 | 0.891 | 0.897 | +0.006 | 64 |
| Ultrasound · Obs. uncertain | 0.595 | 0.546 | 0.615 | +0.020 | 23 |
| Ultrasound · Modify | 0.616 | 0.685 | 0.730 | +0.045 | 148 |
| Ultrasound · Located at | 0.555 | 0.629 | 0.655 | +0.026 | 58 |
| Ultrasound · Suggestive of | 0.426 | 0.377 | 0.576 | +0.150 | 24 |
* Lead: the fine-tuned model minus the better released model. ≈ in test: 15% of the label’s total in the paper’s Table 1, our estimate.
Three things stand out:
- Relations are the weak half. Per modality, the released models average 0.41 to 0.56 relation F1 and 0.60 to 0.79 entity F1 derived, and RG_ER keys each entity on whether it has a relation.
- Cardiac MRI is where they fall furthest. Over all seven labels, both released models score lowest there derived. Their relation scores are 0.26 to 0.50 for RadGraph and 0.44 to 0.64 for RadGraph-XL; fine-tuned on 70 reports, 0.44 to 0.85.
- Elsewhere the lead is smaller, sometimes inside the noise. Per modality, the fine-tuned model leads the better released one by 0.019 on ultrasound entities and 0.044 on CTPA entities derived. RadGraph-XL is highest in three scores, by 0.002, 0.013 and 0.114; the paper puts the last, head CT suggestive-of, at three true positives, “a minimal practical difference”. With about ten items, one item moves F1 by about 0.1 derived, whichever model leads.
The paper sums Table 3 up in two averages. Recomputed:
| Entity mean* | in the text | Relation mean* | in the text | |
|---|---|---|---|---|
| RadGraph + fine-tuning | 0.8006 | 0.801 | 0.6305 | 0.621† |
| RadGraph-XL | 0.7247 | 0.712† | 0.5516 | 0.541† |
| RadGraph | 0.7122 | 0.725† | 0.4839 | 0.484 |
Table 2 — Table 3’s averages, recomputed. * Plain means of the 16 entity and 12 relation scores; four decimals because one is exactly 0.6305. † Differs by more than rounding: the text swaps the released models’ entity averages and prints two relation averages about 0.01 low. Sources: Table 3, §3.3.
The paper calls its averages macro F1. In the text, 0.801 is a plain mean of Table 3’s entity scores, which fits. But seven of the eight F1 values in the paper’s Table 2 equal, within rounding, the harmonic mean of the printed precision and recall, which is how the DyGIE++ metric code in the release computes its pooled, micro F1; a mean of per-label F1s would usually come out lower. That fits micro but doesn’t prove it: macro precision and recall combined the same way would match too.
Tables 4 and 5 give the same models values Table 3 doesn’t: RadGraph scores 0.795 on entities in Table 4 against 0.712 averaged over Table 3, and Table 5’s 70% run, presumably the whole head CT training split, 0.849 and 0.733 against Table 3’s 0.827 and 0.600. Table 4 calls its values macro F1 too; neither table says what it averages over.
04How many reports does a new modality need?
Table 5 retrains the head CT model on growing shares of the data. If the shares are of the 300 head CT reports, 10% is 30 reports and 70% the whole training split derived.
Show as table
| Share of head CT reports | Reports* | Entity F1 | Relation F1 |
|---|---|---|---|
| 10% | 30 | 0.777 | 0.628 |
| 20% | 60 | 0.762 | 0.675 |
| 30% | 90 | 0.830 | 0.684 |
| 40% | 120 | 0.827 | 0.690 |
| 50% | 150 | 0.819 | 0.691 |
| 60% | 180 | 0.792 | 0.659 |
| 70% | 210 | 0.849 | 0.733 |
Entity F1 gets close with few reports, as the paper notes: 30 give 0.777 and 90 give 0.830, against 0.849 with 210. Relations lag: 0.684 at 90 reports against 0.733 at 210. And on 45 test reports the curve is noisy: 180 reports score 0.027 and 0.032 below 150, then 210 jump by 0.057 and 0.074 derived. With no released model on Table 5’s scale, it can’t say how few reports it takes to beat them.
Table 4’s single model for all four modalities scores 0.854 on entities and 0.694 on relations, against 0.795 and 0.526 for RadGraph and 0.745 and 0.541 for RadGraph-XL. Its numbers can’t be set against Table 3’s, so whether one model does better than four stays open.
05Reading it for RadGraph-F1
The paper scores extraction against human labels. It never runs RadGraph-F1 and leaves downstream tasks to future work; after the rebuttal, one reviewer still missed a downstream validation.
An extraction error doesn’t enter RadGraph-F1 one for one, because the same extractor reads both reports. A finding it misses in both drops out, unscored. A finding it labels differently in two wordings, present in one and uncertain in the other, or related in only one, turns a match into an error.
So relation F1 of 0.26 to 0.64 on cardiac MRI means that RG_ER there compares two unreliable readings. How much that moves scores, and the ranking of report generators, is the experiment the paper doesn’t run.
06What the paper leaves out
- The release: the abstract promises the model and data, and the rebuttal code, weights and annotations, the last under a data use agreement. On 6 October 2026 the paper page lists no dataset, and the repository, filled in February 2026, holds a copy of the DyGIE++ framework and an inference script adapted from RadGraph’s, whose last line calls a function the file never defines. There are no weights, no training configuration for these models and no evaluation script for the paper’s tables.
- A like-for-like arm: a reviewer called the comparison potentially less equitable because neither released model was fine-tuned on the new reports; the rebuttal answers that the authors’ model, pre-trained on RadGraph’s data, is in effect a fine-tuned RadGraph, and that RadGraph-XL’s annotations weren’t released (the MIMIC part went on PhysioNet in September 2025, after the review). The fine-tuned model also learned the annotators’ conventions, which the rebuttal says are finer where earlier RadGraph annotations erred, and a hit needs the exact span. The annotators were tested on part of RadGraph’s data, but their agreement with its labels isn’t given.
- Uncertainty: no repeat runs or intervals, no inter-annotator agreement, and no rule for the 70/15/15 split beyond the rebuttal’s word that the parts don’t overlap.
- Which checkpoints: none are named. The README points to RadGraph 1.0.0 and to version 0.1.18 of the radgraph package (June 2025), which ships four extractors, two of them RadGraph-XL models.
- The reports themselves: where the 800 come from, from which years and which sections were labelled are not stated.
07Takeaways for scoring CT and MRI reports
- Check the extractor on your modality before trusting RadGraph-F1 there. Off the shelf, the released models averaged 0.41 to 0.56 relation F1 per modality here. A few dozen labelled reports are enough to measure it.
- Report RG_E next to RG_ER on a new modality. Relations were the weaker half of extraction in every modality, and RG_ER depends on them.
- If you can label about a hundred reports, fine-tune. On head CT, 90 reports came within 0.02 entity F1 of 210, with no repeat runs reported.
- Name the extractor and its version. On the same reports, the two released models’ mean relation F1 differs by up to 0.13 per modality, and one package version ships four extractors.
ASources
- Guan, H., Dai, Y., Afyouni, S., …, Jiao, Z., Jones, C., Bai, H. Enhancing Radiology Report Interpretation through Modality-Specific RadGraph Fine-Tuning. MICCAI 2025, LNCS 15966, pp. 216–225. The page holds the reviews, the rebuttal and the meta-reviews. PDF · doi. No preprint found.
- Code: tonikroos7/RadGraph-Multimodality (MIT), read on 6 Oct 2026 at commit b851e5a.
- Jain, S., et al. RadGraph: arXiv:2106.14463 (2021); PhysioNet, version 1.0.0.
- Delbrouck, J.-B., et al. RadGraph-XL: Findings of ACL 2024; PhysioNet, version 1.0.0 (12 Sep 2025).
- Delbrouck, J.-B., et al. RadGraph-F1 (semantic rewards): Findings of EMNLP 2022; the radgraph package,
rewards.py. - Wadden, D., et al. DyGIE++: arXiv:1909.03546 (2019).
- Our posts on SPEC-CXR, phrase-grounded fact-checking, Merlin and Cervical-RG.