* Diff-RRG as retrained by that paper’s authors.
- Two frozen models, about 3.4M trained parameters. In the released code, the LLM reads 197 tokens of the current image, the prior report, and fourteen status words.
- The prior study gives most of the gain. Adding the prior image and report lifts CheXbert F1 from 0.389 to 0.453; the two modules take it to 0.474. No intervals, on the dataset’s 2,058 test pairs from 266 patients.
- The lead depends on whose run you read. HERGen’s row is copied from a paper with its own data curation and macro averaging; HC-LLM’s, evidently a re-run, sits below HC-LLM’s published BLEU-4.
- In the released code, the patch choice is mostly noise (our arithmetic). A year later, a temporal metric by three of the authors gives Diff-RRG 0.405 on “stable” statements and under 0.10 on every change.
01Why this paper matters for report generation
Radiologists read a follow-up chest X-ray against the previous one, and their reports say so: unchanged, improved, new. Diff-RRG’s idea is direct: compare the two images finding by finding, classify each finding as worsening, stable or improving, and tell the language model.
This post also reads the released code, the open reviews, and a 2026 follow-up, STRIVE, in which three of the four authors score Diff-RRG again.
768-d, per image*
≤ 4 patches each†
stable, improving, N/A
80 to 120 new tokens*
Two things are easy to miss. Only the current image’s tokens enter the prompt, in the paper’s Eq. 6 as in the code. And nothing lines the two images up: in the paper’s own Figure 1, and in one of the two examples of its Figure 2, the prior image is a lateral view and the current one frontal.
02The data and the scores
Longitudinal-MIMIC keeps the MIMIC-CXR patients with at least two studies that have a findings section: 26,625 of the 60,547 who have one, so 56% are left out derived. Each sample pairs a study with the one before it, and the model writes the later one’s findings. There is no path for a first visit.
The paper gives the total, 94,169 pairs, and no split sizes. The dataset paper lists 92,374 for training, 737 for validation and 2,058 for testing, the last from 266 patients. The paper reports no interval, repeat run or significance test.
The scores are BLEU-1 to 4, METEOR, ROUGE-L, and micro-averaged precision, recall and F1 over CheXbert labels. Which labels, and how “uncertain” counts, is not stated, and the release has no CheXbert scoring.
How far does carrying the prior forward go? In the dataset paper, when a label was the same in both reports the model’s report agreed 88.96% of the time; when it had changed, the report was wrong 84.42% of the time. None of the five papers read for this post scores the prior report handed back unchanged.
03The fourteen words, as released
The classifier’s targets come from “ground truth disease annotations”, the paper says. The repository’s README is specific: CheXbert labels both reports, and a finding that appears is “worsening”, one that disappears “improving”, one whose state is unchanged “stable”. An effusion that grows between two reports that both mention it is therefore “stable”. The clinical score is computed with the same labeler. (The release disagrees with itself on the sign: the README maps 1 to worsening, the model’s prompt dictionary prints 1 as “improving”.)
The classifier’s input is the difference map. Condensed from the released model file:
# one image, 14 findings
# sim: 14 × 196, a softmax over patches, so each row sums to 1
sim = softmax(scale * names @ patches.T, dim=-1)
noise = Gumbel(0, 1).sample(sim.shape) # at test time too
keep = softmax((sim + noise) / 0.3, dim=-1) > 0.2
F = stack([patches[k].mean(0) for k in keep]) + pos
D = F_current - F_prior # 14 × 768; pos cancels
state = classifier(D).argmax(-1) # one of 3, per finding
Three properties follow derived:
- At most four patches per finding. Five shares above 0.2 would exceed 1. Our simulation of the rule keeps 1.3 on average.
- The choice is mostly noise. The standard Gumbel-softmax adds noise to log-probabilities; the release adds it to probabilities that sum to 1 over 196 patches. Whatever the similarity says, a patch’s chance of being the top pick stays between 0.5% and 1.4%. The noise is redrawn at test time.
- The classifier isn’t told which finding it is judging. The positional term, which the rebuttal says carries the finding’s identity, cancels in the subtraction, in the paper’s Eqs. 3 and 4 as in the code, where one small network serves all 14 findings.
So in the release, each of the fourteen words typically rests on one or two near-random patches per image. We did not run the model, and the code behind the paper’s numbers may differ; the selection function has not changed since the first commit, in February 2025.
04Results: where the gain comes from
The ablation is the paper’s cleanest evidence: five rows that build on one single-image baseline.
Show as table
| Setting | F1 | Precision | Recall | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | ROUGE-L | METEOR |
|---|---|---|---|---|---|---|---|---|---|
| One image (baseline) | 0.389 | 0.434 | 0.353 | 0.390 | 0.231 | 0.150 | 0.104 | 0.262 | 0.146 |
| + prior image and report | 0.453 | 0.510 | 0.407 | 0.397 | 0.242 | 0.160 | 0.113 | 0.272 | 0.161 |
| … + difference map (DDM) | 0.469 | 0.518 | 0.429 | 0.402 | 0.248 | 0.167 | 0.119 | 0.275 | 0.163 |
| … + classifier (DPG) | 0.459 | 0.516 | 0.414 | 0.401 | 0.246 | 0.165 | 0.117 | 0.272 | 0.161 |
| … + both: Diff-RRG | 0.474 | 0.528 | 0.430 | 0.405 | 0.251 | 0.169 | 0.120 | 0.276 | 0.164 |
| Prior study adds | +0.064 | +0.076 | +0.054 | +0.007 | +0.011 | +0.010 | +0.009 | +0.010 | +0.015 |
| Both modules add, on top | +0.021 | +0.018 | +0.023 | +0.008 | +0.009 | +0.009 | +0.007 | +0.004 | +0.003 |
- The prior study does most of the work. Adding the prior image and report is worth +0.064 F1 and 0.009 BLEU-4. The row adds both at once; in the dataset paper’s ablation of its own model, the report alone added 0.076 F1 and the image alone 0.022.
- The two modules add +0.021 F1 and 0.007 BLEU-4. Alone, the difference map adds 0.016 F1 and the classifier 0.006. The paper doesn’t say what either single-module row feeds the LLM; the release holds only the full model.
- Nothing says these gaps exceed noise. For scale, HERGen’s bootstrap on its MIMIC-CXR test set of 2,210 pairs puts a BLEU-4 gain of 0.011 between +0.005 and +0.016 (95% interval). A reviewer called the modules’ gains limited; the rebuttal answered that such gains are considered meaningful in this field.
05Whose runs are in the table?
The comparison table marks one row, HERGen, as cited from its paper. HERGen built its own longitudinal subset (lateral images and duplicates removed) and macro-averaged its clinical scores over 14 classes; here they sit in a micro-averaged column. The arithmetic shows it: a micro F1 is the harmonic mean of precision and recall, which holds for every row except HERGen’s, printed as 0.295 where 0.343 would follow derived.
The other rows carry no mark and match neither the baselines’ own papers nor the three earlier tables on this benchmark that we checked: evidently the authors’ runs, which the paper doesn’t describe. HC-LLM is the baseline it measures itself against, in its Figure 2 and in the rebuttal:
Show as table
| Row | F1 | Precision | Recall | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | ROUGE-L | METEOR |
|---|---|---|---|---|---|---|---|---|---|
| Prefilling, in this paper’s Table 1 no mark; CheXbert, micro | 0.423 | 0.506 | 0.364 | 0.343 | 0.210 | 0.141 | 0.100 | 0.274 | 0.137 |
| Prefilling, in its own paper CheXpert labeler | 0.480 | 0.538 | 0.434 | 0.343 | 0.210 | 0.140 | 0.099 | 0.271 | 0.137 |
| HERGen, in this paper’s Table 1 † cited from its paper | 0.295 | 0.421 | 0.289 | 0.389 | 0.242 | 0.163 | 0.117 | 0.282 | 0.155 |
| HERGen, in its own paper its own curation; macro over 14 classes | 0.295 | 0.421 | 0.289 | 0.389 | 0.242 | 0.163 | 0.117 | 0.282 | 0.155 |
| HC-LLM, in this paper’s Table 1 no mark; CheXbert, micro | 0.448 | 0.488 | 0.415 | 0.404 | 0.247 | 0.164 | 0.116 | 0.271 | 0.163 |
| HC-LLM, in its own paper R2GenGPT variant; averaging not stated | 0.357 | 0.417 | 0.357 | 0.404 | 0.260 | 0.178 | 0.128 | 0.287 | 0.160 |
| Diff-RRG CheXbert, micro | 0.474 | 0.528 | 0.430 | 0.405 | 0.251 | 0.169 | 0.120 | 0.276 | 0.164 |
The leads over the re-run are 0.001 to 0.005; the deficits against HC-LLM’s published numbers go up to 0.011. BLEU-4 is 0.120 for Diff-RRG, 0.116 for the re-run and 0.128 in HC-LLM’s paper, whose two other variants report 0.127 and 0.142. The clinical columns can’t be compared across papers: HC-LLM’s doesn’t state its averaging, and the dataset paper’s model, “Prefilling” in the table (0.480 F1 there, 0.423 here), was scored with the older CheXpert labeler.
But the ranking then rests on how the baselines were run, and that can move a row a long way: R2GenGPT’s BLEU-4 is 0.087 in this table and 0.108 in STRIVE’s, for the same benchmark.
06A later check, by three of the authors
The paper has no metric for the comparison itself. In August 2026, three of its four authors posted STRIVE on arXiv, with a temporal metric, LCC: an LLM extracts each change statement from a report (a finding and one of five labels), and a rule-based scorer checks whether the generated report states the same change as the reference. They retrained Diff-RRG, like every baseline with public code, and scored it on the same 2,058 pairs.
Show as table
| System | new | increased | decreased | resolved | stable | worsened | improved | LCC-C (95% CI) | omitted |
|---|---|---|---|---|---|---|---|---|---|
| Diff-RRG | 0.000 | 0.097 | 0.066 | 0.000 | 0.405 | 0.083 | 0.074 | 0.187 0.171–0.203 | 0.71 |
| HC-LLM | 0.000 | 0.052 | 0.076 | 0.000 | 0.318 | 0.049 | 0.099 | 0.155 0.140–0.172 | 0.78 |
| Reference statements | 175 | 410 | 384 | 87 | 2,219 | 585 | 471 | 3,275 | – |
On the labels that describe a change, Diff-RRG reaches 0.097 and 0.066 and recovers none of the “new” or “resolved” statements. Overall, 71% of the reference statements have no counterpart in its reports. That is still the lowest omission rate of the ten systems retrained, and its coarse score of 0.187 (95% CI 0.171 to 0.203) sits above HC-LLM’s 0.155 (0.140 to 0.172).
Two cautions: LCC is the group’s own new metric, in a preprint whose new system wins by a wide margin; and “stable” is the common label, on 2,219 of 3,275 reference statements.
07What the paper leaves out
- The trivial row: the prior report handed back unchanged, and scores split by whether a label changed.
- The classifier: a reviewer asked for its performance. The rebuttal gave a test F1 of 0.84 and promised to add it; the final paper has neither that number nor the share of “stable” labels.
- The patches: one example pair, and no score for where the selected patches fall.
In the paper’s Figure 3, the generated report follows the current reference almost sentence by sentence, then adds that bibasilar opacities are unchanged from the prior exam. Those opacities are in the prior report; the current reference doesn’t mention them, and reports no focal consolidation or effusion. The figure highlights the phrase as text that matches a boxed patch.
08Takeaways for building a longitudinal report generator
- Give the LLM the prior report first, then measure what the image adds. Both together were worth +0.064 F1; the two modules on top, +0.021.
- Print the copy-the-prior row, and score changed and unchanged labels apart. The dataset paper’s model was wrong on 84.42% of changed labels.
- If patch selection uses Gumbel noise, add it to logits, switch it off at test time, and compare with random patches.
- Train and test a progression head on progression. MS-CXR-T holds 1,326 radiologist-checked labels of improving, stable and worsening.
- Mark every row as copied or re-run, and say how. One baseline’s BLEU-4 differs by 0.021 between two tables, three times what the two modules add to BLEU-4.
ASources
- Yun, H., Maeng, J., Kang, E., Suk, H.-I. Diff-RRG: Longitudinal Disease-wise Patch Difference as Guidance for LLM-based Radiology Report Generation. MICCAI 2025, LNCS 15966, pp. 152–161 (Korea University, Kangwon National University). The page holds the reviews and the rebuttal. PDF · doi. No arXiv version was found.
- Code: ku-milab/Diff-RRG (BSD-3-Clause), read on 5 Oct 2026 at commit
c25d90fd. - Maeng, J., Kang, E., Suk, H.-I. STRIVE: arXiv:2608.24237 v1 (25 Aug 2026).
- Zhu, Q., et al. Longitudinal-MIMIC (“Prefilling”), MICCAI 2023: arXiv:2306.08749.
- Liu, T., et al. HC-LLM, AAAI 2025: arXiv:2412.11070 v1. Wang, F., et al. HERGen, ECCV 2024: arXiv:2407.15158.
- Model cards: BiomedCLIP, BioMistral-7B. Dataset card: MS-CXR-T.
- CheXbert: arXiv:2004.09167. Gumbel-softmax: arXiv:1611.01144.