* Our comparison with Med-ST’s own Table 1.
- A captioner that sees both films. CoCa reads the current and the prior X-ray, and one new block lets each current-image token look at a window of the prior’s. Gemini split the MIMIC-CXR reports into descriptions and comparisons; Chest ImaGenome added boxes.
- The report score is no higher with the comparisons. RadGraph F1 is 23.7 with them and 24.2 without. The abstract’s “comparable to Med-Gemini” is the second, set against rows copied from the Med-Gemini paper; neither CoCa-CXR row has a test size or interval.
- The evidence of reading change is a classification table. 65.0 against 60.2 for BioViL-T’s published runs, mostly on pneumothorax, on a test part the paper doesn’t define; the ablation credits the windowed block with 4.0 points.
- What we can’t tell. Whether the prior image improves the report at all, and what text the reports were scored against. Neither code nor data is released.
01Why this paper matters for report generation
Radiologists read a chest X-ray against the previous one and say what changed. In MIMIC-CXR’s official test split, 88.6% of the 2,461 samples with a findings section have a prior study, by MAIRA-2’s count. A single-image generator can only drop those sentences or invent them; Flamingo-CXR, which shares three authors with this paper, removed every training report that mentioned a prior study for that reason.
CoCa-CXR gives the model the prior image. Diff-RRG, read earlier in this series, passes the prior image to a frozen LLM as fourteen status words; here the prior’s image tokens reach the report decoder.
tokens per film
→ findings + impression
The paper describes no step that aligns the two films, and in the windowed block each token looks at most five tokens away in the prior: 80 of 768 pixels, about a tenth of the image’s width derived. A patient positioned differently can move the matching region out of the window. The model’s size isn’t given.
02The training text: Gemini and Chest ImaGenome
The four training sets come from MIMIC-CXR’s official training split, without the MS-CXR-T pairs used for testing.
| Set | Samples | Target text | Share, stages 2–3 |
|---|---|---|---|
| 1. Clean image–report pairs (one image) | 224,487 | Findings and impression, rewritten by Gemini to what this image alone shows | 20% |
| 2. Image pairs and filtered reports | 132,320 | Findings and impression as written, kept when the report compares with the prior | 25% |
| 3. Image pairs and comparison sentences | 259,562 | The comparison sentences only, extracted by Gemini; each pair also reversed, with the text reversed | 25% |
| 4. Image pairs and abnormal organs | 758,344 | From Chest ImaGenome: “[condition] [progress] at [organ]” with a box in each image; pairs also reversed | 30% |
Table 1 — Four training sets; Gemini wrote the target text of two. Stage 1 uses set 1 alone; the shares are the batch mix of stages 2 and 3. Sources: Table 1, Sections 2 and 4.1.
Gemini, shown worked examples, rewrites the reports: as what the current image alone shows (a worsening effusion becomes a present one; an unchanged finding whose current state can’t be inferred is dropped), and as their comparison sentences, also reversed for the swapped pair (“new” becomes “resolved”). The prompts appear only in the arXiv version, the Gemini version isn’t named, and no check of either rewrite is reported. Set 3’s 259,562 samples are 98% of twice set 2’s 132,320 pairs derived, so about half its comparison text is likely Gemini’s reversals.
Set 4 comes from Chest ImaGenome’s automatically built scene graphs: boxes from an atlas-based extraction, and “improved”, “worsened” or “no change” from a rule-based reading of the reports. MS-CXR-T, the progression benchmark of section 04, was built from the same Chest ImaGenome labels, each checked by a radiologist against the report.
03Report generation: what the table can say
The paper’s report table compares CoCa-CXR with six published systems on MIMIC-CXR. All six rows are copied cell for cell from the Med-Gemini paper’s Table 8, a phrase of its caption included, and six of CoCa-CXR’s nine authors also sign that paper.
Show as table
| Row | Sections | RadGraph F1 | BLEU-4 | ROUGE-L | Source |
|---|---|---|---|---|---|
| Med-Gemini-2D | F + I | 24.4 | 20.5 | 28.3 | Med-Gemini paper |
| CoCa-CXR, description only | F + I | 24.2 | 18.6 | 27.8 | this paper |
| CoCa-CXR, with comparisons | F + I | 23.7 | 18.7 | 27.5 | this paper |
| Flamingo-CXR | F + I | 20.5 | 10.1 | 29.7 | Med-Gemini paper |
| CvT-21DistillGPT2 | F + I | 15.4 | 12.4 | 28.5 | Med-Gemini paper |
| Med-PaLM M, 12B | F | 25.2 | 10.4 | 26.2 | Med-Gemini paper |
| M² Transformer | F | 22.0 | 11.4 | – | Med-Gemini paper |
| CXR-RePaiR | F | 9.1 | 2.1 | 14.3 | Med-Gemini paper |
Two things limit the comparison:
- The test sets differ. Med-Gemini-2D’s 24.4 comes from 912 test cases (frontal views), scored with Yu et al.’s RadGraph package after lowercasing; Flamingo-CXR’s 20.5 from 1,931. CoCa-CXR gives no split, size or scorer for its rows and cites no RadGraph paper; its rebuttal points to Yu et al.
- No intervals. For scale, MAIRA-1’s bootstrap interval for RadGraph-F1 on 2,461 MIMIC-CXR test samples (findings only) runs from 23.7 to 24.8; on 912 cases it would be about 1.6 times as wide derived. Gaps of 0.2 and 0.7 sit inside it.
What the comparisons cost
The longitudinal claim sits in the last two rows: 23.7 with the comparison sentences, 24.2 without, the number in the abstract. The authors’ reading is that at the 65% accuracy they measure on progression classification, leaving the comparisons out does as well. F1’s arithmetic says more:
(their entities found in the reference) / (entities they add) > F1 / 2
# our derivation, for F1 over one report’s set of entities, or of relations
For a report scoring near 0.24, the bar is about one in eight derived. RadGraph F1 averages reports (and, in Yu et al.’s scorer, entity and relation F1), so over a test set the bar is approximate: if the comparisons were simply added to the same descriptions and scored against the same references, about one in eight or fewer of the entities and relations they added matched. Two readings fit. RadGraph F1 matches exact words, and counts change words such as “increased” as observations, so a right “progression” against the reference’s “worsening” would be a miss on both sides. Or writing comparisons also changed the descriptions. The paper can’t separate the two, and doesn’t say whether its references keep their comparison sentences.
In the paper’s Figure 2, from the MIMIC-CXR validation set, the reference has similar cardiomediastinal contours, worse edema and slightly larger effusions. The generated report has a stable heart and mediastinum, progressing edema and unchanged effusions: two comparisons right in meaning but not in wording, one wrong. With the films swapped, the edema statement reverses, as the reversed training pairs teach.
04Progression: the evidence that it reads change
The paper’s test of reading change is a classification: the decoder finishes “[condition] is” with worsened, unchanged or improved, for five findings of MS-CXR-T. The score is macro-accuracy, the mean recall over the three classes, so a constant answer gets 33.3.
Show as table
| Row | Con. | Pl. eff. | Pneumon. | Pneumoth. | Edema | Avg |
|---|---|---|---|---|---|---|
| BioViL-T | 61.1 | 67.0 | 61.9 | 42.6 | 68.5 | 60.2 |
| BioViL w/reg (printed as “BioViL”) | 56.0 | 63.0 | 60.2 | 42.5 | 67.5 | 57.8 |
| CNN + Transformer | 44.0 | 61.3 | 45.1 | 31.5 | 65.5 | 49.5 |
| CheXRelNet | 47 | 47 | 47 | 36 | 49 | 45.2 |
| Med-ST, its own Table 1 (accuracy, cross-validation) | 60.6 | 58.5 | 65.0 | 54.2 | 67.4 | 61.1 |
| Med-ST, as printed in this paper | 60.6 | 67.4 | 58.5 | 65.0 | 54.2 | 61.1 |
| CoCa-CXR (test) | 69.6± 2.5 | 68.1± 1.5 | 56.4± 0.8 | 59.3± 2.6 | 71.8± 0.8 | 65.0 |
| CoCa-CXR (val.) | 70.4± 0.5 | 69.6± 1.7 | 61.4± 1.6 | 72.8± 1.1 | 71.8± 0.3 | 69.2 |
| Without filtering the comparing pairs (set 2) | 65.2 | 71.2 | 54.9 | 57.9 | 70.9 | 64.0 |
| Without cleaning the descriptions (set 1) | 64.8 | 70.0 | 58.5 | 56.4 | 68.4 | 63.6 |
| Without stage 2 | 61.3 | 69.1 | 55.8 | 52.4 | 67.9 | 61.3 |
| Without the windowed block | 58.8 | 70.5 | 58.8 | 47.2 | 69.6 | 61.0 |
| Without the comparison sentences (set 3) | 59.8 | 65.3 | 60.6 | 47.7 | 69.9 | 60.7 |
| Without set 4’s organs and boxes | 54.2 | 69.1 | 59.9 | 53.2 | 65.9 | 60.5 |
| Without stages 1 and 2 | 58.5 | 65.5 | 58.9 | 46.2 | 62.9 | 58.4 |
| Without the contrastive loss | 57.4 | 68.5 | 49.8 | 45.0 | 70.7 | 58.3 |
CoCa-CXR is tested on a “test” part of MS-CXR-T, which is released as one set of 1,326 pairs with no split; the paper says neither how the part was drawn nor how big it is. Its validation part scores 4.2 points higher on average and 13.5 on pneumothorax, which shows how far these numbers can move between parts of the benchmark; the ± in the paper’s table are spreads over four seeds. The lead is uneven too. Pneumothorax gives 16.7 of the net 24.1 points summed over the five findings, 69%, and CoCa-CXR is 5.5 points behind on pneumonia derived.
| BioViL-T, its paper | Med-ST, its paper | CoCa-CXR | |
|---|---|---|---|
| Classifier | an image encoder and a head, fine-tuned per finding | an SVM on frozen features of both images | the next word after “[condition] is” |
| Trained on | Chest ImaGenome labels, 3,795 to 26,320 pairs per finding | MS-CXR-T itself, by 10-fold cross-validation | CXR-4, Chest ImaGenome statements included |
| Tested on | the MS-CXR-T benchmark, 1,326 pairs | the same pairs, in rotation | a “test” part of MS-CXR-T, size not given |
| Score | macro-accuracy | “accuracy” | macro-accuracy |
| Runs | 4 seeds | 3 seeds | 4 seeds |
| Average | 60.2 | 61.1 | 65.0 |
Table 2 — Three protocols in one comparison table. BioViL-T prints no average; 60.2 is the paper’s. Sources: Table 2; BioViL-T (arXiv:2301.04558), Table 2 and Appendix F; Med-ST (arXiv:2405.19654), Table 1 and Section 4.
The Med-ST row is also pasted in Med-ST’s column order (consolidation, edema, effusion, pneumonia, pneumothorax) under headers in another order, so four of its five values sit under the wrong finding derived. Its edema, 67.4, is printed as pleural effusion, and its pneumothorax, 54.2, as edema; the average is unaffected.
Within the paper, the ablation is the cleanest evidence: without the windowed block the average falls by 4.0 points, without set 4’s organs and boxes by 4.5, without the comparison sentences by 4.3. All of it is classification; no ablation touches the reports.
05What the paper leaves out
- The report test: split, size, scorer, decoding, and whether the references keep their comparison sentences.
- A report ablation: no single-image row and no report score without the windowed block, so nothing shows that the prior image improves the report.
- The MS-CXR-T split: how the 1,326 pairs became a validation and a test part, their sizes, and which pairs were kept out of training.
- The release: no code, and the data awaits institutional approval, per the rebuttal. The MICCAI text points to a supplement for per-structure box accuracy that wasn’t submitted; arXiv v1 has it.
- The reviews: the reviewer who scored 2 found the windowed block incremental to CoCa and the purpose of each training set unclear, and still recommended rejection after the rebuttal. The others (4 and 5) asked for the Gemini prompts, BLEU and ROUGE, and MAIRA-2; the rebuttal promised all three, and the final paper adds BLEU and ROUGE.
06Takeaways for building a longitudinal report generator
- Score the comparison sentences on their own. Whole-report RadGraph F1 gave 23.7 with them and 24.2 without. Finding-level checks of change exist: the temporal metric in our Diff-RRG post, and the comparison attribute of SPEC-CXR, which also scored Med-Gemini.
- Print the single-image row, and score both settings against the same references. Otherwise a longitudinal generator can’t show what the prior image adds.
- A local window over the prior counted on classification: 4.0 points in the ablation. Size it against how far patients move between films; here it reaches a tenth of the image’s width.
- Check LLM-rewritten targets before training on them. Gemini wrote the targets of 484,049 training samples here, with no check reported derived.
- Copy baselines with their protocol and their column order. One row here comes from cross-validation on the benchmark itself, with four values under the wrong finding.
ASources
- Chen, Y., Xu, S., Sellergren, A., Matias, Y., Hassidim, A., Shetty, S., Golden, D., Yuille, A. L., Yang, L. CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding. MICCAI 2025, LNCS 15965, pp. 78–88 (Johns Hopkins University, Google Research). The page holds the reviews and the rebuttal. PDF · doi · arXiv:2502.20509 v1 (27 Feb 2025; adds the Gemini prompts and per-structure box accuracy; its report table has RadGraph F1 only).
- Yang, L., Xu, S., Sellergren, A., et al. Advancing multimodal medical capabilities of Gemini (Med-Gemini): arXiv:2405.03162 (Tables 2 and 8, Section A.3).
- Tanno, R., et al. Flamingo-CXR: arXiv:2311.18260 v3. Hyland, S. L., et al. MAIRA-1: arXiv:2311.13668 (Table 5). Bannur, S., et al. MAIRA-2: arXiv:2406.04449 (Table 1).
- Bannur, S., et al. BioViL-T (CVPR 2023): arXiv:2301.04558 v2. MS-CXR-T on PhysioNet. Yang, J., et al. Med-ST: arXiv:2405.19654 v1. Chest ImaGenome on PhysioNet.
- Jain, S., et al. RadGraph: arXiv:2106.14463. Yu, F., et al. Evaluating progress in automatic chest X-ray radiology report generation. Patterns 4(9), 2023: doi; code rajpurkarlab/CXR-Report-Metric. Delbrouck, J.-B., et al. RadGraph-F1: arXiv:2210.12186.
- Yu, J., et al. CoCa: arXiv:2205.01917.
- Our posts on Diff-RRG and SPEC-CXR.