Jean-Benoit Delbrouck / Research blog

Research blog · Chest X-ray report generation

CoCa-CXR, read for report generation

CoCa-CXR reads a chest X-ray next to the previous one and writes what it sees and what changed. Its report score is no higher when it writes the changes, so this post reads what that score can see, how the training text was rewritten, and whose runs fill the comparison tables.

Paper Yixiong Chen, Shawn Xu, … Alan L. Yuille, Lin Yang. CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding. MICCAI 2025, LNCS 15965, pp. 78–88 By Jean-Benoit Delbrouck Published 6 Oct 2026 Reading time ~8 min Tags MICCAI 2025 · Main Disclosure I am the first author of RadGraph-F1, of the same family as the “RadGraph F1” this paper reports. The paper doesn’t name its implementation; its rebuttal cites another group’s version (Yu et al., 2023).
11 × 11
prior-image tokens each current-image token can attend to, on a 48 × 48 grid
23.7
RadGraph F1 with the comparisons written; 24.2 without, 24.4 for Med-Gemini-2D (copied)
+4.8
points of macro-accuracy over BioViL-T’s published runs on MS-CXR-T progression
4 of 5
Med-ST values printed under the wrong finding in the comparison table*

* Our comparison with Med-ST’s own Table 1.

TL;DR
  • A captioner that sees both films. CoCa reads the current and the prior X-ray, and one new block lets each current-image token look at a window of the prior’s. Gemini split the MIMIC-CXR reports into descriptions and comparisons; Chest ImaGenome added boxes.
  • The report score is no higher with the comparisons. RadGraph F1 is 23.7 with them and 24.2 without. The abstract’s “comparable to Med-Gemini” is the second, set against rows copied from the Med-Gemini paper; neither CoCa-CXR row has a test size or interval.
  • The evidence of reading change is a classification table. 65.0 against 60.2 for BioViL-T’s published runs, mostly on pneumothorax, on a test part the paper doesn’t define; the ablation credits the windowed block with 4.0 points.
  • What we can’t tell. Whether the prior image improves the report at all, and what text the reports were scored against. Neither code nor data is released.

01Why this paper matters for report generation

Radiologists read a chest X-ray against the previous one and say what changed. In MIMIC-CXR’s official test split, 88.6% of the 2,461 samples with a findings section have a prior study, by MAIRA-2’s count. A single-image generator can only drop those sentences or invent them; Flamingo-CXR, which shares three authors with this paper, removed every training report that mentioned a prior study for that reason.

CoCa-CXR gives the model the prior image. Diff-RRG, read earlier in this series, passes the prior image to a frozen LLM as fourteen status words; here the prior’s image tokens reach the report decoder.

Figure 1 — The prior film reaches the report decoder twice: as its own tokens and through one windowed block. States as in stage 3, read off the icons of the paper’s Figure 1: stage 1 trains CoCa alone, stage 2 the new block alone, stage 3 the block and the multimodal decoder. Sources: Sections 3 and 4.1, Figure 1.

The paper describes no step that aligns the two films, and in the windowed block each token looks at most five tokens away in the prior: 80 of 768 pixels, about a tenth of the image’s width derived. A patient positioned differently can move the matching region out of the window. The model’s size isn’t given.

02The training text: Gemini and Chest ImaGenome

The four training sets come from MIMIC-CXR’s official training split, without the MS-CXR-T pairs used for testing.

SetSamplesTarget textShare, stages 2–3
1. Clean image–report pairs (one image)224,487Findings and impression, rewritten by Gemini to what this image alone shows20%
2. Image pairs and filtered reports132,320Findings and impression as written, kept when the report compares with the prior25%
3. Image pairs and comparison sentences259,562The comparison sentences only, extracted by Gemini; each pair also reversed, with the text reversed25%
4. Image pairs and abnormal organs758,344From Chest ImaGenome: “[condition] [progress] at [organ]” with a box in each image; pairs also reversed30%

Table 1 — Four training sets; Gemini wrote the target text of two. Stage 1 uses set 1 alone; the shares are the batch mix of stages 2 and 3. Sources: Table 1, Sections 2 and 4.1.

Gemini, shown worked examples, rewrites the reports: as what the current image alone shows (a worsening effusion becomes a present one; an unchanged finding whose current state can’t be inferred is dropped), and as their comparison sentences, also reversed for the swapped pair (“new” becomes “resolved”). The prompts appear only in the arXiv version, the Gemini version isn’t named, and no check of either rewrite is reported. Set 3’s 259,562 samples are 98% of twice set 2’s 132,320 pairs derived, so about half its comparison text is likely Gemini’s reversals.

Set 4 comes from Chest ImaGenome’s automatically built scene graphs: boxes from an atlas-based extraction, and “improved”, “worsened” or “no change” from a rule-based reading of the reports. MS-CXR-T, the progression benchmark of section 04, was built from the same Chest ImaGenome labels, each checked by a radiologist against the report.

03Report generation: what the table can say

The paper’s report table compares CoCa-CXR with six published systems on MIMIC-CXR. All six rows are copied cell for cell from the Med-Gemini paper’s Table 8, a phrase of its caption included, and six of CoCa-CXR’s nine authors also sign that paper.

CoCa-CXR with its comparisonsCoCa-CXR, description onlyCopied from Med-Gemini’s Table 8RadGraph F10102030BLEU-40102030ROUGE-L0102030Findings + impressionMed-Gemini-2D24.420.528.3CoCa-CXR, description only24.218.627.8CoCa-CXR, with comparisons23.718.727.5Flamingo-CXR20.510.129.7CvT-21DistillGPT215.412.428.5Findings onlyMed-PaLM M, 12B25.210.426.2M² Transformer22.011.4not reportedCXR-RePaiR9.12.114.3
Figure 2 — In RadGraph F1, CoCa-CXR scores 0.2 below Med-Gemini-2D’s copied row without its comparisons, and 0.7 below with them. All three metrics on one 0–35 scale, in %. CoCa-CXR’s two rows are single numbers, with no interval or test-set size. Sources: Table 3; Med-Gemini (arXiv:2405.03162), Table 8.
Show as table
RowSectionsRadGraph F1BLEU-4ROUGE-LSource
Med-Gemini-2DF + I24.420.528.3Med-Gemini paper
CoCa-CXR, description onlyF + I24.218.627.8this paper
CoCa-CXR, with comparisonsF + I23.718.727.5this paper
Flamingo-CXRF + I20.510.129.7Med-Gemini paper
CvT-21DistillGPT2F + I15.412.428.5Med-Gemini paper
Med-PaLM M, 12BF25.210.426.2Med-Gemini paper
M² TransformerF22.011.4–Med-Gemini paper
CXR-RePaiRF9.12.114.3Med-Gemini paper

Two things limit the comparison:

  1. The test sets differ. Med-Gemini-2D’s 24.4 comes from 912 test cases (frontal views), scored with Yu et al.’s RadGraph package after lowercasing; Flamingo-CXR’s 20.5 from 1,931. CoCa-CXR gives no split, size or scorer for its rows and cites no RadGraph paper; its rebuttal points to Yu et al.
  2. No intervals. For scale, MAIRA-1’s bootstrap interval for RadGraph-F1 on 2,461 MIMIC-CXR test samples (findings only) runs from 23.7 to 24.8; on 912 cases it would be about 1.6 times as wide derived. Gaps of 0.2 and 0.7 sit inside it.

What the comparisons cost

The longitudinal claim sits in the last two rows: 23.7 with the comparison sentences, 24.2 without, the number in the abstract. The authors’ reading is that at the 65% accuracy they measure on progression classification, leaving the comparisons out does as well. F1’s arithmetic says more:

added sentences raise a report’s F1 exactly when
(their entities found in the reference) / (entities they add) > F1 / 2
# our derivation, for F1 over one report’s set of entities, or of relations

For a report scoring near 0.24, the bar is about one in eight derived. RadGraph F1 averages reports (and, in Yu et al.’s scorer, entity and relation F1), so over a test set the bar is approximate: if the comparisons were simply added to the same descriptions and scored against the same references, about one in eight or fewer of the entities and relations they added matched. Two readings fit. RadGraph F1 matches exact words, and counts change words such as “increased” as observations, so a right “progression” against the reference’s “worsening” would be a miss on both sides. Or writing comparisons also changed the descriptions. The paper can’t separate the two, and doesn’t say whether its references keep their comparison sentences.

The authors’ own example our observation

In the paper’s Figure 2, from the MIMIC-CXR validation set, the reference has similar cardiomediastinal contours, worse edema and slightly larger effusions. The generated report has a stable heart and mediastinum, progressing edema and unchanged effusions: two comparisons right in meaning but not in wording, one wrong. With the films swapped, the edema statement reverses, as the reversed training pairs teach.

04Progression: the evidence that it reads change

The paper’s test of reading change is a classification: the decoder finishes “[condition] is” with worsened, unchanged or improved, for five findings of MS-CXR-T. The score is macro-accuracy, the mean recall over the three classes, so a constant answer gets 33.3.

Average40506070BioViL-T 60.2Published rows, copied into the paper’s tableBioViL-T60.2BioViL w/reg (printed as “BioViL”)57.8CNN + Transformer49.5CheXRelNet45.2Med-ST (another protocol, see Table 2)61.1CoCa-CXR, four seedsTest part65.0Validation part69.2CoCa-CXR without one part (test part)Without filtering the comparing pairs (set 2)64.0Without cleaning the descriptions (set 1)63.6Without stage 261.3Without the windowed block61.0Without the comparison sentences (set 3)60.7Without set 4’s organs and boxes60.5Without stages 1 and 258.4Without the contrastive loss58.3
Figure 3 — Without the windowed block, set 4’s organs and boxes, or the comparison sentences, CoCa-CXR scores within a point of BioViL-T. Average over five findings, in %; the dashed line is the paper’s average of BioViL-T’s published runs. The paper gives seed spreads per finding for its own two rows, none for the ablation. Sources: Tables 2 and 4.
Show as table
RowCon.Pl. eff.Pneumon.Pneumoth.EdemaAvg
BioViL-T61.167.061.942.668.560.2
BioViL w/reg (printed as “BioViL”)56.063.060.242.567.557.8
CNN + Transformer44.061.345.131.565.549.5
CheXRelNet474747364945.2
Med-ST, its own Table 1 (accuracy, cross-validation)60.658.565.054.267.461.1
Med-ST, as printed in this paper60.667.458.565.054.261.1
CoCa-CXR (test)69.6± 2.568.1± 1.556.4± 0.859.3± 2.671.8± 0.865.0
CoCa-CXR (val.)70.4± 0.569.6± 1.761.4± 1.672.8± 1.171.8± 0.369.2
Without filtering the comparing pairs (set 2)65.271.254.957.970.964.0
Without cleaning the descriptions (set 1)64.870.058.556.468.463.6
Without stage 261.369.155.852.467.961.3
Without the windowed block58.870.558.847.269.661.0
Without the comparison sentences (set 3)59.865.360.647.769.960.7
Without set 4’s organs and boxes54.269.159.953.265.960.5
Without stages 1 and 258.565.558.946.262.958.4
Without the contrastive loss57.468.549.845.070.758.3

CoCa-CXR is tested on a “test” part of MS-CXR-T, which is released as one set of 1,326 pairs with no split; the paper says neither how the part was drawn nor how big it is. Its validation part scores 4.2 points higher on average and 13.5 on pneumothorax, which shows how far these numbers can move between parts of the benchmark; the ± in the paper’s table are spreads over four seeds. The lead is uneven too. Pneumothorax gives 16.7 of the net 24.1 points summed over the five findings, 69%, and CoCa-CXR is 5.5 points behind on pneumonia derived.

BioViL-T, its paperMed-ST, its paperCoCa-CXR
Classifieran image encoder and a head, fine-tuned per findingan SVM on frozen features of both imagesthe next word after “[condition] is”
Trained onChest ImaGenome labels, 3,795 to 26,320 pairs per findingMS-CXR-T itself, by 10-fold cross-validationCXR-4, Chest ImaGenome statements included
Tested onthe MS-CXR-T benchmark, 1,326 pairsthe same pairs, in rotationa “test” part of MS-CXR-T, size not given
Scoremacro-accuracy“accuracy”macro-accuracy
Runs4 seeds3 seeds4 seeds
Average60.261.165.0

Table 2 — Three protocols in one comparison table. BioViL-T prints no average; 60.2 is the paper’s. Sources: Table 2; BioViL-T (arXiv:2301.04558), Table 2 and Appendix F; Med-ST (arXiv:2405.19654), Table 1 and Section 4.

The Med-ST row is also pasted in Med-ST’s column order (consolidation, edema, effusion, pneumonia, pneumothorax) under headers in another order, so four of its five values sit under the wrong finding derived. Its edema, 67.4, is printed as pleural effusion, and its pneumothorax, 54.2, as edema; the average is unaffected.

Within the paper, the ablation is the cleanest evidence: without the windowed block the average falls by 4.0 points, without set 4’s organs and boxes by 4.5, without the comparison sentences by 4.3. All of it is classification; no ablation touches the reports.

05What the paper leaves out

  • The report test: split, size, scorer, decoding, and whether the references keep their comparison sentences.
  • A report ablation: no single-image row and no report score without the windowed block, so nothing shows that the prior image improves the report.
  • The MS-CXR-T split: how the 1,326 pairs became a validation and a test part, their sizes, and which pairs were kept out of training.
  • The release: no code, and the data awaits institutional approval, per the rebuttal. The MICCAI text points to a supplement for per-structure box accuracy that wasn’t submitted; arXiv v1 has it.
  • The reviews: the reviewer who scored 2 found the windowed block incremental to CoCa and the purpose of each training set unclear, and still recommended rejection after the rebuttal. The others (4 and 5) asked for the Gemini prompts, BLEU and ROUGE, and MAIRA-2; the rebuttal promised all three, and the final paper adds BLEU and ROUGE.

06Takeaways for building a longitudinal report generator

  1. Score the comparison sentences on their own. Whole-report RadGraph F1 gave 23.7 with them and 24.2 without. Finding-level checks of change exist: the temporal metric in our Diff-RRG post, and the comparison attribute of SPEC-CXR, which also scored Med-Gemini.
  2. Print the single-image row, and score both settings against the same references. Otherwise a longitudinal generator can’t show what the prior image adds.
  3. A local window over the prior counted on classification: 4.0 points in the ablation. Size it against how far patients move between films; here it reaches a tenth of the image’s width.
  4. Check LLM-rewritten targets before training on them. Gemini wrote the targets of 484,049 training samples here, with no check reported derived.
  5. Copy baselines with their protocol and their column order. One row here comes from cross-validation on the benchmark itself, with four values under the wrong finding.

ASources

  1. Chen, Y., Xu, S., Sellergren, A., Matias, Y., Hassidim, A., Shetty, S., Golden, D., Yuille, A. L., Yang, L. CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding. MICCAI 2025, LNCS 15965, pp. 78–88 (Johns Hopkins University, Google Research). The page holds the reviews and the rebuttal. PDF · doi · arXiv:2502.20509 v1 (27 Feb 2025; adds the Gemini prompts and per-structure box accuracy; its report table has RadGraph F1 only).
  2. Yang, L., Xu, S., Sellergren, A., et al. Advancing multimodal medical capabilities of Gemini (Med-Gemini): arXiv:2405.03162 (Tables 2 and 8, Section A.3).
  3. Tanno, R., et al. Flamingo-CXR: arXiv:2311.18260 v3. Hyland, S. L., et al. MAIRA-1: arXiv:2311.13668 (Table 5). Bannur, S., et al. MAIRA-2: arXiv:2406.04449 (Table 1).
  4. Bannur, S., et al. BioViL-T (CVPR 2023): arXiv:2301.04558 v2. MS-CXR-T on PhysioNet. Yang, J., et al. Med-ST: arXiv:2405.19654 v1. Chest ImaGenome on PhysioNet.
  5. Jain, S., et al. RadGraph: arXiv:2106.14463. Yu, F., et al. Evaluating progress in automatic chest X-ray radiology report generation. Patterns 4(9), 2023: doi; code rajpurkarlab/CXR-Report-Metric. Delbrouck, J.-B., et al. RadGraph-F1: arXiv:2210.12186.
  6. Yu, J., et al. CoCa: arXiv:2205.01917.
  7. Our posts on Diff-RRG and SPEC-CXR.