* In the released code.
- A classifier, a lookup table and a retriever feed an off-the-shelf LLM. RadAlign fine-tunes a ResNet-50 to predict five findings. In the released code, the LLM gets those five labels, the typical signs attached to them and the seven nearest training reports.
- The headline is 0.678 GREEN against 0.634, with no interval. Eight of twelve RadAlign settings, ablations included, land within 0.015 of that 2021 baseline.
- The gain is fewer omissions, not fewer false findings. In the first arXiv version’s counts, false findings go from 384 to 701 and missed findings from 880 to 630; the total barely moves (1,429 to 1,453).
- What we can’t tell. The test split and its size are not given, both baselines rest on one 2021 generator, and no radiologist read the reports.
01Why this paper matters for report generation
Most report generators train or adapt a language model on image–report pairs. RadAlign doesn’t: it fine-tunes an image classifier for under six hours on eight GPUs (the rebuttal’s figure), then asks an LLM that never sees the image to write the report. If that is enough, any team with a good classifier and an 8B open model has a report generator.
This post reads the paper’s report-generation half with three more sources: its first arXiv version, which printed more than the final one; the open reviews; and the code, released seven months after the conference.
→ 5 probabilities
pool: 22,037 images*
3 to 5 sentences asked*
The 14 criteria (heart size, lung opacity, fissure displacement…) and their 65 descriptions, as counted in the released code, came from prompting GPT-4, starting from the training reports.
02What the LLM is told about the image
The paper says the LLM is prompted with the criteria the model recognised and the predicted class. A reviewer asked whether the concept tokens reach the LLM, warning that a writer given only the class, general criteria and retrieved reports would be likely to invent signs. The rebuttal: tokens were tried and did not work without fine-tuning the LLM, so the prompt is text only.
The released code shows what that text is. Condensed from chat_bot.py:
# pred: five 0/1 labels, after thresholding the classifier
# TABLE[criterion][c]: which description goes with class c,
# e.g. "Mediastinal Shift": [4, 0, 1, 2, 2, 3]
def get_findings(TABLE, DESCRIPTIONS, pred):
present = [int(sum(pred) == 0)] + pred # normal + 5 findings
text = "Network's diagnosis findings shows that"
for criterion, index_of in TABLE.items():
# chosen by the predicted labels, not by the image
chosen = {index_of[c] for c in range(6) if present[c]}
signs = [DESCRIPTIONS[criterion][i] for i in chosen]
text += f" the {criterion} is {', '.join(signs)},"
return text
# the image-specific message; style instructions come before it
prompt = (labels_sentence(pred) + " " + get_findings(...)
+ " Report examples: " + ", ".join(seven_reports))
In this code, the similarity scores computed for each image never reach the prompt; they only feed the classifier. The description of each criterion is chosen by the predicted labels, so five yes/no labels give at most 32 different paragraphs derived. And the labels are far from certain: per-finding F1 runs from 0.473 (consolidation) to 0.820 (pleural effusion) in Table 2.
arXiv v1 shows two generated reports with the judge’s analysis (Figure 4; the MICCAI version has none). In the second, the reference describes low lung volumes with linear opacities at both bases, read as atelectasis. The generated report describes a localised opacity with volume loss in one lobe, displaced fissures, a raised hemidiaphragm and a mediastinal shift: the table’s entries for atelectasis, one per criterion.
Where the table says the mediastinum “may shift”, the report states a shift, and the judge’s analysis says it could be a clinically significant error. arXiv v1 presents the case as a sign of thoroughness.
03Results: every GREEN number on one axis
GREEN is the paper’s only report metric. Figure 2 puts all its values on one axis: Table 1, the retrieval sweep of Figure 2d and a component ablation given only in the rebuttal.
Show as table
| Setting | GREEN | Change | Source |
|---|---|---|---|
| R2GenCMN (the baseline) | 0.634 | – | Table 1 |
| GPT-4o writes: Five labels only | 0.609 | −0.025 | rebuttal |
| GPT-4o writes: + criteria descriptions | 0.647 | +0.013 | rebuttal |
| GPT-4o writes: + 7 retrieved reports | 0.678 | +0.044 | Table 1; rebuttal |
| GPT-4o mini writes: K = 0 | 0.633 | −0.001 | Figure 2d |
| GPT-4o mini writes: K = 5 | 0.635 | +0.001 | Figure 2d |
| GPT-4o mini writes: K = 6 | 0.642 | +0.008 | Figure 2d |
| GPT-4o mini writes: K = 7 | 0.646 | +0.012 | Figure 2d; Table 1 (0.648 in its left half) |
| GPT-4o mini writes: K = 8 | 0.624 | −0.010 | Figure 2d |
| GPT-4o mini writes: K = 7, ImageNet weights | 0.629 | −0.005 | Table 1 |
| Another LLM writes: GPT-3.5 Turbo | 0.648 | +0.014 | Table 1 |
| Another LLM writes: Claude 3.5 Sonnet | 0.658 | +0.024 | Table 1 |
| Another LLM writes: Llama 3.1 8B | 0.695 | +0.061 | Table 1 |
| ChatCAD with GPT-4o mini | 0.633 | −0.001 | Table 1 |
| ChatCAD with GPT-4o | 0.634 | 0.000 | Table 1 |
Five things stand out:
- Without retrieval, RadAlign ties the baseline: 0.633 against 0.634 in the sweep (GPT-4o mini, by our reading), 0.647 with GPT-4o in the rebuttal.
- Labels alone score lower. With GPT-4o and only the five labels the score is 0.609, 0.025 below the baseline; the descriptions add 0.038 and the seven reports 0.031.
- The LLM moves the score as much as the method. Five LLMs span 0.646 to 0.695, a range of 0.049 against a headline gain of 0.044. The top score is Llama 3.1, the only open-weight model (8B in the code, 7B in the paper).
- From ImageNet weights, RadAlign is below the baseline: 0.629 with GPT-4o mini. The headline starts from BioViL, itself trained on MIMIC-CXR images and reports.
- One more retrieved report moves the score by 0.022. Seven give 0.646, eight give 0.624: half the headline gain. The paper blames less relevant reports, but with sampled outputs (the released code leaves the API’s temperature at its default) and no repeat run, that drop cannot be told from noise. K = 7 was picked on these same scores.
04What the gain is made of
GREEN has an LLM judge list, for each generated report, the findings it shares with the reference and its clinically significant errors, in six categories:
# matched: findings in both reports · Ostmeier et al., 2024, Eq. 1
arXiv v1 printed the six error counts for every system. Its text points to the four categories that improved and says nothing about false findings. The MICCAI version keeps the scores, which are identical, and drops the counts.
Show as table, with every system and LLM
| System | (a) | (b) | (c) | (d) | (e) | (f) | All six | GREEN |
|---|---|---|---|---|---|---|---|---|
| R2GenCMN | 384 | 880 | 18 | 40 | 90 | 17 | 1,429 | 0.634 |
| ChatCAD + GPT-4o mini | 384 | 884 | 18 | 40 | 92 | 18 | 1,436 | 0.633 |
| ChatCAD + GPT-4o | 384 | 880 | 18 | 40 | 90 | 17 | 1,429 | 0.634 |
| RadAlign + GPT-4o mini | 645 | 701 | 29 | 55 | 44 | 10 | 1,484 | 0.648 / 0.646 |
| RadAlign + GPT-4o | 701 | 630 | 17 | 54 | 40 | 11 | 1,453 | 0.678 |
| RadAlign + GPT-3.5 Turbo | 733 | 668 | 17 | 54 | 57 | 12 | 1,541 | 0.648 |
| RadAlign + Claude 3.5 Sonnet | 749 | 651 | 11 | 63 | 48 | 8 | 1,530 | 0.658 |
| RadAlign + Llama 3.1 | 560 | 624 | 14 | 40 | 51 | 12 | 1,301 | 0.695 |
(a) false finding, (b) missed finding, (c) wrong location, (d) wrong severity, (e) comparison not in the reference, (f) omitted comparison. “All six” is our sum. arXiv v1 prints the GPT-4o mini row with two scores. Sources: arXiv v1, Tables 1 and 5.
RadAlign with GPT-4o reports 83% more false findings than the baseline and misses 28% fewer derived. It also makes fewer comparisons that the reference doesn’t make (90 to 40). Every LLM shows the pattern: false findings run from 560 to 749 against the baseline’s 384, and four of the five make more errors in total. Only Llama 3.1 makes fewer (1,301).
With the error total flat, a higher mean score has to come from the matched findings, or from how the errors fall across reports, since GREEN is averaged per report. The released prompt points to more matches: it lists what to mention, from the heart to the bones, and every normal statement the reference shares is a match (GREEN’s own worked example counts a clear costophrenic sinus). arXiv v1 printed the errors, not the matches.
The paper says its prompting joins a classifier’s accuracy to an LLM’s language while “significantly reducing hallucinations” (Section 3.3); it does not measure them. By GREEN’s closest category, false findings, arXiv v1’s counts put RadAlign above the baseline, not below, and section 02 shows one route: typical signs of the predicted class, written as observations.
05The comparison: one baseline, one unknown test set
The 0.634 that the abstract presents as the state of the art is R2GenCMN, from 2021. The other baseline, ChatCAD, hands R2GenCMN’s report and a classifier’s output to an LLM for rewriting; in arXiv v1’s table its GPT-4o row has the same score and the same six error counts as R2GenCMN. No other generator is compared.
The baselines’ scores are evidently the authors’ own runs (GREEN is more recent than both), which is the right way to compare. But the paper does not describe its test set, and the released code calls it “custom”. On ReXrank’s MIMIC-CXR board, which uses the official split, the best of 29 entries on the findings track scores 0.388 GREEN (read on 5 Oct 2026). A 2021 model at 0.634 points to a far easier set or a different scoring our observation, so 0.678 cannot be set beside any published MIMIC-CXR number.
06What the paper leaves out
- The test set: which split, how many reports, whether patients are kept apart from the retrieval pool, and any interval or repeat run. The code points to a “custom” test file of about 7,345 images, more than the 5,159 of MIMIC-CXR’s official test split, of which the first 700 are scored likely, not confirmed.
- A second view of quality: GREEN’s error counts were cut, no radiologist read the reports, and a reviewer’s suggestion to add BLEU or ROUGE was declined, although the released scorer computes both.
- The release: no weights, no split files, no GREEN step. The generation script stops after 100 images, the thresholds that turn probabilities into labels are fitted on the file passed as the test set, and the model class that the training and generation scripts import loads BioViL-T where the paper cites BioViL. Seven issues are open with no reply (five asked for the code before it came).
07Takeaways for building a chest X-ray report generator
- Keep a label-conditioned LLM as a baseline. A classifier trained in under 48 GPU-hours derived and an off-the-shelf LLM matched a trained 2021 generator here, before any retrieval.
- Give the LLM what was measured, not what is typical. RadAlign scores every description against every image; its released code then passes on the one chosen by the label. Pass the per-image choice, or mark typical signs as priors the LLM may not assert.
- Report GREEN with its parts: matched findings and the six error counts. Here a score that rose by 0.044 came with 83% more false findings, by arXiv v1’s counts.
- Try an open 8B model before an API. Llama 3.1 scored 0.695 against 0.678 for GPT-4o, and an open model can run wherever your reports have to stay.
- State the split, the count and an interval. When one more retrieved report moves the score by 0.022, a gain of 0.044 needs repeat runs.
ASources
- Gu, D., Gao, Y., Zhou, Y., Zhou, M., Metaxas, D. RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment. MICCAI 2025, LNCS 15966, pp. 484–494. The page holds the reviews and the rebuttal. PDF · doi · arXiv:2501.07525: v1 (January 2025) has the error counts and the example reports; v2 has the MICCAI text.
- Code: difeigu/RadAlign (MIT), released on 25 Apr 2026 and read on 5 Oct 2026.
- Ostmeier, S., et al. GREEN: arXiv:2405.03595; Findings of EMNLP 2024.
- Chen, Z., et al. R2GenCMN: Cross-modal Memory Networks for Radiology Report Generation, ACL-IJCNLP 2021. Wang, S., et al. ChatCAD: arXiv:2302.07257 (2023).
- Boecking, B., et al. BioViL: arXiv:2204.09817, ECCV 2022; model card.
- ReXrank leaderboard, MIMIC-CXR findings track, read on 5 Oct 2026. Johnson, A., et al. MIMIC-CXR-JPG: arXiv:1901.07042, Table 3, for the official split.