* Our arithmetic from the Methods, and our measurement of the paper’s chart (section 05).
- A proof of concept inside a much larger paper. Report generation is one of six task types. A frozen Merlin encoder passes 490 visual tokens through one linear layer into RadLlama-7B, tuned with LoRA, and the report is written one organ section at a time.
- It beats RadFM everywhere, which says little. RadFM, the only baseline, is near zero on three of the four metrics, and the paper doesn’t say it was fine-tuned on these reports. Merlin’s scores are highest on sections that are usually normal: RadGraph-F1 is 0.88 for the adrenal glands and 0.17 for the gastrointestinal tract.
- It under-reports, by its own account. The paper says Merlin misses positive findings, and a later re-scoring of its test reports puts recall over 30 findings at 0.10. The paper itself reports text metrics only.
- The lasting contribution to report generation is the benchmark. Jolia and NV-Reason-CT are scored on the released 5,125-scan test split, and nnFoundation uses the dataset as its abdominal benchmark.
01Why this paper matters for report generation
Merlin is a 3D vision–language model for abdominal CT, pretrained on a single GPU with two kinds of supervision that hospitals already have: diagnosis codes and radiology reports. The paper evaluates it on six task types and 752 tasks. Report generation is one of them, and it takes two paragraphs of results, one extended-data figure, and a table and a figure in the supplement. It still deserves a close read. Later CT report generators quote Merlin as their baseline, and the CT–report pairs released with the paper are the public abdominal benchmark they are scored on. The next two posts in this series, on nnFoundation and NV-Reason-CT, both lean on it.
→ 7 × 7 × 10 × 2,048
02The data: 25,528 scans and their findings
The dataset comes from one academic medical centre: consecutive abdominal CT exams from December 2012 to October 2018, 25,528 scans from 18,321 patients, mostly inpatients (37%) and emergency patients (35%). Only the findings section of each report is used. Together the findings hold 10,051,571 tokens, about 394 per report derived, and 21% of them run past 512 tokens, which is why the text encoder is a Clinical Longformer and not a CLIP text encoder.
The reports follow an organ-by-organ template: lower thorax, liver and biliary tree, gallbladder, and so on. That structure does a lot of work in this paper. It lets regular expressions cut a report into anatomical sections, first for pretraining and then for generation.
| Split | Scans (paper) | Findings tokens (paper) | Patients (paper) | Scans (release) |
|---|---|---|---|---|
| Training | 15,331 | 6,036,645 | 11,010 | 15,314 |
| Validation | 5,060 | 1,985,925 | 3,644 | 5,055 |
| Test | 5,137 | 2,029,001 | 3,667 | 5,125 |
| Total | 25,528 | 10,051,571 | 18,321 | 25,494 |
Table 1 — The paper’s splits and the released ones. Split by patient, 60/20/20. The release has 34 fewer scans than the paper used derived. Sources: Supplementary Table 1; Data availability; released split sizes as reported by NV-Reason-CT (arXiv:2609.27511).
The release is what matters to everyone else: the scans with their findings and a fixed test split, under Stanford’s research-use agreement. The diagnosis codes aren’t part of it yet. The paper’s own report-generation numbers are on its internal 5,137-scan test set; twelve of those exams aren’t in the release, and the maintainers recommend the released 5,125 for evaluation.
03What the encoder learned from reports
The image encoder is trained with two losses at once: a binary cross-entropy over 1,692 phenotypes derived from the visit’s diagnosis codes, and a contrastive loss against the findings text.
- Encoders
- Image: 3D ResNet152, ImageNet weights inflated to 3D · Text: Clinical Longformer, 4,096-token context
- Losses
- Binary cross-entropy on 1,692 phenotypes + InfoNCE on findings, 512-dimensional embeddings, trained jointly
- Optimiser
- AdamW · LR 1×10⁻⁵ · cosine schedule · FP16 with gradient checkpointing
- Compute
- One 48 GB A6000 · batch 18 · about 160 h
Two results matter for our lens. Reports carry most of the signal: with reports alone, zero-shot F1 on 30 findings is 0.730, against 0.741 with the diagnosis codes added, and the authors say the extra complexity may not be worth it where such data are hard to get. And splitting helps: alternating full reports with a single anatomical section at each training step lifts that F1 from 0.656 to 0.741.
04The report generator: 490 tokens, one linear layer
For report generation the encoder’s last hidden layer is kept as a 7 × 7 × 10 grid of 2,048-dimensional features, 490 visual tokens derived. The input covers 33.6 × 33.6 × 48 cm, so each token stands for a cube of about 48 mm derived. One linear layer (8.4M parameters derived) maps the tokens to RadLlama-7B’s 4,096-dimensional width, and LoRA adapts 5% of the language model’s parameters.
The report is written one organ system at a time. Each pass gets the same 490 tokens and a prompt naming the section, and the 13 outputs are concatenated. Here is the head, condensed from the released code:
class ReportHead(nn.Module):
def __init__(self, encoder, llm): # RadLLaMA-7b + LoRA, r = 512
super().__init__()
self.encoder = encoder.requires_grad_(False) # frozen
self.adapter = nn.Linear(2048, 4096) # the only new layer
self.llm = llm
def forward(self, ct, organ, target=""):
feats = self.encoder(ct) # (B, 2048, 10, 7, 7)
vis = feats.flatten(2).transpose(1, 2) # (B, 490, 2048)
vis = self.adapter(vis) # (B, 490, 4096)
prompt = f"Generate a radiology report for {organ}###\n"
txt = self.llm.embed(prompt + target)
return self.llm(torch.cat([vis, txt], dim=1))
# inference: one pass per organ system
report = " ".join(generate(head, ct, o) for o in ORGAN_SYSTEMS)
Training uses AdamW at 1×10⁻⁴ with 500 warm-up steps, a cosine schedule set to decay over 500 epochs, and an effective batch of 48. The released demo decodes greedily, with a repetition penalty of 1.2 and at most 128 new tokens per section. The training split and the number of epochs actually run aren’t given.
| Merlin | nnFoundation stack | NV-Reason-CT | |
|---|---|---|---|
| Input | 1.5 × 1.5 × 3 mm, 33.6 × 33.6 × 48 cm | 1 mm, 19.2 cm cube | 2 mm, 38.4 cm cube |
| Encoder | 3D ResNet152, frozen | 3D ViT, frozen | 3D ViT, fine-tuned |
| Tokens to the LLM | 490 · about 48 mm each | 1,728 · 16 mm each | 13,824 · 16 mm each |
| Bridge | one linear layer | 3D adapter, 2×2×2 merge | per-token projector |
| Language model | RadLlama-7B, LoRA | Qwen2.5-VL-3B, LoRA | Qwen3.5-4B, full fine-tuning |
| Output | one section per pass | whole report | reasoning, report, finding list |
Table 2 — What the language model sees, here and in two later systems. Merlin’s tokens are three times coarser along each axis derived. The nnFoundation column is its chest setting; on Merlin’s data it resizes the whole abdomen into the same 192³ grid, so its tokens are coarser there. This is context, not a controlled comparison. Sources: Methods; our posts on nnFoundation and NV-Reason-CT.
05Results: every section, four metrics
The comparison is with RadFM, a generalist model whose training data include abdominal CT, on the internal test set, with four text metrics: BLEU, ROUGE-2, BERT score and RadGraph-F1.
Show as table
| Section | BLEU: RadFM | BLEU: Merlin | ROUGE-2: RadFM | ROUGE-2: Merlin | BERT score: RadFM | BERT score: Merlin | RadGraph-F1: RadFM | RadGraph-F1: Merlin |
|---|---|---|---|---|---|---|---|---|
| Lower thorax | 0.001 | 0.019 | 0.070 | 0.332 | 0.406 | 0.615 | 0.020 | 0.319 |
| Liver and biliary tree | 0.001 | 0.269 | 0.025 | 0.389 | 0.328 | 0.641 | 0.080 | 0.380 |
| Gallbladder | 0.000 | 0.006 | 0.006 | 0.632 | 0.534 | 0.851 | 0.152 | 0.721 |
| Spleen | 0.000 | 0.002 | 0.004 | 0.710 | 0.382 | 0.853 | 0.283 | 0.805 |
| Pancreas | 0.000 | 0.001 | 0.010 | 0.700 | 0.447 | 0.849 | 0.091 | 0.748 |
| Adrenal glands | 0.006 | 0.030 | 0.067 | 0.882 | 0.490 | 0.942 | 0.106 | 0.879 |
| Kidneys and ureters | 0.005 | 0.269 | 0.040 | 0.385 | 0.368 | 0.654 | 0.091 | 0.387 |
| Gastrointestinal tract | 0.001 | 0.013 | 0.037 | 0.152 | 0.398 | 0.531 | 0.092 | 0.167 |
| Peritoneal cavity | 0.000 | 0.206 | 0.005 | 0.390 | 0.387 | 0.702 | 0.050 | 0.335 |
| Pelvic organs | 0.000 | 0.233 | 0.009 | 0.358 | 0.328 | 0.656 | 0.036 | 0.432 |
| Vasculature | 0.000 | 0.026 | 0.004 | 0.485 | 0.232 | 0.748 | 0.006 | 0.548 |
| Lymph nodes | 0.003 | 0.023 | 0.119 | 0.601 | 0.502 | 0.775 | 0.031 | 0.542 |
| Musculoskeletal | 0.001 | 0.046 | 0.018 | 0.303 | 0.449 | 0.689 | 0.008 | 0.293 |
| Full report | 0.000 | 0.102 | 0.011 | 0.262 | 0.224 | 0.588 | 0.008 | 0.293 |
Four things stand out:
- Merlin wins all 56 comparisons, against a baseline near zero. On full reports RadFM gets BLEU 0.000, ROUGE-2 0.011 and RadGraph-F1 0.008. The paper doesn’t say whether RadFM was fine-tuned on these reports. Scores this low suggest it was used as released likely, not confirmed: CT-FineBench gets a similar near-zero overlap from RadFM on this dataset, also without saying how it was run. If so, the gap measures training on this hospital’s report format more than image understanding.
- Scores are highest where “normal” is the usual answer. RadGraph-F1 is 0.88 for the adrenal glands, 0.81 for the spleen and 0.75 for the pancreas, three sections that read “normal” in all three reference reports the paper shows our observation. It is 0.17 for the gastrointestinal tract and about 0.3 for the musculoskeletal system, the lower thorax and the peritoneal cavity.
- Low BLEU here often means short, not wrong. BLEU is below 0.05 on 9 of the 13 sections derived, including the spleen (0.002), where ROUGE-2 is 0.71. The scorer the authors shared is a BLEU-4, and under it a perfect match scores 0.001 on a two-token section and 0.03 on a three-token one derived. “Spleen: normal” can’t score higher, and the adrenal glands’ 0.030 is close to that ceiling. BERT score has the opposite problem: RadFM stays between 0.22 and 0.53 while its other scores are mostly near zero.
- The full report scores below most of its parts. Its RadGraph-F1 is 0.293 in the table, against a mean of 0.504 over the 13 sections derived. A report-level score isn’t an average of section scores: the long, finding-heavy sections hold most of the entities our observation.
In Supplementary Table 8, the full-report RadGraph-F1 pair (0.008 for RadFM, 0.293 for Merlin) is identical to the musculoskeletal row above it. In the chart of the same data, Extended Data Fig. 3b, Merlin’s full-report bar is clearly taller than the musculoskeletal one. We measured the bars: the 13 section bars reproduce the table to within 0.003, and the full-report bar comes out at about 0.37. The other three full-report values agree between table and chart, and the pair is already in the June 2024 preprint. Until this is clarified, quote the value with its source.
06What the errors look like
The paper annotates three generated reports phrase by phrase, and states the main weakness itself: Merlin tends to under-report positive findings. The examples show more:
- Missed and partial findings: gallstones that are in the image and in the radiologist’s report; one side of bilateral pleural effusions.
- Invented references: series and image numbers, such as “3/297”, that the model has no way to know.
- Contradictions: a post-cholecystectomy state in one section and a normal gallbladder in the next; a renal cyst in one section and normal kidneys in another.
- Length: 274 words, against 154 for the radiologist, in one case.
The annotations catch the radiologists too. One reference report calls the pancreas normal after a Whipple procedure, a template default that was never edited. Reference reports are noisy, and every metric above inherits that noise.
Two of these errors follow from the design. Each section is generated in its own pass, with no view of what the other passes wrote, so nothing ties the gallbladder section to the biliary one. And a token that summarises a 48 mm cube can’t point to a slice: the series and image numbers are habits copied from the training reports. Removing them from the training targets would fix the second problem at no cost.
07What happened next: the benchmark
Merlin’s test split became the place where abdominal CT report generators are compared. Its baseline numbers travel with it, and they don’t travel unchanged.
How far the later systems lead depends on which Merlin they are compared with. Against the Jolia authors’ re-run (0.237), Jolia and NV-Reason-CT are 0.08 ahead. Against Merlin’s own table they are 0.02 to 0.03 ahead, and against its chart they are behind. Only the first comparison stays within one evaluation setup: the Jolia team scored Merlin and Jolia, and NV-Reason-CT followed that setup for its own score. It is the one to trust. At least two things can move the number: RadGraph-F1 has three levels (simple, partial, complete), which papers rarely name, and two generator checkpoints have been released, in September 2025 and May 2026.
RadGraph-F1 is still a text metric. The finding-level score came from DKFZ’s Resolution Meets Reduction paper, which ran a 30-finding classifier over Merlin’s and Jolia’s official test reports. Merlin’s get a macro-F1 of 0.097, with recall 0.096 and precision 0.216, against 0.238 for Jolia and 0.47 to 0.49 for that paper’s own models. That puts a number on the remark about under-reporting.
The nnFoundation paper took a different route: it put Merlin’s encoder and eight others behind one fixed language-model stack. That isolates the encoder, which the report experiment in this paper doesn’t.
08What the paper leaves out
- Clinical accuracy: no finding-level precision or recall for the generated reports, although the paper labels 30 findings on the same test set for its classification task. Others had to supply it (section 07).
- A trivial baseline: no “always normal” output per section, so high section scores can’t be read.
- The RadFM protocol: fine-tuned or not, its prompt and its decoding.
- Generator ablations: another encoder, an unfrozen one, more tokens, or the whole report in one pass.
- Uncertainty and scope: confidence intervals appear only as error bars on the chart’s section bars, the RadGraph-F1 variant isn’t named, and report generation isn’t part of the external validation.
The authors don’t claim otherwise. They call this an early demonstration, and list the adapter and the language model as the parts to tune next.
09Takeaways for building a CT report generator
- Quote Merlin’s report numbers with their source. The same model has a RadGraph-F1 of 0.187, 0.237, 0.293 or about 0.37, depending on who computed it.
- Score findings, not only text. A recall of 0.10 over 30 findings is invisible in BLEU, ROUGE-2, BERT score and RadGraph-F1.
- Score per section, next to a trivial baseline, and mind section length. A 0.88 on a section that is nearly always normal says little, and BLEU-4 tops out at 0.03 on a three-token section.
- Generate the report as a whole, and clean the targets. Independent passes contradicted each other, and slice numbers in the training reports came back as inventions.
- Use the dataset. A public abdominal set with a fixed 5,125-scan test split is this paper’s durable contribution to report generation.
ASources
- Blankemeier, L., Kumar, A., et al. Merlin: a computed tomography vision–language foundation model and dataset. Nature 652, 1318–1328 (2026). Preprint with the supplement: arXiv:2406.06512.
- Merlin code · weights · dataset (Stanford AIMI).
- Jolia: arXiv:2606.24570 (2026), Table 4. CT-FineBench: arXiv:2604.24001 (2026), Table 3. Resolution Meets Reduction: arXiv:2608.08713 (2026), Table 7.
- NV-Reason-CT: arXiv:2609.27511 (2026), Table 6; our post, NV-Reason-CT, read for report generation.
- nnFoundation: arXiv:2609.26924 (2026); our post, nnFoundation, read for report generation.
- Wu, C., et al. RadFM: arXiv:2308.02463 (2023).
- Delbrouck, J.-B., et al. RadGraph-F1: arXiv:2210.12186 (2022).