Jean-Benoit Delbrouck / Research blog

Research blog · CT report generation

Merlin, read for report generation

Merlin is the abdominal CT foundation model that later report generators are measured against. Its own report generator is a small, early part of the paper: 490 visual tokens, one linear layer, and a LoRA-tuned language model writing one organ section at a time. What it showed, how it was scored, and why the dataset matters more than the scores.

Paper Louis Blankemeier, Ashwin Kumar, … Sergios Gatidis, Akshay S. Chaudhari. Merlin: a computed tomography vision–language foundation model and dataset. Nature 652, 1318–1328 (2026), published online 4 March 2026 By Jean-Benoit Delbrouck Published 24 Sep 2026 Reading time ~11 min Disclosure I am a co-author of this paper and of the RadGraph-F1 metric it uses.
25,494
CT–report pairs released: the abdominal benchmark later models use
490
visual tokens per scan reach the language model*
0.10
recall over 30 findings, when a later paper scored Merlin’s test reports
0.293
full-report RadGraph-F1 in the table; the chart shows about 0.37*

* Our arithmetic from the Methods, and our measurement of the paper’s chart (section 05).

TL;DR
  • A proof of concept inside a much larger paper. Report generation is one of six task types. A frozen Merlin encoder passes 490 visual tokens through one linear layer into RadLlama-7B, tuned with LoRA, and the report is written one organ section at a time.
  • It beats RadFM everywhere, which says little. RadFM, the only baseline, is near zero on three of the four metrics, and the paper doesn’t say it was fine-tuned on these reports. Merlin’s scores are highest on sections that are usually normal: RadGraph-F1 is 0.88 for the adrenal glands and 0.17 for the gastrointestinal tract.
  • It under-reports, by its own account. The paper says Merlin misses positive findings, and a later re-scoring of its test reports puts recall over 30 findings at 0.10. The paper itself reports text metrics only.
  • The lasting contribution to report generation is the benchmark. Jolia and NV-Reason-CT are scored on the released 5,125-scan test split, and nnFoundation uses the dataset as its abdominal benchmark.

01Why this paper matters for report generation

Merlin is a 3D vision–language model for abdominal CT, pretrained on a single GPU with two kinds of supervision that hospitals already have: diagnosis codes and radiology reports. The paper evaluates it on six task types and 752 tasks. Report generation is one of them, and it takes two paragraphs of results, one extended-data figure, and a table and a figure in the supplement. It still deserves a close read. Later CT report generators quote Merlin as their baseline, and the CT–report pairs released with the paper are the public abdominal benchmark they are scored on. The next two posts in this series, on nnFoundation and NV-Reason-CT, both lean on it.

Figure 1 — A frozen encoder, one linear layer and a LoRA-tuned language model. The Methods don’t say the encoder is frozen; the icon in Extended Data Fig. 3a and the released code do. Sources: Methods; Extended Data Fig. 3a; the released model code.

02The data: 25,528 scans and their findings

The dataset comes from one academic medical centre: consecutive abdominal CT exams from December 2012 to October 2018, 25,528 scans from 18,321 patients, mostly inpatients (37%) and emergency patients (35%). Only the findings section of each report is used. Together the findings hold 10,051,571 tokens, about 394 per report derived, and 21% of them run past 512 tokens, which is why the text encoder is a Clinical Longformer and not a CLIP text encoder.

The reports follow an organ-by-organ template: lower thorax, liver and biliary tree, gallbladder, and so on. That structure does a lot of work in this paper. It lets regular expressions cut a report into anatomical sections, first for pretraining and then for generation.

SplitScans (paper)Findings tokens (paper)Patients (paper)Scans (release)
Training15,3316,036,64511,01015,314
Validation5,0601,985,9253,6445,055
Test5,1372,029,0013,6675,125
Total25,52810,051,57118,32125,494

Table 1 — The paper’s splits and the released ones. Split by patient, 60/20/20. The release has 34 fewer scans than the paper used derived. Sources: Supplementary Table 1; Data availability; released split sizes as reported by NV-Reason-CT (arXiv:2609.27511).

The release is what matters to everyone else: the scans with their findings and a fixed test split, under Stanford’s research-use agreement. The diagnosis codes aren’t part of it yet. The paper’s own report-generation numbers are on its internal 5,137-scan test set; twelve of those exams aren’t in the release, and the maintainers recommend the released 5,125 for evaluation.

03What the encoder learned from reports

The image encoder is trained with two losses at once: a binary cross-entropy over 1,692 phenotypes derived from the visit’s diagnosis codes, and a contrastive loss against the findings text.

Encoders
Image: 3D ResNet152, ImageNet weights inflated to 3D · Text: Clinical Longformer, 4,096-token context
Losses
Binary cross-entropy on 1,692 phenotypes + InfoNCE on findings, 512-dimensional embeddings, trained jointly
Optimiser
AdamW · LR 1×10⁻⁵ · cosine schedule · FP16 with gradient checkpointing
Compute
One 48 GB A6000 · batch 18 · about 160 h

Two results matter for our lens. Reports carry most of the signal: with reports alone, zero-shot F1 on 30 findings is 0.730, against 0.741 with the diagnosis codes added, and the authors say the extra complexity may not be worth it where such data are hard to get. And splitting helps: alternating full reports with a single anatomical section at each training step lifts that F1 from 0.656 to 0.741.

04The report generator: 490 tokens, one linear layer

For report generation the encoder’s last hidden layer is kept as a 7 × 7 × 10 grid of 2,048-dimensional features, 490 visual tokens derived. The input covers 33.6 × 33.6 × 48 cm, so each token stands for a cube of about 48 mm derived. One linear layer (8.4M parameters derived) maps the tokens to RadLlama-7B’s 4,096-dimensional width, and LoRA adapts 5% of the language model’s parameters.

The report is written one organ system at a time. Each pass gets the same 490 tokens and a prompt naming the section, and the 13 outputs are concatenated. Here is the head, condensed from the released code:

class ReportHead(nn.Module):
    def __init__(self, encoder, llm):   # RadLLaMA-7b + LoRA, r = 512
        super().__init__()
        self.encoder = encoder.requires_grad_(False)    # frozen
        self.adapter = nn.Linear(2048, 4096)   # the only new layer
        self.llm = llm

    def forward(self, ct, organ, target=""):
        feats = self.encoder(ct)                 # (B, 2048, 10, 7, 7)
        vis = feats.flatten(2).transpose(1, 2)   # (B, 490, 2048)
        vis = self.adapter(vis)                  # (B, 490, 4096)
        prompt = f"Generate a radiology report for {organ}###\n"
        txt = self.llm.embed(prompt + target)
        return self.llm(torch.cat([vis, txt], dim=1))

# inference: one pass per organ system
report = " ".join(generate(head, ct, o) for o in ORGAN_SYSTEMS)

Training uses AdamW at 1×10⁻⁴ with 500 warm-up steps, a cosine schedule set to decay over 500 epochs, and an effective batch of 48. The released demo decodes greedily, with a repetition penalty of 1.2 and at most 128 new tokens per section. The training split and the number of epochs actually run aren’t given.

MerlinnnFoundation stackNV-Reason-CT
Input1.5 × 1.5 × 3 mm, 33.6 × 33.6 × 48 cm1 mm, 19.2 cm cube2 mm, 38.4 cm cube
Encoder3D ResNet152, frozen3D ViT, frozen3D ViT, fine-tuned
Tokens to the LLM490 · about 48 mm each1,728 · 16 mm each13,824 · 16 mm each
Bridgeone linear layer3D adapter, 2×2×2 mergeper-token projector
Language modelRadLlama-7B, LoRAQwen2.5-VL-3B, LoRAQwen3.5-4B, full fine-tuning
Outputone section per passwhole reportreasoning, report, finding list

Table 2 — What the language model sees, here and in two later systems. Merlin’s tokens are three times coarser along each axis derived. The nnFoundation column is its chest setting; on Merlin’s data it resizes the whole abdomen into the same 192³ grid, so its tokens are coarser there. This is context, not a controlled comparison. Sources: Methods; our posts on nnFoundation and NV-Reason-CT.

05Results: every section, four metrics

The comparison is with RadFM, a generalist model whose training data include abdominal CT, on the internal test set, with four text metrics: BLEU, ROUGE-2, BERT score and RadGraph-F1.

MerlinRadFMRadGraph-F100.51ROUGE-200.51BERT score00.51BLEU00.51Adrenal glandsSpleenPancreasGallbladderVasculatureLymph nodesPelvic organsKidneys and uretersLiver and biliary treePeritoneal cavityLower thoraxMusculoskeletalGastrointestinal tractFull report
Figure 2 — Merlin beats RadFM everywhere, and scores track how routine the section is. One row per report section, sorted by Merlin’s RadGraph-F1; the full findings are the last row. Means on the internal test set (n = 5,137); the table gives no confidence intervals. Source: Supplementary Table 8.
Show as table
SectionBLEU: RadFMBLEU: MerlinROUGE-2: RadFMROUGE-2: MerlinBERT score: RadFMBERT score: MerlinRadGraph-F1: RadFMRadGraph-F1: Merlin
Lower thorax0.0010.0190.0700.3320.4060.6150.0200.319
Liver and biliary tree0.0010.2690.0250.3890.3280.6410.0800.380
Gallbladder0.0000.0060.0060.6320.5340.8510.1520.721
Spleen0.0000.0020.0040.7100.3820.8530.2830.805
Pancreas0.0000.0010.0100.7000.4470.8490.0910.748
Adrenal glands0.0060.0300.0670.8820.4900.9420.1060.879
Kidneys and ureters0.0050.2690.0400.3850.3680.6540.0910.387
Gastrointestinal tract0.0010.0130.0370.1520.3980.5310.0920.167
Peritoneal cavity0.0000.2060.0050.3900.3870.7020.0500.335
Pelvic organs0.0000.2330.0090.3580.3280.6560.0360.432
Vasculature0.0000.0260.0040.4850.2320.7480.0060.548
Lymph nodes0.0030.0230.1190.6010.5020.7750.0310.542
Musculoskeletal0.0010.0460.0180.3030.4490.6890.0080.293
Full report0.0000.1020.0110.2620.2240.5880.0080.293

Four things stand out:

  1. Merlin wins all 56 comparisons, against a baseline near zero. On full reports RadFM gets BLEU 0.000, ROUGE-2 0.011 and RadGraph-F1 0.008. The paper doesn’t say whether RadFM was fine-tuned on these reports. Scores this low suggest it was used as released likely, not confirmed: CT-FineBench gets a similar near-zero overlap from RadFM on this dataset, also without saying how it was run. If so, the gap measures training on this hospital’s report format more than image understanding.
  2. Scores are highest where “normal” is the usual answer. RadGraph-F1 is 0.88 for the adrenal glands, 0.81 for the spleen and 0.75 for the pancreas, three sections that read “normal” in all three reference reports the paper shows our observation. It is 0.17 for the gastrointestinal tract and about 0.3 for the musculoskeletal system, the lower thorax and the peritoneal cavity.
  3. Low BLEU here often means short, not wrong. BLEU is below 0.05 on 9 of the 13 sections derived, including the spleen (0.002), where ROUGE-2 is 0.71. The scorer the authors shared is a BLEU-4, and under it a perfect match scores 0.001 on a two-token section and 0.03 on a three-token one derived. “Spleen: normal” can’t score higher, and the adrenal glands’ 0.030 is close to that ceiling. BERT score has the opposite problem: RadFM stays between 0.22 and 0.53 while its other scores are mostly near zero.
  4. The full report scores below most of its parts. Its RadGraph-F1 is 0.293 in the table, against a mean of 0.504 over the 13 sections derived. A report-level score isn’t an average of section scores: the long, finding-heavy sections hold most of the entities our observation.
One value, two versions our observation

In Supplementary Table 8, the full-report RadGraph-F1 pair (0.008 for RadFM, 0.293 for Merlin) is identical to the musculoskeletal row above it. In the chart of the same data, Extended Data Fig. 3b, Merlin’s full-report bar is clearly taller than the musculoskeletal one. We measured the bars: the 13 section bars reproduce the table to within 0.003, and the full-report bar comes out at about 0.37. The other three full-report values agree between table and chart, and the pair is already in the June 2024 preprint. Until this is clarified, quote the value with its source.

06What the errors look like

The paper annotates three generated reports phrase by phrase, and states the main weakness itself: Merlin tends to under-report positive findings. The examples show more:

  • Missed and partial findings: gallstones that are in the image and in the radiologist’s report; one side of bilateral pleural effusions.
  • Invented references: series and image numbers, such as “3/297”, that the model has no way to know.
  • Contradictions: a post-cholecystectomy state in one section and a normal gallbladder in the next; a renal cyst in one section and normal kidneys in another.
  • Length: 274 words, against 154 for the radiologist, in one case.

The annotations catch the radiologists too. One reference report calls the pancreas normal after a Whipple procedure, a template default that was never edited. Reference reports are noisy, and every metric above inherits that noise.

What one section at a time costs our observation

Two of these errors follow from the design. Each section is generated in its own pass, with no view of what the other passes wrote, so nothing ties the gallbladder section to the biliary one. And a token that summarises a 48 mm cube can’t point to a slice: the series and image numbers are habits copied from the training reports. Removing them from the training targets would fix the second problem at no cost.

07What happened next: the benchmark

Merlin’s test split became the place where abdominal CT report generators are compared. Its baseline numbers travel with it, and they don’t travel unchanged.

Merlin, four waysOther systems00.10.20.30.4Merlin, in its own chartpaper · 5,137 scans≈0.37NV-Reason-CTNV paper · 5,1250.322JoliaJolia paper · 5,1250.317Merlin, in its own tablepaper · 5,137 scans0.293Merlin, re-run by JoliaJolia paper · 5,1250.237Merlin, scored by CT-FineBenchCT-FineBench · 5,0820.187MedGemma 1.5Jolia paper · 5,1250.032Med3DVLMJolia paper · 5,1250.010RadFMpaper · 5,137 scans0.008
Figure 3 — One model, four RadGraph-F1 values. RadGraph-F1 for generated findings on Merlin’s test split, by who computed it. NV-Reason-CT was first tuned for one epoch on Merlin’s training reports. Sources: Supplementary Table 8 and Extended Data Fig. 3b; Jolia (arXiv:2606.24570), Table 4; CT-FineBench (arXiv:2604.24001), Table 3; NV-Reason-CT (arXiv:2609.27511), Table 6.

How far the later systems lead depends on which Merlin they are compared with. Against the Jolia authors’ re-run (0.237), Jolia and NV-Reason-CT are 0.08 ahead. Against Merlin’s own table they are 0.02 to 0.03 ahead, and against its chart they are behind. Only the first comparison stays within one evaluation setup: the Jolia team scored Merlin and Jolia, and NV-Reason-CT followed that setup for its own score. It is the one to trust. At least two things can move the number: RadGraph-F1 has three levels (simple, partial, complete), which papers rarely name, and two generator checkpoints have been released, in September 2025 and May 2026.

RadGraph-F1 is still a text metric. The finding-level score came from DKFZ’s Resolution Meets Reduction paper, which ran a 30-finding classifier over Merlin’s and Jolia’s official test reports. Merlin’s get a macro-F1 of 0.097, with recall 0.096 and precision 0.216, against 0.238 for Jolia and 0.47 to 0.49 for that paper’s own models. That puts a number on the remark about under-reporting.

The nnFoundation paper took a different route: it put Merlin’s encoder and eight others behind one fixed language-model stack. That isolates the encoder, which the report experiment in this paper doesn’t.

08What the paper leaves out

  • Clinical accuracy: no finding-level precision or recall for the generated reports, although the paper labels 30 findings on the same test set for its classification task. Others had to supply it (section 07).
  • A trivial baseline: no “always normal” output per section, so high section scores can’t be read.
  • The RadFM protocol: fine-tuned or not, its prompt and its decoding.
  • Generator ablations: another encoder, an unfrozen one, more tokens, or the whole report in one pass.
  • Uncertainty and scope: confidence intervals appear only as error bars on the chart’s section bars, the RadGraph-F1 variant isn’t named, and report generation isn’t part of the external validation.

The authors don’t claim otherwise. They call this an early demonstration, and list the adapter and the language model as the parts to tune next.

09Takeaways for building a CT report generator

  1. Quote Merlin’s report numbers with their source. The same model has a RadGraph-F1 of 0.187, 0.237, 0.293 or about 0.37, depending on who computed it.
  2. Score findings, not only text. A recall of 0.10 over 30 findings is invisible in BLEU, ROUGE-2, BERT score and RadGraph-F1.
  3. Score per section, next to a trivial baseline, and mind section length. A 0.88 on a section that is nearly always normal says little, and BLEU-4 tops out at 0.03 on a three-token section.
  4. Generate the report as a whole, and clean the targets. Independent passes contradicted each other, and slice numbers in the training reports came back as inventions.
  5. Use the dataset. A public abdominal set with a fixed 5,125-scan test split is this paper’s durable contribution to report generation.

ASources

  1. Blankemeier, L., Kumar, A., et al. Merlin: a computed tomography vision–language foundation model and dataset. Nature 652, 1318–1328 (2026). Preprint with the supplement: arXiv:2406.06512.
  2. Merlin code · weights · dataset (Stanford AIMI).
  3. Jolia: arXiv:2606.24570 (2026), Table 4. CT-FineBench: arXiv:2604.24001 (2026), Table 3. Resolution Meets Reduction: arXiv:2608.08713 (2026), Table 7.
  4. NV-Reason-CT: arXiv:2609.27511 (2026), Table 6; our post, NV-Reason-CT, read for report generation.
  5. nnFoundation: arXiv:2609.26924 (2026); our post, nnFoundation, read for report generation.
  6. Wu, C., et al. RadFM: arXiv:2308.02463 (2023).
  7. Delbrouck, J.-B., et al. RadGraph-F1: arXiv:2210.12186 (2022).