Jean-Benoit Delbrouck / Research blog

Research blog · CT report generation

The 3D MLLM design space, read for report generation

A MICCAI 2025 study swaps the language model, the projector, the fine-tuning method and the input size of a 3D CT report generator, one at a time, on 1,287 training scans. No change to the model lifts its GREEN score by more than 0.006, while adding canned “normal” sentences to its output lifts it by 0.078. What that says about small-data training, and about the metric.

Paper Mohammed Baharoon, Jun Ma, … Augustin Toma, Bo Wang. Exploring the Design Space of 3D MLLMs for CT Report Generation. MICCAI 2025, LNCS 15965, pp. 237–246 By Jean-Benoit Delbrouck Published 5 Oct 2026 Reading time ~8 min Tags MICCAI 2025 Disclosure I am a co-author of GREEN, the metric this paper optimises and this post examines.
1,287
training scans, each seen 150 times; 400 more for testing
256
visual tokens per scan reach the language model*
+0.006
GREEN: the most any change to the model adds to the baseline’s 0.366
+0.078
GREEN from canned “normal” sentences added to the output, after the yes/no answers

* From the M3D paper and the released code.

TL;DR
  • One stack, one change at a time, no repeat runs reported. A 3D ViT, a projector and an LLM, 1,287 AMOS-MM training scans, and 400 test scans scored by GREEN, an LLM judge.
  • No change to the model beats the baseline by more than 0.006. LLMs from 2B to 14B land within 0.029 GREEN of each other, larger inputs lower every metric, and a frozen LLM beats LoRA, DoRA and full fine-tuning.
  • The big gain is post-processing, and mostly one metric sees it. A yes/no model’s answers and canned normal sentences lift GREEN from 0.366 to 0.470. The canned sentences give 0.078 of that, while RaTEScore moves 0.002 and BLEU falls 0.057.
  • What we can’t tell. There are no seeds or intervals, and no radiologist read the augmented reports. In the authors’ own examples, two of them call the uterus normal for patients whose reference reports describe a prostate.

01Why this paper matters for report generation

The systems read so far in this series were built on datasets of 25,000 scans or more. Many teams start with a thousand or so. This paper works at that scale and asks their first questions: which language model, which projector, whether to fine-tune it, and what input size.

The design is one stack, M3D’s, with one change at a time. Every table repeats the same baseline row, so nearly every other row differs from it in one thing. The work began as an entry to the MICCAI 2024 AMOS-MM challenge, where the authors report second place on the hidden test set. The code and the reviews are public, and this post uses both.

Figure 1 — The baseline: 256 visual tokens into a frozen LLM. The paper gives the volume and patch sizes. Which weights train, the HU window, the per-region prompt and the token count come from the released code (training, preprocessing) and the M3D paper. Sources: Sections 2 and 3.2.

Three things in it are easy to miss:

  • A 4 × 8 × 8 grid. Each of the 256 tokens stands for an eighth of the scan’s width and height and a quarter of its length derived. It is the smallest budget in this series: Merlin passes 490 tokens, the nnFoundation stack 1,728 and NV-Reason-CT 13,824.
  • One window for every region. −160…240 HU is the abdominal soft-tissue window, which saturates air-filled lung, chest scans included our observation.
  • “Frozen” is not untouched. The embeddings and output head that keep training hold about 0.2B parameters in Phi-3 mini derived.

02The data, the recipe and the judge

AMOS-MM pairs CT scans with findings written per region. Nearly every case has abdominal findings (99.8%), most have pelvic ones (86.9%) and fewer than a third have chest findings (30.3%). The paper trains on the challenge’s 1,287 training cases and tests on its 400 validation cases, because the test reports were never released. There is no third split for development.

Every row gets the same recipe: 150 epochs, batch size 4, a learning rate of 5×10⁻⁵ with a cosine schedule and, in the released code, one region’s findings per pass.

The score is GREEN, the challenge’s metric. A fine-tuned LLM reads the reference and the generated findings, lists the findings they share and the clinically significant errors, and returns:

GREEN = matched / (matched + significant errors)
# matched: findings in both reports · Ostmeier et al., 2024, Eq. 1

Normal findings count as matches: the GREEN paper’s own training example counts a clear costophrenic sinus. The judge the challenge prescribes is a 7B checkpoint whose model card describes it as tuned for chest X-ray reports; the GREEN paper tests abdominal CT on 15 report pairs.

The paper’s average is the plain mean of the three regions derived, so the chest weighs a third although only 30% of cases have chest findings. It also reports RaTEScore, an entity-level metric, and BLEU, ROUGE and METEOR (BLEU-2 and ROUGE-L in the released code).

03Results: the design space on one axis

Since all four tables share the baseline row, every setting can be drawn as a change from it.

Change to the model or its trainingPost-processing of the generated reportbaseline 0.366GREENChange−0.1−0.050+0.05+0.1LLMbaseline: Phi-3 mini 4BPhi-3 medium 14B0.370+0.004M3D Phi-30.365−0.001Gemma 2B0.359−0.007Llama 3.1 8B0.359−0.007Mistral v0.3 7B0.353−0.013Qwen2.5 3B0.341−0.025Projectorbaseline: M3D poolingMLP0.346−0.020TokenPacker0.343−0.023LLM tuningbaseline: frozenLoRA0.336−0.030DoRA0.321−0.045Full fine-tuning0.271−0.095Input volumebaseline: 32 × 256²64 × 512², AnyRes0.344−0.02232 × 512²0.328−0.038Additionsbaseline: none+ segmentation mask0.372+0.006+ impressions as target0.356−0.010Post-processingbaseline: none+ BQ0.392+0.026+ BQ + naive normality0.470+0.104
Figure 2 — No change to the model beats the baseline by more than 0.006; post-processing adds 0.104. Each bar is the change in GREEN when one thing is changed in the baseline. The two post-processing rows are cumulative. No repeat runs or intervals are reported; 400 test cases. Sources: Tables 1–4.
Show as table
SettingGREENChangeRaTEScoreBLEUROUGEMETEOR
Baseline0.366–0.5730.2720.3840.357
LLM: Phi-3 medium 14B0.370+0.0040.5720.2690.3820.356
LLM: M3D Phi-30.365−0.0010.5710.2760.3940.360
LLM: Gemma 2B0.359−0.0070.5720.2740.3920.365
LLM: Llama 3.1 8B0.359−0.0070.5700.2640.3790.355
LLM: Mistral v0.3 7B0.353−0.0130.5690.2580.3740.346
LLM: Qwen2.5 3B0.341−0.0250.5540.2520.3780.342
Projector: MLP0.346−0.0200.5630.2680.3860.359
Projector: TokenPacker0.343−0.0230.5510.2590.3720.349
LLM tuning: LoRA0.336−0.0300.5490.2530.3610.343
LLM tuning: DoRA0.321−0.0450.5460.2510.3600.343
LLM tuning: Full fine-tuning0.271−0.0950.5280.2320.3410.323
Input volume: 64 × 512², AnyRes0.344−0.0220.5600.2560.3790.344
Input volume: 32 × 512²0.328−0.0380.5510.2500.3700.339
Additions: + segmentation mask0.372+0.0060.5810.2710.3920.357
Additions: + impressions as target0.356−0.0100.5710.2730.3880.366
Post-processing: + BQ0.392+0.0260.5990.2640.4100.384
Post-processing: + BQ + naive normality0.470+0.1040.6010.2070.3900.384

Five things stand out:

  1. The LLM hardly matters, while it is frozen. Seven LLMs from 2B to 14B span 0.029 GREEN. The cleanest pair is in one family: Phi-3 medium (14B) gains 0.004 over Phi-3 mini (4B), a row added after a reviewer asked for it. M3D’s Phi-3, already trained on 3D medical image–text pairs, scores 0.365 against 0.366. All seven are frozen, so this says nothing about fine-tuned LLMs with more data.
  2. Freezing beats tuning, on every metric. LoRA costs 0.030 GREEN, DoRA 0.045 and full fine-tuning 0.095, and no metric reverses that order. The four settings share one schedule, and the authors point to overfitting. For scale, the nnFoundation stack tunes LoRA for about 11 epochs on CT-RATE, a dataset of 25,692 scans.
  3. More voxels made it worse. Four times the patches (32 × 512²) cost 0.038 GREEN. AnyRes at 64 × 512², eight crops plus a downsized overview, stays 0.022 below. All five metrics drop in both. The authors point to the mismatch with a ViT pretrained at 32 × 256².
  4. The projectors compress differently, not more. In the released code all three hand the LLM 256 tokens. M3D’s average pooling beats a learned mix over token positions (the “MLP”, −0.020) and TokenPacker (−0.023) on GREEN.
  5. The mask result is too small to call. A TotalSegmentator mask, encoded by the same ViT in the released code, adds 0.006 GREEN and 0.008 RaTEScore. The abstract lists it as a finding, but with no repeat runs reported it can’t be told from noise.

04The augmentation: where +0.104 comes from

Two steps run on the generated text, in this order. Binary questioning (BQ) asks a second Phi-3 model, which also sees the scan, a fixed list of yes/no questions about common findings, 22 in the released code. Each answer becomes a fixed sentence, unless the report already mentions the finding. Naive normality (NN) uses no model. A list pairs keywords with a canned normal sentence, 45 in the release (13 chest, 15 abdomen, 17 pelvis), and every sentence whose keywords are missing is added. Condensed from the released code:

def naive_normality(report, knowledge_base):
    # knowledge_base: {keywords: canned normal sentence}
    parts = report.split(".")
    for keywords, sentence in knowledge_base.items():
        mentioned = any(all(k in p.lower() for k in keywords)
                        for p in parts)
        if not mentioned:
            parts.insert(0, " " + sentence)   # the scan is never read
    return ".".join(parts).strip()
Step 1: binary questioning (BQ)change from the baseline−0.050+0.05+0.1+0.15+0.2+0.017+0.033+0.027+0.026+0.026+0.026+0.027−0.008Step 2: naive normality (NN)change from step 1−0.050+0.05+0.1+0.15+0.2+0.027+0.024+0.182+0.078+0.002−0.0200.000−0.057GREEN · chestGREEN · abdomenGREEN · pelvisGREEN · mean of regionsRaTEScoreROUGEMETEORBLEU
Figure 3 — Binary questioning lifts four of five metrics; naive normality lifts GREEN, mostly in the pelvis. Both panels share one scale. The paper reports p < 0.05 for GREEN and RaTEScore, baseline against both steps together; the steps are not tested separately. Source: Table 4.
Show as table
GREEN
StageChestAbdomenPelvisMeanRaTEScoreBLEUROUGEMETEOR
Baseline0.2430.3580.4990.3660.5730.2720.3840.357
+ BQ0.2600.3910.5260.3920.5990.2640.4100.384
+ BQ + NN0.2870.4150.7080.4700.6010.2070.3900.384
Step 1 (BQ), change+0.017+0.033+0.027+0.026+0.026−0.008+0.026+0.027
Step 2 (NN), change+0.027+0.024+0.182+0.078+0.002−0.057−0.0200.000

The pelvis supplies 78% of naive normality’s 0.078, and naive normality three quarters of the headline gain derived. The abstract’s 10% is ten points, 0.366 to 0.470, a 28% relative gain.

The formula in section 02 can explain it. A canned sentence that the reference also contains adds a match and removes no error: a report with 2 matches and 3 errors scores 0.40, and four shared normal sentences take it to 0.67 derived. The authors suspect that pelvic reports hold more normal findings, and the pelvis list is the longest. They see the problem themselves: RaTEScore’s stability, they write, may show its robustness to “these types of tricks”.

The authors’ own examples our observation

The MICCAI version has no examples of generated reports; the arXiv version adds three. In both pelvis examples the reference describes the prostate, and the augmented report also declares the uterus and the adnexal regions normal. The chest example opens with a sentence about the seminal vesicles. The released lists explain it: the pelvis list holds a normal prostate and a normal uterus with no check of the patient’s sex, and the chest list has a seminal-vesicle entry.

BQ earns its place in the first example: it adds the enlarged, calcified prostate that the generator had missed.

All three reviewers asked whether the GREEN gain means better reports; one warned that added sentences could contradict the model’s own text. The authors replied that lexical metrics miss clinical meaning and that clinical templates list every organ. But RaTEScore isn’t lexical: the paper’s tables file it under clinical metrics, and naive normality didn’t move it.

05What the paper leaves out

  • Uncertainty: no seeds or intervals for the design-space tables. A reviewer asked for tests; the final version adds the single p-value quoted under Figure 3.
  • Training length: every setting gets 150 epochs. An open issue on the repository reports GREEN 0.448 for the Llama setting after 10 epochs, against 0.359 in the paper’s table, and adds that the reports may not be very good. Nobody has answered or confirmed it. If it holds, training length moves GREEN more than the choice of LLM or projector.
  • The yes/no model: no accuracy for its 22 answers, although they are written into the report as findings.
  • Readers and other data: no radiologist read the augmented reports, and everything comes from two hospitals in one city. The rebuttal cites compute limits for leaving out CT-RATE.
  • The challenge result: no test-set score is given, and the public leaderboards still show placeholders (5 Oct 2026).

06Takeaways for building a CT report generator

  1. With about a thousand scans, start with the LLM frozen. LoRA cost 0.030 GREEN and full fine-tuning 0.095 under one schedule. Train the encoder, the projector, the embeddings and the output head, as the released code does.
  2. Don’t pay for a bigger LLM at this scale. 14B against 4B in one family: +0.004.
  3. Match the input to the encoder’s pretraining size. Four times the patches cost 0.038. If you need resolution, plan to retrain the encoder.
  4. Stress-test a metric before optimising it. Pad a baseline with canned normal sentences. If one metric jumps (GREEN, +0.078) while an entity-level one stays put (RaTEScore, +0.002), report both.
  5. For completeness, ask a model, and gate the answers. BQ’s checklist lifted four of five metrics. Canned normals need at least the patient’s sex, the scanned region and what the report already says.

ASources

  1. Baharoon, M., Ma, J., Fang, C., Toma, A., Wang, B. Exploring the Design Space of 3D MLLMs for CT Report Generation. MICCAI 2025, LNCS 15965, pp. 237–246. The page holds the reviews and the rebuttal. PDF · doi · arXiv:2506.21535 (same tables; adds the example reports as Figure 3).
  2. Code: bowang-lab/AMOS-MM-Solution (MIT), read on 5 Oct 2026.
  3. AMOS-MM challenge: Codabench competition 3137 (data, evaluation, rules) and the AMOS site.
  4. Ostmeier, S., et al. GREEN: arXiv:2405.03595 (2024); GREEN-RadLlama2-7b model card.
  5. Zhao, W., et al. RaTEScore: arXiv:2406.16845 (2024).
  6. Bai, F., et al. M3D: arXiv:2404.00578 (2024).
  7. Our posts on Merlin, nnFoundation and NV-Reason-CT.