* From the released code and configuration. The paper’s text says 1,024 queries.
- A question-aware bottleneck in front of a small LLM. Eight 32-slice blocks pass through M3D’s 3D ViT; the µ²Tokenizer mixes their tokens down to 256 for Llama-3.2-1B. The released training scripts for that LLM freeze nothing.
- Then DPO on GREEN, and the word-overlap scores follow. GREEN rises by 0.045 to 0.065 on each dataset, and ROUGE-1 and METEOR rise with it, unlike the GREEN-only gain of the design-space paper. All three compare the text with one reference.
- The larger models were run as released, going by the evaluation scripts in the repository’s history. LaMed-Phi-3-4B scores 0.009 GREEN on AMOS-MM here; the same stack, trained on AMOS-MM, scores 0.365 in the design-space paper.
- What we can’t tell. No test-set size, interval or repeat run is given, the released scripts decode by sampling, and one printed score, a METEOR of 0.876, is almost certainly a misprint.
01Why this paper matters for report generation
Two choices shape every system in this series: how much of the scan reaches the language model, and what the model is trained on. µ²LLM makes an unusual one of each. It reads 256 slices but hands the LLM only as many tokens as one 32-slice block yields, chosen with the question in view. Then it trains on its judge: the model drafts reports, GREEN ranks them, and direct preference optimisation (DPO) teaches it to prefer the winners.
So GREEN is trainer and examiner at once. The design-space paper read earlier raised GREEN by 0.078 with canned sentences that its other metrics did not reward; does training on GREEN move anything else?
→ 8 × 2,048 × 768
→ 1,024 → 1,792
→ 256 tokens*
Two of these facts come from the repository, not the text:
- 256 tokens, not 1,024. The text mentions 1,024 learnable queries, but its notation gives the tokenizer one block’s worth of output tokens, and the released configuration sets 256. That is the design-space baseline’s budget, drawn from eight times as many slices; Merlin passes 490 tokens and NV-Reason-CT 13,824.
- The training scripts freeze nothing. The paper does not say what trains. The released stage-1 scripts for Llama-3.2-1B update the ViT, the projector, the tokenizer and the whole LLM, one model per dataset (4 epochs at a learning rate of 4×10⁻⁶ in the scripts of July 2025). “1B” counts the LLM alone; as released, the tokenizer adds about 0.36B parameters derived.
02Three datasets, and what the baselines were given
The abstract announces four CT report datasets; the introduction says three, and three are described and scored. The paper gives sizes but no split and no test-set size; Table 1 adds what the evaluation scripts in the repository’s history show.
| Dataset | Scans | Who wrote the references | Scored on, in the released scripts |
|---|---|---|---|
| AMOS-MM | 2,088 chest-to-pelvis CTs, two hospitals in Shenzhen (challenge split: 1,288 / 400 / 400) | Radiologists, per region | The challenge’s validation cases, region by region; the committed script stops after 100 |
| CT-RATE | 50,188 chest CT volumes (25,692 scans), one hospital in Istanbul | Radiologists, translated from Turkish | The validation split (3,039 volumes), findings section |
| AbdomenAtlas 3.0 | 9,262 abdominal CTs from 17 public datasets | RadGPT, from tumour masks; radiologists revised masks and reports | A list named IID_test.csv, narrative reports |
Table 1 — The three datasets, all public; one has machine-written references. The last column is from scripts kept in the repository from March 2025 to February 2026, which may not be the exact runs behind the paper. Sources: Section 3, the challenge page, the CT-RATE and RadGPT papers.
evalscipt/, not the paper; dark grey is CT-CHAT, which its authors trained on CT-RATE. No intervals are reported.Show as table, with the other three metrics
| Model | ROUGE-1 | METEOR | BERTScore | GREEN | How it was run |
|---|---|---|---|---|---|
| AbdomenAtlas 3.0 · references written by RadGPT from tumour masks | |||||
| LaMed-Phi-3-4B | 0.136 | 0.058 | 0.807 | 0.011 | run as released |
| LaMed-Llama-2-7B | 0.139 | 0.060 | 0.810 | 0.009 | run as released |
| RadFM-14B | 0.037 | 0.013 | 0.794 | 0.000 | run as released |
| RadGPT-N | 0.247 | 0.112 | – | – | matches RadGPT’s Table 4 |
| µ²LLM-1B (SFT) | 0.529 | 0.295 | 0.891 | 0.281 | fine-tuned on it |
| µ²LLM-1B (SFT & DPO) | 0.567 | 0.319 | 0.895 | 0.346 | fine-tuned on it |
| CT-RATE · radiologists’ reports, translated from Turkish | |||||
| LaMed-Phi-3-4B | 0.130 | 0.050 | 0.814 | 0.002 | run as released |
| LaMed-Llama-2-7B | 0.103 | 0.048 | 0.815 | 0.001 | run as released |
| RadFM-14B | 0.054 | 0.017 | 0.812 | 0.014 | run as released |
| CT-CHAT-8B | 0.294 | 0.221 | 0.815 | 0.113 | trained on it by its authors |
| µ²LLM-1B (SFT) | 0.517 | 0.330 | 0.879 | 0.384 | fine-tuned on it |
| µ²LLM-1B (SFT & DPO) | 0.539 | 0.359 | 0.890 | 0.429 | fine-tuned on it |
| AMOS-MM · radiologists’ reports, one per body region | |||||
| LaMed-Phi-3-4B | 0.126 | 0.047 | 0.821 | 0.009 | run as released |
| LaMed-Llama-2-7B | 0.163 | 0.065 | 0.823 | 0.009 | run as released |
| RadFM-14B | 0.046 | 0.015 | 0.812 | 0.001 | run as released |
| µ²LLM-1B (SFT) | 0.421 | 0.249 | 0.881 | 0.339 | fine-tuned on it |
| µ²LLM-1B (SFT & DPO) | 0.459 | 0.876 | 0.881 | 0.400 | fine-tuned on it |
Three things stand out:
- In the scripts, the baselines are public checkpoints, used as they are. The paper does not say how they were run. The evaluation scripts load LaMed, RadFM and CT-CHAT as published, LaMed and RadFM with generic prompts.
- One row matches another paper’s table. RadGPT-N’s scores (ROUGE-1 0.247, METEOR 0.112) match the 24.7 and 11.2 of the RadGPT preprint’s Table 4, which were measured on a private set from another hospital.
- Trained on the data, the same baseline is another model. LaMed-Phi-3-4B scores 0.009 GREEN on AMOS-MM here. In the design-space paper, the M3D stack with that Phi-3, trained on AMOS-MM’s 1,287 scans, scores 0.365, with the same judge checkpoint in both releases (400 validation cases, mean over regions; prompts and decoding differ).
So, by these scripts, the claim that a 1B model outperforms models of 7B to 14B compares fine-tuned with as-released.
03The tokenizer: what the ablation shows
The µ²Tokenizer starts from LinVT, a module that lets image LLMs read video: score the frame tokens, keep the top k, pool them at several scales, and let learned queries, which also read the text, aggregate them. The paper treats the eight blocks as frames and changes three steps: a learned bias per relative position in attention (RPE); 1,024 soft selections, each a weighted sum of all tokens, in place of the hard top-k (DTS); and learned weights on the pooled scales (DMTP).
linvt in the released scripts); the middle column is the change from it. One number per row, no intervals. Source: the paper’s Table 2, which names no dataset; its two µ²LLM rows are Table 1’s AMOS-MM rows.Show as table, with all five metrics
| Row | BLEU | ROUGE-1 | METEOR | BERTScore | GREEN | GREEN, change |
|---|---|---|---|---|---|---|
| Baseline (LinVT design) | 0.190 | 0.405 | 0.210 | 0.864 | 0.204 | – |
| + relative positions (RPE) | 0.281 | 0.421 | 0.236 | 0.880 | 0.277 | +0.073 |
| + soft selection (DTS) | 0.271 | 0.411 | 0.240 | 0.888 | 0.299 | +0.095 |
| + weighted pooling (DMTP) | 0.254 | 0.401 | 0.220 | 0.874 | 0.233 | +0.029 |
| All three (µ²Tokenizer) | 0.279 | 0.421 | 0.249 | 0.881 | 0.339 | +0.135 |
| All three + DPO on GREEN | 0.336 | 0.459 | 0.876 | 0.881 | 0.400 | +0.196 |
Bold: the top value of each column among the five rows before DPO. The last row’s METEOR is as printed. Source: the paper’s Table 2.
- Only GREEN and METEOR put the full tokenizer first. RPE alone matches it on ROUGE-1 (0.421) and on BLEU (0.281 against 0.279); DTS alone has the top BERTScore (0.888 against 0.881). A reviewer noted the uneven pattern.
- The text and the table disagree on DTS. The text credits it with “up to 0.2 points” of GREEN; the table shows 0.095. The rebuttal said the sentence belonged to a second, subtractive ablation and promised to print both. The final paper keeps the sentence and one table.
- The RPE row may measure more than positions. In the repository’s code until a commit of January 2026, the settings without RPE called PyTorch’s attention with the batch axis where it expects the sequence, so a query could not attend to the other 255. Whether Table 2 ran on that code is not stated our observation.
- No row removes the tokenizer, drops the question or uses one block instead of eight.
The reviewer who scored the paper lowest doubted two claims: that a question can steer feature extraction when report prompts hardly vary, and that a model fed one input shape is multi-scale. The rebuttal placed the scales in the tokenizer’s pooling; the reviewer was not convinced. The released code gives both doubts some ground our observation. The pooling averages 1, 2 and 4 neighbouring entries of a list of soft selections that has no spatial order, so a scale is not a size in the scan. And the evaluation scripts ask one of three region prompts on AMOS-MM, and one fixed sentence on CT-RATE.
04DPO on GREEN: what moves, and what can’t be told
Stage 2 needs no reward model and no reinforcement learning, only a better and a worse report for the same input. Condensed from the released code (drafting and ranking, pairing):
# 1. draft, judge, rank: for each example of the training split
drafts = []
while len(drafts) < 8:
d = model.generate(scan, question, do_sample=True,
temperature=1.0, top_p=0.9)
if usable(d): # 20+ characters, none Chinese
drafts.append(d)
scores = {d: green(reference, d) for d in drafts}
ranked = sorted(scores, key=scores.get, reverse=True)
# 2. one preference pair per example, unless every draft scores 0
if scores[ranked[0]] != 0:
pairs.append({"image": scan, "question": question,
"chosen": ranked[0], "rejected": ranked[-1]})
Each pair is the model’s own best and worst draw (0.63 and 0.1 in the paper’s schematic). DPO then raises the likelihood of the chosen report over the rejected one, relative to a frozen copy of the stage-1 model.
Which GREEN? The abstract names GREEN-RedLlama (sic), presumably the 7B RadLlama2 checkpoint, as DPO’s guide; the checkpoint that scores the results is not named. In the released scripts, a checkpoint from a folder named GREEN-RadPhi2 (2.8B) ranks the drafts and GREEN-RadLlama2-7b (6.7B) scores the results. Either way, one family of judges, fitted mostly on chest X-ray reports, trains the model and examines it.
Show as table
| Dataset | Stage | ROUGE-1 | METEOR | BERTScore | GREEN |
|---|---|---|---|---|---|
| AbdomenAtlas 3.0 | SFT | 0.529 | 0.295 | 0.891 | 0.281 |
| SFT & DPO | 0.567 | 0.319 | 0.895 | 0.346 | |
| change | +0.038 | +0.024 | +0.004 | +0.065 | |
| CT-RATE | SFT | 0.517 | 0.330 | 0.879 | 0.384 |
| SFT & DPO | 0.539 | 0.359 | 0.890 | 0.429 | |
| change | +0.022 | +0.029 | +0.011 | +0.045 | |
| AMOS-MM | SFT | 0.421 | 0.249 | 0.881 | 0.339 |
| SFT & DPO | 0.459 | 0.876 | 0.881 | 0.400 | |
| change | +0.038 | +0.627* | 0.000 | +0.061 |
* As printed; almost certainly a misprint. Source: the paper’s Table 1.
GREEN gains 12 to 23% of its starting value and ROUGE-1 4 to 9% derived. This is not the pattern of the design-space paper. There, canned normal sentences added 0.078 GREEN while RaTEScore moved 0.002 and BLEU fell 0.057: a gain only the judge saw. Here the model’s own text changes, and no score the paper prints goes down.
All four scores compare the generated text with one reference report. None checks findings on its own terms, as a label classifier would on CT-RATE’s 18 abnormalities, and no radiologist read the output.
And the released evaluation scripts sample µ²LLM’s reports at temperature 1.0, like the eight drafts. A model taught to prefer its best draw over its worst should draw fewer bad reports. Whether its best reports improved, only a greedy-decoded row before and after DPO would show.
05What the paper leaves out
- Test sets and uncertainty: no split, no count, no decoding setting, no interval or repeat run, and one example report without its reference. The committed AMOS-MM script for µ²LLM stops after 100 validation cases. The lead over CT-CHAT is still called significant, though a reviewer asked for a test or another word.
- The recipe: what trains, for how long, how many preference pairs, and which β between 0.1 and 0.5 (0.1 in the scripts). The released training transform also rotates and flips volumes while the reports keep their lefts and rights our observation.
- The release: the tokenizer, both training stages and the pair-building scripts are public. The weights are three Qwen3-based successors, not the paper’s Llama-3.2-1B checkpoints, and the data are CT-RATE derivatives, not the preference pairs.
06Takeaways for building a CT report generator
- DPO is a cheap way to put a judge in the loop. The paper reports 0.045 to 0.065 more GREEN with no reward model; the released recipe is eight samples per training example and one scoring pass.
- If you train on a metric, add a ruler it doesn’t share. ROUGE-1 and METEOR followed GREEN, but all three compare text with the same reference; label F1 on CT-RATE’s 18 findings compares labels.
- Decode greedily before and after. When drafts and test reports are both sampled at temperature 1.0, fewer bad draws and better reports look the same.
- Train your baselines on your data. The same stack scores 0.009 GREEN here, run as published going by the scripts, and 0.365 in a study that trained it on AMOS-MM.
ASources
- Li, S., Qin, P., Wu, H., Nie, D., Thirunavukarasu, A. J., Yu, J., Zhang, L. µ² Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation. MICCAI 2025, LNCS 15964, pp. 3–12. The page holds the reviews and the rebuttal. PDF · doi · arXiv:2507.00316: v1 still has the typos the reviewers flagged; v2 has the final text plus a section and an appendix on synthetic question–answer data. The table values are the same in all three.
- Code: Siyou-Li/u2Tokenizer (MIT), read on 5 Oct 2026 at commit
71ed693; the per-dataset training and evaluation scripts are in the history, at commit7d1ecb7(2 Jul 2025). Weights and data: AlpachinoNLP on Hugging Face. - Ostmeier, S., et al. GREEN: arXiv:2405.03595, Findings of EMNLP 2024; checkpoints GREEN-RadLlama2-7b and GREEN-RadPhi2.
- Rafailov, R., et al. Direct Preference Optimization: arXiv:2305.18290, NeurIPS 2023. Gao, L., et al. LinVT: arXiv:2412.05185 (2024). Bai, F., et al. M3D: arXiv:2404.00578 (2024).
- Datasets: AMOS-MM challenge (Codabench); Hamamci, I. E., et al. CT-RATE and CT-CHAT: arXiv:2403.17834; Bassi, P. R. A. S., et al. RadGPT and AbdomenAtlas 3.0: arXiv:2501.04678.
- Our posts on the 3D MLLM design space, Merlin and NV-Reason-CT.