* From the released code and test outputs; the last assumes a per-scan average (section 02).
- Diagnose first, then write. A classifier fills one sentence per CheXbert label, and LLaMA-2 reads them after 32 visual tokens. In the released code, the training prompt is filled from the reference report’s own labels.
- The printed gain is small and uneven. It is +0.027 F1 over the same stack without the prompt. The prompt alone adds 0.002, and the paper’s table and chart rank its two modules in opposite order.
- In the released outputs the diagnosis over-calls, and most of it doesn’t reach the report. The prompt marks 5.5 labels present per scan, the references 1.6. Pneumothorax is asserted for 148 of 362 scans and written in no report.
- The rulers are loose. Read as a per-scan average, the clinical score puts one fixed answer above every row of the table, and a typical fixed report reaches BLEU-4 23.6. Which run produced the tables isn’t stated.
01Why this paper matters for report generation
An LLM can write a fluent report around a finding it missed. Dia-LLaMA decides first and writes second: a classifier reads the scan, and its verdict goes into the prompt as plain sentences. The same group’s PromptMRG did this for chest X-rays; this paper moves it to chest CT and LLaMA-2. Code, a checkpoint, test outputs and reviews are public; this post uses all four.
→ 1,024 × 768
- The attention doesn’t look at the scan. The weights of Eq. 5 are a learned table, one row of 1,024 per label, the same for every patient. It can favour a region, not follow a lesion our observation.
- The memory bank is 28 vectors in the released code. With two prototypes per label, the loss of Eq. 6 is that of a linear classifier with no bias our derivation.
- 14 labels, not eight. The paper puts eight diseases in the prompt; every public version of the code writes and scores all 14 CheXbert labels, “no finding” included.
- The training prompt is the answer key. In training the sentences are filled from the reference report’s own CheXbert labels, as in PromptMRG’s code; at test time, by the classifier. The paper doesn’t say so. Condensed from the released dataset and model code:
def diagnosis_prompt(present): # one 0/1 per CheXbert label
return "".join(
f'The "{name}" is {"positive" if on else "negative"}. '
for name, on in zip(LABELS, present))
# training: the reference report's own labels
prompt = diagnosis_prompt(chexbert(reference) == PRESENT)
# test: the scan's features against two prototypes per label
prompt = diagnosis_prompt(sim_present > sim_absent)
text = IMAGE_TOKENS + prompt + INSTRUCTION # then the report
02The data and the ruler
CTRG-Chest-548K has 1,804 chest CT scans with reports, 548,696 JPEG slices in all, so the model sees display intensities, not Hounsfield units. The reports were written in Chinese and come with an English version, according to a later paper. This one doesn’t say. It trains on 80% and tests on 20%, drawn at random per its arXiv version; the public split has 362 test scans.
The clinical score is CheXbert, a chest X-ray labeler, run on both reports. The label file in the release shows what it finds here: “lung lesion” in 57.2% of reports, “enlarged cardiomediastinum” in 35.0%, pneumonia and pneumothorax never, and no present label in 14.1%. The paper asserts that the labeler stays valid on CT reports. A reviewer noted this wasn’t validated; the reply was that it worked “reasonably well”, with a CT labeler left to future work.
The paper follows PromptMRG’s setting, whose scorer, extended in the release, averages per scan over the 14 labels. The paper’s Table 1 fits: its F1 (0.372) is below both precision and recall, which a micro-average can’t be, and it isn’t the mean over labels, which the paper’s chart puts near 0.33 for its eight.
Under the per-scan formula, the 54 test scans whose reference has no present label score zero whatever the report says (ceiling 0.851). And a report read as “lung lesion” alone, for every scan, would score 0.452 F1, with precision 0.591 and recall 0.401 derived: above every row of Table 1. (Micro-averaged: 0.455. As a mean over labels: 0.053.)
03The results as printed
Dia-LLaMA leads the strongest of six baselines, RadFM, by 0.018 precision, 0.026 recall and 0.027 F1 and by 4.9 BLEU-4 points. On ROUGE-L it is fourth of seven. RadFM’s row is also the ablation’s first: the same encoder, perceiver and LoRA-tuned LLaMA-2, without the prompt. So the ablation is the test of the idea.
Show as table
| Setting | F1 | Fig. 2 | Prec. | Rec. | BL-1 | BL-4 | MTR | RG-L |
|---|---|---|---|---|---|---|---|---|
| No prompt (baseline) | 0.345 | 0.199 | 0.403 | 0.361 | 46.70 | 24.70 | 24.01 | 38.98 |
| + prompt, plain classifier | 0.347 | 0.224 | 0.415 | 0.336 | 45.74 | 27.05 | 24.80 | 42.29 |
| + prompt, attention (DAA) | 0.358 | 0.213 | 0.424 | 0.347 | 44.22 | 26.38 | 24.34 | 42.68 |
| + prompt, prototypes (DPM) | 0.339 | 0.256 | 0.437 | 0.313 | 44.06 | 27.10 | 24.46 | 44.5 |
| + prompt, both: Dia-LLaMA | 0.372 | 0.325 | 0.421 | 0.387 | 51.16 | 29.64 | 26.28 | 42.15 |
- The prompt alone adds 0.002 F1, and with a plain classifier behind it recall falls from 0.361 to 0.336.
- The two rulers disagree about the modules. By the table, attention helps (0.358) and prototypes alone hurt (0.339, below the baseline). By the chart’s mean, prototypes help (0.256) and attention hardly does (0.213). The four rarest of the chart’s eight labels have 13 to 25 positive scans each in the released test set.
No interval or second run is reported.
04Inside the released test outputs
The paper scores the report, never the classifier. The repository lets us: results/dia-llama.csv holds, for 362 test scans, the prompt the model wrote, its report, the reference and its labels. Re-scored the release’s way, the reports reach BLEU-4 30.0; an earlier version of the file, with the same prompts and different reports, gives 29.5. The paper’s 29.64 lies between, so this is evidently the paper’s system, though perhaps not the run behind its tables.
Show as table
| CheXbert label | Present in the 1,804 reports | Test references | Prompt says positive | Both |
|---|---|---|---|---|
| Pleural effusion | 221 (12.3%) | 48 | 331 | 46 |
| Lung lesion | 1,031 (57.2%) | 214 | 307 | 182 |
| Edema | 10 (0.6%) | 3 | 263 | 3 |
| Fracture | 18 (1.0%) | 6 | 240 | 5 |
| Pleural other | 443 (24.6%) | 94 | 225 | 63 |
| Consolidation | 76 (4.2%) | 17 | 166 | 13 |
| Pneumothorax | 0 (0.0%) | 0 | 148 | 0 |
| Enlarged cardiomediastinum | 631 (35.0%) | 122 | 134 | 52 |
| Lung opacity | 28 (1.6%) | 6 | 64 | 1 |
| Atelectasis | 12 (0.7%) | 5 | 47 | 1 |
| Pneumonia | 0 (0.0%) | 0 | 47 | 0 |
| Cardiomegaly | 144 (8.0%) | 25 | 17 | 3 |
| Support devices | 112 (6.2%) | 25 | 10 | 1 |
| No finding | 81 (4.5%) | 13 | 0 | 0 |
| All 14 labels | 2,807 | 578 | 1,999 | 370 |
The prompts assert 1,999 labels, the references hold 578, and 370 coincide: a precision of 0.19 and a recall of 0.64. Pneumothorax is asserted for 148 scans, although no report in the dataset carries the label. Scored as if it were the report, the prompt would get a per-scan F1 of 0.269, below the no-prompt baseline.
Most of it doesn’t reach the report, by a word count. Seven labels have an unambiguous word: effusion, edema, fracture, consolidation, pneumothorax, pneumonia and atelectasis. They carry 1,242 of the 1,999 assertions; the report states the finding, with no negation before the word, in 27 or 28 of those cases, depending on the rule. The word pneumothorax is in no report.
Show the seven labels as a table
| Label | Prompt says positive | … and the report affirms it | … and the reference has it | Reference label present |
|---|---|---|---|---|
| Pleural effusion | 331 | 22 or 23 | 7 | 48 |
| Edema | 263 | 1 | 0 | 3 |
| Fracture | 240 | 2 | 0 | 6 |
| Consolidation | 166 | 2 | 0 | 17 |
| Pneumothorax | 148 | 0 | 0 | 0 |
| Pneumonia | 47 | 0 | 0 | 0 |
| Atelectasis | 47 | 0 | 0 | 5 |
| Seven labels | 1,242 | 27 or 28 | 7 | 79 |
Table 1 — What the reports do with the prompt’s assertions. “Affirms”: the report has the word with no negation word earlier in the sentence or, as a second rule, among the six words before it. On the references, the two rules find 71 and 74 of the 79 labelled reports, and fire on 9 unlabelled ones. Pneumothorax and pneumonia are never present in the label file, so the language model saw no positive prompt for either in training. Our count on the released outputs.
In training, each label’s feature is the attention-weighted sum over the 1,024 patches. At test time the released code sums over the 14 labels instead, so label i is judged on the feature at patch i, against prototypes fitted to pooled features (lines 130–134, in every public version of the file; the comment there gives the intended shape). It may explain the over-calling; we did not run the model. Whether the paper’s numbers came from this path isn’t stated.
So here the score evidently belongs to the language model, whose reports are more cautious than their prompts. What produces the 0.027 is open: perhaps the auxiliary loss, which also trains the encoder in the released code. No row of the paper trains the modules and leaves the prompt out.
05Fixed answers and moving baselines
The paper notes that CT reports follow a rigid template. One of the test set’s own reference reports, as the answer for every test scan (its own too) and scored the release’s way, gets BLEU-4 23.6 at the median of the 362 choices derived. The six baselines score 21.9 to 24.7; their tokenisation isn’t stated.
Show as table
| Row | Prec. | Recall | F1 | BLEU-1 | BLEU-4 | METEOR | ROUGE-L |
|---|---|---|---|---|---|---|---|
| R2Gen | 0.207 | 0.121 | 0.144 | 34.11 | 23.39 | 21.40 | 47.75 |
| R2GenCMN | 0.158 | 0.100 | 0.114 | 35.88 | 23.37 | 21.43 | 45.94 |
| M2KT | 0.220 | 0.119 | 0.145 | 46.09 | 21.93 | 25.20 | 36.47 |
| PromptMRG | 0.290 | 0.330 | 0.290 | 47.73 | 23.02 | 22.87 | 37.35 |
| SL-DG * | – | – | – | – | 23.70 | 21.90 | 43.80 |
| RadFM | 0.403 | 0.361 | 0.345 | 46.70 | 24.70 | 24.01 | 38.98 |
| Dia-LLaMA | 0.421 | 0.387 | 0.372 | 51.16 | 29.64 | 26.28 | 42.15 |
| Fixed: “lung lesion” label | 0.591 | 0.401 | 0.452 | – | – | – | – |
| Fixed: one reference report | – | – | – | 43.02 | 23.57 | – | – |
| Released prompts as reports | 0.198 | 0.555 | 0.269 | – | – | – | – |
| Released reports, re-scored | – | – | – | 51.23 | 29.98 | – | 42.77 |
What rises above the template isn’t checked by these scores. 82 generated reports cite a slice number in the form IM141; where the reference cites one too (17 scans), none matches. In the paper’s one example, the model’s report shrinks the reference’s 15 × 10 mm lobulated nodule to 6 × 4 mm in another segment, which a “lung lesion” label can’t see.
Who ran the baselines? Only SL-DG’s row is marked as copied. It comes from the dataset paper, on another split: the arXiv caption said so, the MICCAI caption doesn’t. The others are evidently the authors’ runs, the four chest X-ray models on 30 slices per scan, per the arXiv version. And in a later paper from the group (Reg2RG), RadFM, described the same way on this dataset, scores BLEU-4 30.89, above Dia-LLaMA’s 29.64 here, with a METEOR of 49 against 24: the scorer evidently changed too.
06What the paper leaves out
- The classifier: no accuracy for the verdict the method rests on.
- Readers and spread: no radiologist read a report, and there is no repeat run or interval. Decoding isn’t described; the released file’s versions suggest it is sampled.
- The data: one dataset, as a reviewer objected; no patient count, and nothing on a patient-level split, the reports’ language or the JPEG source.
- The recipe: 2,000 steps in the arXiv version, 20 epochs in the final text, 25.4 passes over 1,262 scans in the released training log; nothing on what is frozen.
07Takeaways for building a CT report generator
- Score the classifier before crediting the prompt. The released verdicts have a precision of 0.19; the paper reports the report’s F1 only.
- Train the writer on prompts it will meet. In the release, training prompts come from the reference labels (about 1.6 present per scan), test prompts from a classifier (5.5). Mix in predicted labels.
- Put fixed answers in the table. One label scores 0.452 per-scan F1 and one typical report BLEU-4 23.6.
- Use labels your reports contain. Two of the 14 never occur in 1,804 reports, and 14.1% of reports have none.
- Freeze the protocol before comparing. RadFM scores BLEU-4 24.70 and 30.89 in two papers from one group, and a second decode moves Dia-LLaMA’s BLEU-4 by 0.5.
ASources
- Chen, Z., Luo, L., Bie, Y., Chen, H. Dia-LLaMA: Towards Large Language Model-driven CT Report Generation. MICCAI 2025, LNCS 15966, pp. 141–151. The page holds the three reviews and the authors’ reply. PDF · doi · arXiv:2403.16386 v1 (March 2024: the same three tables; a different account of the split, the baselines and the training length).
- Code: zhi-xuan-chen/Dia-LLaMA at commit 6c278fea, read on 6 Oct 2026, with its test outputs and label file. Checkpoint and training log: Trusure/Dia-LLaMA.
- Dataset: CTRG; Tang, Y., et al. Work like a doctor. Expert Systems with Applications 237, 121442 (2024), the source of the SL-DG row.
- Jin, H., et al. PromptMRG: arXiv:2308.12604 (AAAI 2024) and its scorer.
- Chen, Z., et al. Reg2RG: arXiv:2411.15539 (IEEE TMI 2025), Table I and Section IV.
- Li, D., Liang, J., Li, W. ALTER: arXiv:2608.05615 (2026), Appendix A: the dataset’s language and split sizes.
- Smit, A., et al. CheXbert: arXiv:2004.09167 (EMNLP 2020). Wu, C., et al. RadFM: arXiv:2308.02463.
- Our posts on the 3D MLLM design space and µ²LLM.