Jean-Benoit Delbrouck / Research blog

Research blog · CT report generation

Dia-LLaMA, read for report generation

A MICCAI 2025 paper asks a classifier for a diagnosis, writes it into the prompt and lets a LoRA-tuned LLaMA-2 report a chest CT. This post reads the method, its 0.027 F1 gain, and the released test outputs, where the report leaves out most of the diagnosis.

Paper Zhixuan Chen, Luyang Luo, Yequan Bie, Hao Chen. Dia-LLaMA: Towards Large Language Model-driven CT Report Generation. MICCAI 2025, LNCS 15966, pp. 141–151 By Jean-Benoit Delbrouck Published 6 Oct 2026 Reading time ~8 min Tags MICCAI 2025 · Main
32
visual tokens per scan, plus one sentence per label from a classifier*
+0.027
F1 over the same stack without the prompt; 362 test scans*, no interval
5.5 vs 1.6
labels marked present per scan: the released prompts against the references*
0.452
F1 of one fixed answer, “lung lesion” for every scan; the table’s best is 0.372*

* From the released code and test outputs; the last assumes a per-scan average (section 02).

TL;DR
  • Diagnose first, then write. A classifier fills one sentence per CheXbert label, and LLaMA-2 reads them after 32 visual tokens. In the released code, the training prompt is filled from the reference report’s own labels.
  • The printed gain is small and uneven. It is +0.027 F1 over the same stack without the prompt. The prompt alone adds 0.002, and the paper’s table and chart rank its two modules in opposite order.
  • In the released outputs the diagnosis over-calls, and most of it doesn’t reach the report. The prompt marks 5.5 labels present per scan, the references 1.6. Pneumothorax is asserted for 148 of 362 scans and written in no report.
  • The rulers are loose. Read as a per-scan average, the clinical score puts one fixed answer above every row of the table, and a typical fixed report reaches BLEU-4 23.6. Which run produced the tables isn’t stated.

01Why this paper matters for report generation

An LLM can write a fluent report around a finding it missed. Dia-LLaMA decides first and writes second: a classifier reads the scan, and its verdict goes into the prompt as plain sentences. The same group’s PromptMRG did this for chest X-rays; this paper moves it to chest CT and LLaMA-2. Code, a checkpoint, test outputs and reviews are public; this post uses all four.

Figure 1 — A scan becomes 32 visual tokens and 14 sentences. Input size, LoRA settings and modules are the paper’s; the token count, the 14 labels and what trains come from the released code (encoder and diagnosis, language model). Sources: Sections 2 and 3.2.
  • The attention doesn’t look at the scan. The weights of Eq. 5 are a learned table, one row of 1,024 per label, the same for every patient. It can favour a region, not follow a lesion our observation.
  • The memory bank is 28 vectors in the released code. With two prototypes per label, the loss of Eq. 6 is that of a linear classifier with no bias our derivation.
  • 14 labels, not eight. The paper puts eight diseases in the prompt; every public version of the code writes and scores all 14 CheXbert labels, “no finding” included.
  • The training prompt is the answer key. In training the sentences are filled from the reference report’s own CheXbert labels, as in PromptMRG’s code; at test time, by the classifier. The paper doesn’t say so. Condensed from the released dataset and model code:
def diagnosis_prompt(present):      # one 0/1 per CheXbert label
    return "".join(
        f'The "{name}" is {"positive" if on else "negative"}. '
        for name, on in zip(LABELS, present))

# training: the reference report's own labels
prompt = diagnosis_prompt(chexbert(reference) == PRESENT)
# test: the scan's features against two prototypes per label
prompt = diagnosis_prompt(sim_present > sim_absent)

text = IMAGE_TOKENS + prompt + INSTRUCTION   # then the report

02The data and the ruler

CTRG-Chest-548K has 1,804 chest CT scans with reports, 548,696 JPEG slices in all, so the model sees display intensities, not Hounsfield units. The reports were written in Chinese and come with an English version, according to a later paper. This one doesn’t say. It trains on 80% and tests on 20%, drawn at random per its arXiv version; the public split has 362 test scans.

The clinical score is CheXbert, a chest X-ray labeler, run on both reports. The label file in the release shows what it finds here: “lung lesion” in 57.2% of reports, “enlarged cardiomediastinum” in 35.0%, pneumonia and pneumothorax never, and no present label in 14.1%. The paper asserts that the labeler stays valid on CT reports. A reviewer noted this wasn’t validated; the reply was that it worked “reasonably well”, with a CT labeler left to future work.

Which average is the F1? likely, not confirmed

The paper follows PromptMRG’s setting, whose scorer, extended in the release, averages per scan over the 14 labels. The paper’s Table 1 fits: its F1 (0.372) is below both precision and recall, which a micro-average can’t be, and it isn’t the mean over labels, which the paper’s chart puts near 0.33 for its eight.

Under the per-scan formula, the 54 test scans whose reference has no present label score zero whatever the report says (ceiling 0.851). And a report read as “lung lesion” alone, for every scan, would score 0.452 F1, with precision 0.591 and recall 0.401 derived: above every row of Table 1. (Micro-averaged: 0.455. As a mean over labels: 0.053.)

03The results as printed

Dia-LLaMA leads the strongest of six baselines, RadFM, by 0.018 precision, 0.026 recall and 0.027 F1 and by 4.9 BLEU-4 points. On ROUGE-L it is fourth of seven. RadFM’s row is also the ablation’s first: the same encoder, perceiver and LoRA-tuned LLaMA-2, without the prompt. So the ablation is the test of the idea.

Table 2: F1change from the baseline (0.345)changescore0+0.05+0.10+0.0020.347+0.0130.358−0.0060.339+0.0270.372Fig. 2: mean F1 of eight labelschange from the baseline (0.199), measuredchangescore0+0.05+0.10+0.0250.224+0.0140.213+0.0570.256+0.1260.325+ prompt, plain classifier+ prompt, attention (DAA)+ prompt, prototypes (DPM)+ prompt, both: Dia-LLaMA
Figure 2 — The table and the chart rank the partial settings in opposite order. Changes from the no-prompt baseline. Left: F1 in the paper’s Table 2. Right: the mean of the eight per-label F1s of its Fig. 2, measured off the image (two measurements within 0.05 points). Sources: Table 2 and Fig. 2.
Show as table
SettingF1Fig. 2Prec.Rec.BL-1BL-4MTRRG-L
No prompt (baseline)0.3450.1990.4030.36146.7024.7024.0138.98
+ prompt, plain classifier0.3470.2240.4150.33645.7427.0524.8042.29
+ prompt, attention (DAA)0.3580.2130.4240.34744.2226.3824.3442.68
+ prompt, prototypes (DPM)0.3390.2560.4370.31344.0627.1024.4644.5
+ prompt, both: Dia-LLaMA0.3720.3250.4210.38751.1629.6426.2842.15
  1. The prompt alone adds 0.002 F1, and with a plain classifier behind it recall falls from 0.361 to 0.336.
  2. The two rulers disagree about the modules. By the table, attention helps (0.358) and prototypes alone hurt (0.339, below the baseline). By the chart’s mean, prototypes help (0.256) and attention hardly does (0.213). The four rarest of the chart’s eight labels have 13 to 25 positive scans each in the released test set.

No interval or second run is reported.

04Inside the released test outputs

The paper scores the report, never the classifier. The repository lets us: results/dia-llama.csv holds, for 362 test scans, the prompt the model wrote, its report, the reference and its labels. Re-scored the release’s way, the reports reach BLEU-4 30.0; an earlier version of the file, with the same prompts and different reports, gives 29.5. The paper’s 29.64 lies between, so this is evidently the paper’s system, though perhaps not the run behind its tables.

The model’s prompt says “positive”The reference label is presentpromptreference0100200300362Pleural effusion33148Lung lesion307214Edema2633Fracture2406Pleural other22594Consolidation16617Pneumothorax1480Enlarged cardiomediastinum134122Lung opacity646Atelectasis475Pneumonia470Cardiomegaly1725Support devices1025No finding013
Figure 3 — The prompt marks 5.5 labels present per scan; the references, 1.6. Test scans out of 362, sorted by the prompt’s count. Our count on the released file, whose reference labels are its own. Source: the repository’s test outputs.
Show as table
CheXbert labelPresent in the 1,804 reportsTest referencesPrompt says positiveBoth
Pleural effusion221 (12.3%)4833146
Lung lesion1,031 (57.2%)214307182
Edema10 (0.6%)32633
Fracture18 (1.0%)62405
Pleural other443 (24.6%)9422563
Consolidation76 (4.2%)1716613
Pneumothorax0 (0.0%)01480
Enlarged cardiomediastinum631 (35.0%)12213452
Lung opacity28 (1.6%)6641
Atelectasis12 (0.7%)5471
Pneumonia0 (0.0%)0470
Cardiomegaly144 (8.0%)25173
Support devices112 (6.2%)25101
No finding81 (4.5%)1300
All 14 labels2,8075781,999370

The prompts assert 1,999 labels, the references hold 578, and 370 coincide: a precision of 0.19 and a recall of 0.64. Pneumothorax is asserted for 148 scans, although no report in the dataset carries the label. Scored as if it were the report, the prompt would get a per-scan F1 of 0.269, below the no-prompt baseline.

Most of it doesn’t reach the report, by a word count. Seven labels have an unambiguous word: effusion, edema, fracture, consolidation, pneumothorax, pneumonia and atelectasis. They carry 1,242 of the 1,999 assertions; the report states the finding, with no negation before the word, in 27 or 28 of those cases, depending on the rule. The word pneumothorax is in no report.

Show the seven labels as a table
LabelPrompt says positive… and the report affirms it… and the reference has itReference label present
Pleural effusion33122 or 23748
Edema263103
Fracture240206
Consolidation1662017
Pneumothorax148000
Pneumonia47000
Atelectasis47005
Seven labels1,24227 or 28779

Table 1 — What the reports do with the prompt’s assertions. “Affirms”: the report has the word with no negation word earlier in the sentence or, as a second rule, among the six words before it. On the references, the two rules find 71 and 74 of the 79 labelled reports, and fire on 9 unlabelled ones. Pneumothorax and pneumonia are never present in the label file, so the language model saw no positive prompt for either in training. Our count on the released outputs.

Worth checking: the test path our observation

In training, each label’s feature is the attention-weighted sum over the 1,024 patches. At test time the released code sums over the 14 labels instead, so label i is judged on the feature at patch i, against prototypes fitted to pooled features (lines 130–134, in every public version of the file; the comment there gives the intended shape). It may explain the over-calling; we did not run the model. Whether the paper’s numbers came from this path isn’t stated.

So here the score evidently belongs to the language model, whose reports are more cautious than their prompts. What produces the 0.027 is open: perhaps the auxiliary loss, which also trains the encoder in the released code. No row of the paper trains the modules and leaves the prompt out.

05Fixed answers and moving baselines

The paper notes that CT reports follow a rigid template. One of the test set’s own reference reports, as the answer for every test scan (its own too) and scored the release’s way, gets BLEU-4 23.6 at the median of the 362 choices derived. The six baselines score 21.9 to 24.7; their tokenisation isn’t stated.

Table 1Dia-LLaMAOne fixed answer for every scan (our arithmetic)CheXbert F1scale 0–100.10.20.30.40.50.1440.1140.1450.290not reported–0.3450.3720.452the label “lung lesion”BLEU-4scale 0–100010203023.3923.3721.9323.0223.7024.7029.6423.57a typical reference reportR2GenR2GenCMNM2KTPromptMRGSL-DG *RadFMDia-LLaMAOne fixed answer
Figure 4 — Fixed answers top the F1 column, read per scan, and sit mid-table on BLEU-4. Filled bars: Table 1. Outlined: our arithmetic on the released references, “lung lesion” for every scan (per-scan F1) and one test reference report for every scan (median of 362). * Copied from the dataset paper. Sources: Table 1 and the repository’s test outputs.
Show as table
RowPrec.RecallF1BLEU-1BLEU-4METEORROUGE-L
R2Gen0.2070.1210.14434.1123.3921.4047.75
R2GenCMN0.1580.1000.11435.8823.3721.4345.94
M2KT0.2200.1190.14546.0921.9325.2036.47
PromptMRG0.2900.3300.29047.7323.0222.8737.35
SL-DG *––––23.7021.9043.80
RadFM0.4030.3610.34546.7024.7024.0138.98
Dia-LLaMA0.4210.3870.37251.1629.6426.2842.15
Fixed: “lung lesion” label0.5910.4010.452––––
Fixed: one reference report–––43.0223.57––
Released prompts as reports0.1980.5550.269––––
Released reports, re-scored–––51.2329.98–42.77

What rises above the template isn’t checked by these scores. 82 generated reports cite a slice number in the form IM141; where the reference cites one too (17 scans), none matches. In the paper’s one example, the model’s report shrinks the reference’s 15 × 10 mm lobulated nodule to 6 × 4 mm in another segment, which a “lung lesion” label can’t see.

Who ran the baselines? Only SL-DG’s row is marked as copied. It comes from the dataset paper, on another split: the arXiv caption said so, the MICCAI caption doesn’t. The others are evidently the authors’ runs, the four chest X-ray models on 30 slices per scan, per the arXiv version. And in a later paper from the group (Reg2RG), RadFM, described the same way on this dataset, scores BLEU-4 30.89, above Dia-LLaMA’s 29.64 here, with a METEOR of 49 against 24: the scorer evidently changed too.

06What the paper leaves out

  • The classifier: no accuracy for the verdict the method rests on.
  • Readers and spread: no radiologist read a report, and there is no repeat run or interval. Decoding isn’t described; the released file’s versions suggest it is sampled.
  • The data: one dataset, as a reviewer objected; no patient count, and nothing on a patient-level split, the reports’ language or the JPEG source.
  • The recipe: 2,000 steps in the arXiv version, 20 epochs in the final text, 25.4 passes over 1,262 scans in the released training log; nothing on what is frozen.

07Takeaways for building a CT report generator

  1. Score the classifier before crediting the prompt. The released verdicts have a precision of 0.19; the paper reports the report’s F1 only.
  2. Train the writer on prompts it will meet. In the release, training prompts come from the reference labels (about 1.6 present per scan), test prompts from a classifier (5.5). Mix in predicted labels.
  3. Put fixed answers in the table. One label scores 0.452 per-scan F1 and one typical report BLEU-4 23.6.
  4. Use labels your reports contain. Two of the 14 never occur in 1,804 reports, and 14.1% of reports have none.
  5. Freeze the protocol before comparing. RadFM scores BLEU-4 24.70 and 30.89 in two papers from one group, and a second decode moves Dia-LLaMA’s BLEU-4 by 0.5.

ASources

  1. Chen, Z., Luo, L., Bie, Y., Chen, H. Dia-LLaMA: Towards Large Language Model-driven CT Report Generation. MICCAI 2025, LNCS 15966, pp. 141–151. The page holds the three reviews and the authors’ reply. PDF · doi · arXiv:2403.16386 v1 (March 2024: the same three tables; a different account of the split, the baselines and the training length).
  2. Code: zhi-xuan-chen/Dia-LLaMA at commit 6c278fea, read on 6 Oct 2026, with its test outputs and label file. Checkpoint and training log: Trusure/Dia-LLaMA.
  3. Dataset: CTRG; Tang, Y., et al. Work like a doctor. Expert Systems with Applications 237, 121442 (2024), the source of the SL-DG row.
  4. Jin, H., et al. PromptMRG: arXiv:2308.12604 (AAAI 2024) and its scorer.
  5. Chen, Z., et al. Reg2RG: arXiv:2411.15539 (IEEE TMI 2025), Table I and Section IV.
  6. Li, D., Liang, J., Li, W. ALTER: arXiv:2608.05615 (2026), Appendix A: the dataset’s language and split sizes.
  7. Smit, A., et al. CheXbert: arXiv:2004.09167 (EMNLP 2020). Wu, C., et al. RadFM: arXiv:2308.02463.
  8. Our posts on the 3D MLLM design space and µ²LLM.