* Our arithmetic: 5,159 / 120.
- One sentence per lesion, with the lesion given. Masks come from an nnU-Net and are cleaned by radiologists. The model writes what the CT shows and where; no caption printed in the paper mentions tracer uptake.
- The location reaches the decoder as words and as a CT window. A CLIP-style model names the region from a wider crop. The two modules add 1.2 BLEU-4, 1.9 points of location accuracy and 1.8 of CT-finding accuracy to an X-Transformer baseline.
- BLEU-4 is a poor ruler for these sentences. Nine captioners score between 70.1 and 76.9. By our arithmetic on the paper’s own examples, a caption with the wrong neck node level scores 82.7 and one with a harmless synonym 78.3; with punctuation counted, they tie.
- What we can’t tell. Data and code are private, and no result is split by scanner maker. No reader rated the captions, and the significance test doesn’t say whether it counts lesions or patients.
01Why this paper matters for report generation
This is the blog’s first PET/CT post. For readers who work on CT or chest X-ray reports, three things change:
- Two volumes per exam. PET shows where an injected tracer accumulates, here FDG or a PSMA tracer; CT shows anatomy and density. A report sentence usually joins the two: where the lesion is, what the CT shows, how strong the uptake is.
- The whole body. A scan typically runs from the skull base to the thighs. Neither the location model nor the captioner in this paper sees it whole.
- A sentence, not a report. The model takes one lesion mask at a time and writes one sentence about it. The authors present it as the first lesion captioner for whole-body PET/CT.
So the experiment asks a narrow question: given a confirmed lesion, can a model name its location and its CT appearance? Detection is outside the score, because the masks are given (Figure 1). Unlike the CT systems read so far in this series, the model has no LLM. A 3D CNN encoder feeds the X-Transformer, a captioning model from 2020; the rebuttal counts about 400M parameters.
≈ 19 × 19 × 26 cm
→ top 5 names + confidence
= 9.6 cm cube
120 patients
02The data and the captions
The dataset holds 1,867 whole-body scans of patients with lymphoma or with nasopharyngeal, lung, prostate or liver cancer. The paper doesn’t say where they were acquired. Its authors are at ShanghaiTech University, United Imaging Intelligence and two nuclear medicine departments, in Beijing and Guangzhou, and its acknowledgments thank the Guangzhou one for “validation data”.
The split is by patient ID, which a reviewer had asked about. How patients were assigned isn’t described. Test and validation patients carry about 43 and 45 lesions each, against 28 in training derived, which suggests the draw was not random our observation.
A reference caption reads: “A lymph node is noted in the left neck (level Va) region.” It is probably not dictated text: for the training set, the paper says an LLM turned the original reports into structured ones and radiologists refined the result. The paper names neither the LLM nor the language of the reports; its examples are printed in English.
The paper prints 24 captions, references and model outputs together. They describe CT findings and locations; none mentions uptake or an SUV. The two accuracy measures cover the same two fields, and no run removes the PET input.
So the score is for a CT description of a lesion that PET presumably helped to find.
03How the location reaches the decoder
Naming a location takes context that a close-up lacks: which rib, which vertebra. So the paper crops each lesion twice. The captioner reads 64³ voxels at 1.5 mm, a 9.6 cm cube. The location model reads 96 × 96 × 128 voxels at 2 mm, about 19 × 19 × 26 cm, over ten times the volume derived.
The paper calls stage 1 CLIP-based lesion localisation, but it doesn’t find lesions. It matches the wide crop against the names of 130 anatomical locations, such as right lung, hilum or mediastinum, and returns the five best, each with a confidence. With a closed list, this is in effect a 130-way classifier trained with a contrastive loss.
Two rules pass its answer to the captioner, both with a threshold of 0.95. Our reconstruction from the paper’s Sections 2.2–2.3 and Figure 2, not the authors’ code; the windowing formula, the HU unit and the encoder’s inputs are our assumptions:
SOFT_TISSUE = (40, 350) # CT window: (level, width) in HU
# also named in the paper: lung (-600, 1500), bone (800, 2600)
# window_of: location name -> window; the table isn't published
def guide(names, conf, window_of, thr=0.95):
"""names, conf: stage-1 answers, best first."""
if conf[0] > thr: # confident: trust the top answer
return names[:1], window_of[names[0]]
return names[:3], SOFT_TISSUE # unsure: 3 names, default window
def caption(ct, pet, mask, names, conf, window_of):
prompt, (level, width) = guide(names, conf, window_of)
ct = ((ct - level) / width + 0.5).clip(0, 1) # window to [0, 1]
feats = encoder(ct, pet, mask) # close-up: 64³ at 1.5 mm
words = embed(pad(tokenize(prompt))) # location names as tokens
return decoder(concat(feats, words)) # one sentence
A reviewer asked whether the window was a lookup table. The rebuttal doesn’t say no: it calls the pairing of region and window clinical prior knowledge, and puts the contribution in choosing it automatically. It also explains 0.95: above it, the top answer was right more than 90% of the time on validation data.
The paper gives no accuracy for stage 1 as a whole. A companion paper by the same team reports 84.13% on an internal test set and 80.42% on an external one, for a finer localiser with 387 locations. The label set and the test data differ, so read it as an order of magnitude.
04Results: nine captioners, seven points
The main table sets the full model against eight captioning models from the natural-image literature, the newest from 2020. The paper doesn’t say how they were adapted to 3D PET/CT or trained. The X-Transformer row is the authors’ own baseline: it matches the first row of the ablation to the decimal.
Show as table
| Model | BLEU-3 | BLEU-4 | METEOR | ROUGE-L | CIDEr | SPICE |
|---|---|---|---|---|---|---|
| This paper’s model | 80.1 | 76.9 | 85.8 | 88.3 | 7.92 | 88.7 |
| X-Transformer | 78.9 | 75.7 | 84.9 | 87.5 | 7.80 | 87.9 |
| Up-Down | 78.6 | 75.1 | 85.0 | 87.7 | 7.76 | 88.1 |
| X-LAN | 78.4 | 74.8 | 84.8 | 87.4 | 7.76 | 87.7 |
| Transformer | 78.1 | 74.6 | 84.5 | 87.2 | 7.71 | 87.6 |
| LSTM-A | 77.6 | 74.1 | 84.5 | 87.3 | 7.67 | 87.7 |
| AoANet | 77.3 | 73.4 | 84.5 | 87.2 | 7.62 | 87.6 |
| SCST | 77.0 | 73.3 | 83.9 | 86.7 | 7.60 | 87.1 |
| LSTM | 73.8 | 70.1 | 81.7 | 85.0 | 7.31 | 85.4 |
Percent, except CIDEr, a raw score that tops out at 10 in the COCO toolkit.
Three things stand out:
- Every captioner lands close. A plain CNN-LSTM scores 70.1, and the seven other baselines sit within 2.4 points of each other; which one comes second depends on the metric. The full model’s lead over the best of them is 1.2.
- The text is formulaic. On COCO, with five references per image, the X-Transformer’s own paper reports a BLEU-4 of 39.7. Here, with a single reference, it scores 75.7.
- The test’s unit is missing. The paper reports p < 0.05 on all six metrics, from a paired t-test against the second-best method. If the unit is the lesion, the 5,159 lesions are not independent: they sit in 120 patients. No interval or repeat run is reported.
05Two rulers: BLEU-4 and word matching
What does a BLEU-4 in the high 70s reward? Here are three outputs from the paper’s own examples, each scored against its reference:
| Output | Differs from its reference by | Error | BLEU-4 |
|---|---|---|---|
| X-Transformer, example 1 | neck level IIb for Va | wrong location | 82.7 |
| X-Transformer, example 4 | pulmonary nodule for pleural thickening | wrong finding | 77.6 |
| Full model, example 4 | observed for noted | none: a synonym | 78.3 |
Table 1 — BLEU-4 can’t tell a wrong nodal level from a harmless synonym. Sentence-level BLEU-4 × 100, on the English text printed in the paper’s Figure 3, lower-cased and without punctuation. With punctuation kept as tokens, the three rows score 80.0, 79.2 and 80.0. Our arithmetic.
BLEU counts broken word sequences, so where a changed word sits matters more than what it means. A wrong nodal level near the end of a sentence costs less than a synonym inside it, and a wrong finding costs about the same as the synonym. All three sit at or above the full model’s test-set score of 76.9.
To its credit, the paper’s ablation adds a second ruler: accuracy on location words and on CT-finding words, by word matching, presumably against the reference.
Show as table
| Setting | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | Location acc. | CT-finding acc. |
|---|---|---|---|---|---|---|
| Baseline (X-Transformer) | 86.3 | 82.3 | 78.9 | 75.7 | 81.0 | 65.5 |
| + window (DWS) | 86.5 | 82.6 | 79.3 | 76.0 | 81.2 | 67.3 |
| + prompts (CGLP) | 86.8 | 83.0 | 79.7 | 76.5 | 82.2 | 66.2 |
| + both | 87.1 | 83.3 | 80.1 | 76.9 | 82.9 | 67.3 |
In error terms, the two modules together remove one location error in ten and one finding error in twenty derived.
Two cautions. The baseline reads only the close-up and already scores 81.0 on location, so most locations can be named from a 9.6 cm cube. And the two measures are defined in one sentence, with no vocabulary and no matching rule. In the paper’s third example the authors call their output the same anatomical area as the reference, yet the two share no location word.
06What the paper leaves out
- Scanner makers: the abstract says the gains hold across makers, but no table or figure splits a result by maker, and the maker mix of the test set isn’t given.
- The baselines: only the four ablation rows get the two accuracies. No medical report generator or 3D vision–language model is compared.
- Promised, or never done: the rebuttal promises cross-validation and a failure-case analysis. Neither is in the final paper, and a reviewer’s request for an external test set went unanswered. No reader rated the captions, though the rebuttal says the system already runs in reading software at several centres.
- What can’t be re-run: there is no code, no data, no list of the 130 locations or their windows, and no detail on the encoders.
- What can be reused: the two crops, the 0.95 rule and three window settings. The decoder has public code: the X-Transformer’s repository implements the three baselines shown in the paper’s example figure. For data, PETARSeg-11K pairs 11,356 lesion descriptions with masks, with access on request.
07Takeaways for building a PET/CT report generator
- Score lesion sentences as fields. Location and finding, each with a published vocabulary. By our arithmetic, a wrong nodal level still earns a sentence BLEU-4 of 82.7.
- Count patients, not lesions. With 43 lesions per test patient, resample by patient, and report each scanner maker.
- Use the cheap tricks, and expect small gains. A window chosen by region added 1.8 points of CT-finding accuracy, and location names as a prompt added 1.2 of location accuracy. Gate both on confidence.
- Decide what the PET is for. If the sentence leaves out uptake, measure SUVmax from the mask. If it should state uptake, the reference captions must state it too.
- A sentence per lesion isn’t a report. Forty-three sentences per patient still have to be merged, compared with the previous scan and summed up in an impression.
ASources
- Yu, M., Gao, Y., Shu, Y., et al. Location-Guided Automated Lesion Captioning in Whole-body PET/CT Images. MICCAI 2025, LNCS 15964, pp. 348–357. The page holds the three reviews and the rebuttal. PDF · doi. No code or data is released, and we found no preprint.
- Gao, Y., Shu, Y., Yu, M., et al. Hierarchical Contrastive Learning for Precise Whole-Body Anatomical Localization in PET/CT Imaging. IEEE Trans. Med. Imaging 45(1), 391–405 (2026): abstract. A related conference paper by the same team: Yu, M., et al., IPMI 2025, LNCS 15830, pp. 234–246, doi.
- Pan, Y., Yao, T., Li, Y., Mei, T. X-Linear Attention Networks for Image Captioning. CVPR 2020: arXiv:2003.14080 (Table 1 for the COCO scores); code.
- COCO caption evaluation toolkit: tylin/coco-caption (BLEU, METEOR, ROUGE-L, CIDEr and SPICE).
- Maqbool, D., et al. PETAR: arXiv:2510.27680 (CVPR 2026); PETARSeg-11K on Hugging Face, checked on 5 Oct 2026.