Jean-Benoit Delbrouck / Research blog

Research blog

Long-form reading notes on radiology AI papers, with a focus on report generation.

CT report generationarXiv 202628 Sep 2026~12 min read

NV-Reason-CT, read for report generation

The report-generation side of NVIDIA’s NV-Reason-CT: all 13,824 visual tokens, radiologist-guided reasoning data and GRPO on finding lists, and what its CT-RATE, Merlin and RAD-ChestCT comparisons can and can’t support.

CT report generationarXiv 202625 Sep 2026~15 min read

nnFoundation, read for report generation

The report-generation side of DKFZ’s 2.16M-volume radiology foundation model: where the data came from, how a 674M 3D ViT was pretrained without a single report, how it was wired into Qwen2.5-VL-3B, and what the appendix numbers show.

CT report generationNature 20262 Oct 2026~11 min read

Merlin, read for report generation

The report-generation side of Merlin, the abdominal CT foundation model published in Nature: 490 visual tokens, one linear adapter and a LoRA-tuned language model writing one organ section at a time, what the scores show, and why the released dataset matters more.

Chest X-ray report generationMICCAI 20255 Oct 2026~8 min read

RadAlign, read for report generation

RadAlign trains a small classifier, then lets an off-the-shelf LLM write the chest X-ray report from its predictions and seven retrieved reports. What the LLM is told about the image, and what the GREEN gain over a 2021 baseline is made of.

PET/CT report generationMICCAI 20255 Oct 2026~8 min read

PET/CT lesion captioning, read for report generation

A model writes one sentence per lesion on whole-body PET/CT, guided by a predicted anatomical location. Its captions reach a BLEU-4 of 76.9 while its own word-matching check puts CT-finding accuracy at 67.3%: a lesson in how to score lesion-level reports.

Pathology report generationMICCAI 20255 Oct 2026~8 min read

Melanocytic lesion reports, read for report generation

A small vision–language model writes pathology reports for melanocytic skin lesions, and one pathologist scores 50 of them blind against the originals: level on ordinary moles, far behind on the other lesions. How far that evidence goes.

Report-generation evaluationMICCAI 20255 Oct 2026~8 min read

SPEC-CXR, read for report-generation evaluation

SPEC-CXR scores a generated chest X-ray report finding by finding, over 122 fixed entities. What that adds to report-level metrics, how the scorer was validated, and what it says about three leading generators.

CT report generationMICCAI 20255 Oct 2026~8 min read

The 3D MLLM design space, read for report generation

A MICCAI 2025 study swaps the LLM, the projector, the tuning method and the input size of a 3D CT report generator trained on 1,287 scans. No change to the model adds more than 0.006 GREEN, while canned “normal” sentences add 0.078. What that says about small-data training, and about the metric.