Jean-Benoit Delbrouck / Research blog

Research blog · CT report generation

NV-Reason-CT, read for report generation

NVIDIA’s 3D CT model gives its language model all 13,824 visual tokens of a crop, learns its reasoning style from radiologists’ narrated reads, and is tuned with reinforcement learning to get its finding list right, which is close to what the benchmarks measure. What that buys, and what the paper’s comparisons can’t tell you.

Paper Andriy Myronenko, Dong Yang, … Pengfei Guo, Daguang Xu. NV-Reason-CT: 3D Visual Language Model for CT Analysis. arXiv:2609.27511, v1, 23 Sep 2026 By Jean-Benoit Delbrouck Published 28 Sep 2026 Reading time ~12 min
13,824
visual tokens per CT crop, passed to the LLM unmerged
0.592
report-derived macro-F1 on CT-RATE, 18 findings
550k
instruction examples from 70,111 CT inputs
2 × 10
radiologists × cases behind “50% faster reporting”
TL;DR
  • An open, end-to-end 3D CT reporter. All 13,824 tokens of a 38.4 cm crop reach Qwen3.5-4B with their 3D positions. The whole model is fine-tuned on about 550k instructions, then GRPO, a reinforcement-learning method, rewards the report’s finding list.
  • The best clinical scores in its CT-RATE table, from a mixed comparison. Its report F1 (0.592) beats each of the seven other systems with per-finding numbers, on all 18 findings. But it is the only entry scored on the corrected 3,002-scan set, and no result has a confidence interval.
  • Best on findings, near the bottom on wording. BLEU-4 is 0.11, against 0.21 for 2024’s CT2Rep. The authors point to format, and one epoch of format adaptation puts it first on four of five text metrics on Merlin.
  • What we can’t tell yet. No ablations separate the unmerged tokens, the reasoning data and GRPO. “50% faster reporting” comes from 2 radiologists reading 10 cases, always unaided first, with reader-reported times.

01Why this paper matters for report generation

Many CT report generators squeeze the image into far fewer tokens and train only a thin bridge plus LoRA weights, like the nnFoundation setup in our previous post. NV-Reason-CT does the opposite at every step: it keeps every visual token, tells the language model where each one sits in 3D, fine-tunes the whole model, and then applies reinforcement learning to the finding list. The weights and the training code for both stages, supervised fine-tuning (SFT) and GRPO, are public under the permissive OpenMDW-1.1 license. The paper also covers classification and interactive reasoning; these notes stay on report generation.

Figure 1 — The whole stack is trainable. Unlike frozen-encoder designs, the encoder, projector and language model are all updated. Sources: Sections 3.2–3.9.

02The data: three sources, about eight examples per image

Training uses three CT sources with reports: public chest CT from CT-RATE, an internal NIH set with mostly whole-body coverage, and CancerVerse, whose reports cover only cancer, so it feeds cancer questions and never serves as a report target. Where the field of view allows, one volume yields both a chest and an abdominal crop.

SourceTraining volumesChest cropsAbdominal cropsEvaluation volumes
CT-RATE47,14946,20320,2433,002 (corrected v2; 1,564 studies)
NIH internal15,99112,74212,2435,092
CancerVerse22,7204,6684,438–
Merlin–––5,125 (abdomen)
RAD-ChestCT–––3,630 (chest; labels only)

Table 1 — Where training and evaluation data come from. Crop counts are unique regional crops per task. Sources: Table 1 and Section 4.1.

The crops add up to 100,537 (63,613 chest, 36,924 abdominal) derived, while the text reports 70,111 unique image inputs; the paper doesn’t reconcile the two. From those inputs the authors build about 550,000 instruction examples, roughly eight per input derived: structured reports, first-person reasoning walkthroughs, follow-up questions, yes/no questions per finding, section-specific reports, location and severity questions, and refusals.

Where the reasoning style comes from

Radiologists recorded themselves reading CTs in their usual workflow, narrating what they saw, suspected and were unsure of. The transcripts serve as training targets and as templates for rewriting ordinary reports into synthetic narrations in the same voice, so a small set of expert reads sets the style of a much larger corpus. The paper calls the set limited but gives no count, and doesn’t name the rewriting model. The recipe is adapted from NVIDIA’s chest X-ray model, NV-Reason-CXR, where 100 radiologist-narrated cases guided GPT-OSS-120B in rewriting 100,000 reports.

One curation detail is worth copying. A radiologist-guided ontology of 30 chest and 29 abdominal findings harmonises the sources, with rules for synonyms, negation, uncertainty and temporal status. An unmentioned finding counts as absent only if its region was fully imaged; otherwise it is masked, so a scan that never covered the pelvis can’t teach the model that the pelvis is normal.

03The model: keep every token, and tell the LLM where it is

The Primus encoder turns the crop into a 24 × 24 × 24 grid of 864-dimensional features. The projector lifts each one to Qwen3.5-4B’s 2,560-dimensional width and deliberately doesn’t merge neighbours, and each token carries its (depth, height, width) grid index into the language model’s multimodal rotary position encoding (MRoPE). That is a lot of context for a 4B model: at the 1,300-token completion cap used in GRPO, the image is about 91% of the sequence derived. One nuance from the released config: Qwen3.5-4B has rotary positions only in its 8 full-attention layers (of 32), on a quarter of each head’s dimensions, so that is where the 3D positions act our observation.

nnFoundation stack (post 001)NV-Reason-CT
Voxel size, crop1 mm, 192³ = 19.2 cm cube2 mm, 192³ = 38.4 cm cube
ViT patch8³ voxels = 8 mm cube8³ voxels = 16 mm cube
Tokens to the LLM1,728 (2×2×2 merge) · 16 mm each13,824 (no merge) · 16 mm each
Encoder during report trainingfrozen (MAE-pretrained)fine-tuned (COLIPRI init)
Language modelQwen2.5-VL-3B, LoRAQwen3.5-4B, full fine-tuning
Report supervisionCT-RATE reports (≈12 epochs)~550k instructions + GRPO
CT-RATE F1 / CRG0.437 / 0.444 (split and averaging unstated)0.592 / 0.514 (corrected 3,002 set; macro)

Table 2 — Same 16 mm tokens, eight times the volume. Both stacks use a Primus-style tokenizer and give the language model one token per 16 mm cube: nnFoundation by merging 8 mm patches, NV-Reason-CT with 16 mm patches and no merge. This is context, not a controlled comparison. Sources: this paper’s Sections 3–4 and Table 4; nnFoundation arXiv:2609.26924.

The bigger crop takes in the whole chest, which answers the field-of-view worry we raised about nnFoundation’s 19.2 cm crop our observation. The cost is context length, and anything outside the crop is still lost, as the authors note.

04Training: supervised fine-tuning, then GRPO on finding lists

SFT trains the encoder, projector and language model jointly on the instruction mixture, with loss only on the assistant’s reply. Reasoning answers wrap the clinical analysis in <think>…</think>, and report answers end with exactly one <answer>…</answer> block listing the findings, comma-separated. That list is the hook for the second stage.

SFT
All components · fused AdamW · base LR 2×10⁻⁵ (encoder 0.1×, projector 5×, LLM 1×) · cosine schedule, 3% warm-up · gradient clipping 0.3
GRPO
16 samples per prompt at temperature 1.0 · at most 1,300 tokens · LR 1×10⁻⁶, 5 warm-up steps, cosine to 5% · clipping ε 0.20 / 0.28 · no KL penalty (β = 0) · advantages scaled by the batch-level reward SD
Hardware
16 nodes × 8 H100 · per-device batch 1, 2 accumulation steps (256 sequences per update if data-parallel across all 128 GPUs; derived)

GRPO samples 16 reports per prompt, scores each, and pushes the policy toward those that beat their group’s mean. The release builds it on TRL’s GRPO trainer, adapted to 3D, and three settings stand out: DAPO’s asymmetric clip-higher range, which lets a token’s probability rise further in one update than it can fall; rewards scaled by the batch’s standard deviation rather than the group’s; and no KL term tying the policy to the SFT model. Here is the reward, condensed from the released vlm_rewards.py and Section 3.8:

def reward(text, ref, region, n_tokens):
    """2.0 * set-F1 + 0.5 * structure + length (Sec. 3.8)."""
    r_set = 0.0
    if text.count("<answer>") == 1 == text.count("</answer>"):
        body = text.split("<answer>")[1].split("</answer>")[0]
        pred = {s.strip() for s in body.split(",") if s.strip()}
        pred = pred or {NO_FINDING[region]}   # empty = normal
        if not (NO_FINDING[region] in pred and len(pred) > 1):
            r_set = 2 * len(pred & ref) / (len(pred) + len(ref))
    r_struct = structure_score(text, region)    # in [0, 1]
    if n_tokens < 180:
        r_len = max(-1.0, (n_tokens - 180) / 180)
    elif n_tokens > 900:
        r_len = max(-1.0, (900 - n_tokens) / 300)
    else:
        r_len = 0.0
    return 2.0 * r_set + 0.5 * r_struct + r_len

Four-fifths of the best-case reward (2.0 of 2.5) is finding agreement. structure_score mostly checks that the organ headings are present and in order (weight 0.55); the rest is non-empty section text (0.20), the common report blocks (0.15) and a numbered or bulleted impression (0.10). The length term only subtracts, against bare lists and runaway reports.

What the reward does and doesn’t see our observation

Only the <answer> list is scored for content; the prose is rewarded for structure alone. The two could have drifted apart, but the scores suggest they didn’t: RadBERT reading the prose gives 0.592, and the list itself 0.601. The catch is that the reward and the benchmark both come down to agreement on finding sets, so some of the gain may be the model learning what the metric counts.

05Results on CT-RATE

The evaluation uses the corrected (v2) CT-RATE validation manifest: 3,002 chest reconstructions from 1,564 studies, with 37 non-chest volumes removed from the historical 3,039. The official RadBERT labeler turns each generated Findings + Impression into 18 labels at a 0.5 threshold, and positive-class F1 is averaged with equal weight over the labels. That is macro-F1, stated in Equation 1; for nnFoundation we had to infer it. Reconstructions of one study share its reference report, so systems scored on the 1,564 studies aren’t on quite the same footing.

This paper (corrected 3,002 set)Same labeler and macro-F1, 3,039 setOther or unstated protocol00.10.20.30.40.50.60.7NV-Reason-CT3,002 · RadBERT0.592CT-AGRG3,039 · avg. n/s0.501RMR1,564 reports0.495Astra1,564 · from a chart0.484COLIPRI-CRMcohort not stated0.449VoxelFM3,039 · GPT labels0.432CT-SSG5-fold · avg. n/s0.387MonteRET1,5640.365MS-VLM3,039 · labeler n/s0.261BTB3D3,039 · RadBERT0.258ClinFusion-32Bclaim-level F10.239ClinFusion-8Bclaim-level F10.216CT-CHAT3,039 · RadBERT0.184CT2Rep3,039 · RadBERT0.160M3D-LaMedmicro-F10.148RadFM3,039 · RadBERT0.059
Figure 2 — Highest in the table, but the table mixes protocols. Report-derived F1 on CT-RATE, shaded by how each number was produced; the tags give each row’s cohort and scoring. Source: Table 4.
Show as table
MethodEvaluationF1CRGGREENBLEU-4ROUGE-LMETEOR
NV-Reason-CTauthors' evaluation on the corrected 3,002 set; RadBERT on Findings + Impression0.5920.5140.4070.1150.2140.192
CT-AGRGauthors' evaluation on 3,039; CT-Net variant; F1 averaging not stated0.501––0.1720.2800.196
RMR1,564 test reports; reconstruction handling not stated0.495––0.2030.2800.235
Astra1,564 cases; macro-F1 approximated from a bar chart0.484––0.2500.4410.240
COLIPRI-CRMevaluation cohort size not stated0.449–––––
VoxelFM3,039; GPT-OSS-120B abnormality extraction0.432–––––
CT-SSGfive-fold mean; F1 averaging not stated0.3870.437–0.1070.2460.164
MonteRET1,564 scans; macro-F1 derived from per-finding F10.365––0.2520.3790.454
MS-VLMauthors' evaluation on 3,039; labeler not stated in this paper0.261––0.2320.4380.396
BTB3Dauthors' evaluation on 3,0390.2580.370–0.215–0.223
ClinFusion-32B3,039; GPT-4.1 claim-level F10.239–––––
ClinFusion-8B3,039; GPT-4.1 claim-level F10.216–––––
CT-CHATCRG-paper evaluation on 3,0390.1840.3680.4360.1980.3260.215
CT2RepCRG-paper evaluation on 3,0390.1600.3590.4870.2130.3620.197
M3D-LaMedCT-Agent evaluation; micro-F10.148––0.2450.4000.326
RadFMCRG-paper evaluation of the pretrained checkpoint0.0590.3350.0180.0000.0420.020

Four things stand out:

  1. The lead is clear against the most comparable systems and smaller against the rest. CT2Rep, CT-CHAT, BTB3D and RadFM, scored with the same labeler and averaging on the historical 3,039 set, reach at most 0.258. The next rows use other or unstated protocols: CT-AGRG (0.501) doesn’t state its averaging, RMR (0.495) and Astra (0.484) score the 1,564 studies, and Astra’s value was read off a chart.
  2. It leads on every finding. Its report F1 is higher than the best of seven other systems on all 18 findings (Figure 3), though two of those systems’ values were digitised from figures and the cohorts and labelers differ.
  3. The report is nearly as good as asking. Macro-F1 is 0.614 from 18 yes/no questions, 0.610 from one prompt listing all findings, 0.601 from the report’s own finding list and 0.592 from RadBERT reading its prose: a spread of 0.022. The report adds location, severity and findings beyond the 18 labels.
  4. Best on findings, near the bottom on wording. BLEU-4 is 0.11 and ROUGE-L 0.21, below CT2Rep (0.21 and 0.36), a 2024 model with F1 0.160 (Figure 4). The authors’ explanation is format: structured reports scored against free-narrative references, which Merlin supports (section 06). GREEN, an LLM judge of clinical errors, also puts it below CT2Rep (0.41 against 0.49), though its checkpoint is out of domain for CT.
NV-Reason-CT, reportNV-Reason-CT, yes/no questionBest other published report0.20.30.40.50.60.70.80.9Pleural effusionArterial wall calcificationCoronary artery wall calcificationLung opacityLung noduleConsolidationMedical materialAtelectasisCardiomegalyPericardial effusionLymphadenopathyHiatal herniaEmphysemaPulmonary fibrotic sequelaBronchiectasisMosaic attenuation patternInterlobular septal thickeningPeribronchial thickening
Figure 3 — Above the best published report on all 18 findings. Grey is the best of seven other systems for each finding (* in the table: digitised from the authors’ figures). Sources: Tables B1 and A1; positive counts from the paper’s Figure 7.
Show as table
FindingReportYes/noBest other reportByPositives
Pleural effusion0.8470.8590.724COLIPRI-CRM*370
Arterial wall calcification0.7660.7750.700MonteRET854
Coronary artery wall calcification0.7560.7770.642COLIPRI-CRM*758
Lung opacity0.7340.7450.680COLIPRI-CRM*1,177
Lung nodule0.6870.6590.605VoxelFM*1,351
Consolidation0.6820.7120.569VoxelFM*578
Medical material0.6790.7270.414COLIPRI-CRM*307
Atelectasis0.5970.6040.489COLIPRI-CRM*704
Cardiomegaly0.5850.6390.508MonteRET316
Pericardial effusion0.5610.5780.380MonteRET217
Lymphadenopathy0.5450.5750.451COLIPRI-CRM*773
Hiatal hernia0.5310.5410.480MonteRET417
Emphysema0.5210.5270.464VoxelFM*591
Pulmonary fibrotic sequela0.4910.5410.453COLIPRI-CRM*822
Bronchiectasis0.4610.4680.334COLIPRI-CRM*327
Mosaic attenuation pattern0.4570.4520.425VoxelFM*246
Interlobular septal thickening0.3920.4290.353VoxelFM*246
Peribronchial thickening0.3700.4410.273VoxelFM*348
This paperSame labeler, 3,039 setOther protocol0.00.10.20.30.40.50.60.700.050.10.150.20.250.3BLEU-4 (n-gram overlap with the reference report)report-derived F1NV-Reason-CTCT-AGRGRMRAstraCT-SSGMonteRETMS-VLMBTB3DCT-CHATCT2RepM3D-LaMedRadFM
Figure 4 — Text overlap doesn’t track clinical accuracy. One point per CT-RATE system with both BLEU-4 and report-derived F1; shapes mark the protocol as in Figure 2. Source: Table 4.

06Beyond CT-RATE: Merlin and RAD-ChestCT

Merlin (abdominal CT), after format adaptation

Merlin’s reports use short organ-specific subsections, unlike NV-Reason-CT’s native format. So the authors first fine-tuned the model for one epoch on Merlin’s 15,314 released training reports, with a Findings-only prompt listing the 15 Merlin headings; no test report was used. The comparator numbers come from the Jolia paper, whose authors ran all four models on the same test split, possibly with different evaluator versions.

BLEU00.30.60.077ROUGE-L00.30.60.336BERTScore00.30.60.579RadGraph-F100.30.60.322GREEN00.30.60.366NV-Reason-CT†JoliaMerlinMedGemma 1.5Med3DVLM
Figure 5 — After format adaptation, it tops four of five text metrics. Findings-only reports on all 5,125 Merlin test scans; † fine-tuned for one epoch on Merlin’s training reports. Its RadGraph-F1 lead over Jolia is 0.005, with no confidence intervals. Source: Table 6.
Show as table
MethodBLEUROUGE-LBERTScoreRadGraph-F1GREEN
NV-Reason-CT†0.0770.3360.5790.3220.366
Jolia0.1190.3230.5670.3170.324
Merlin0.0630.2850.5510.2370.316
MedGemma 1.50.1240.1060.3610.0320.038
Med3DVLM0.0040.0790.2560.0100.007

This is the closest thing to a control in the paper: once its reports match the reference format, the model with some of the lowest wording scores on CT-RATE comes out on top, which suggests much of that gap is format. Two caveats: an epoch on 15,314 in-domain reports teaches content too, and the paper leaves out the sixth metric in Jolia’s table, CRIMSON, on which Merlin scores best.

RAD-ChestCT (external chest CT), without tuning

RAD-ChestCT publishes labels but no report text, so reports are scored only through the 16 harmonised findings extracted from them. NV-Reason-CT ran with its default prompt and no RAD-ChestCT tuning.

00.10.20.30.40.50.6NV-Reason-CT · finding list0.510NV-Reason-CT · RadBERT0.485DCP-PD0.395EXACT-CHAT0.289BTB3D0.266CT-CHAT0.182M3D-LaMed0.113RadFM0.069
Figure 6 — The lead carries over to an unseen chest dataset. Macro-F1 over 16 labels on 3,630 volumes (two comparators use a modified label mapping; BTB3D’s count is unstated). RadBERT truncated 288 reports (7.9% derived) at its 512-token limit; reading the finding list directly avoids this. Source: Table 8.
Show as table
MethodMacro-F1PrecisionRecall
NV-Reason-CT · finding list0.5100.5160.552
NV-Reason-CT · RadBERT0.4850.5100.494
DCP-PD0.3950.4430.418
EXACT-CHAT0.2890.4690.298
BTB3D0.2660.2720.329
CT-CHAT0.1820.3820.171
M3D-LaMed0.1130.2690.080
RadFM0.0690.2830.044

07The reader study behind “50% faster”

Two US board-certified radiologists, each with at least 10 years of practice, read 10 cases: 5 chest and 5 abdomen, 2 normal and 8 abnormal. They reported each case unaided first, then, after a short distraction, reviewed it with the model’s reasoning and structured report, which they could copy and edit. Reported time per case fell from 26.25 to 13.13 minutes. Nine of ten agreement statements averaged at least 4.60 on a 1–5 scale; the exception was the AI finding something they had missed (2.60), and the readers said their careful first read had missed nothing.

How much weight to put on it

Treat it as a feasibility signal. It is 20 case-reads, and the assisted read always came second on the same case, so memory of the first read helps; the authors flag these recall and order effects. The times are reported by the readers, and the final reports weren’t scored for accuracy. The paper calls the study preliminary and not an assessment of clinical safety.

08Reading the metrics

  • CRG, re-expressed. CRG weights true positives, false negatives and false positives by prevalence. It works out to CRG = 1/(3 − 2J), with J = TPR − FPR over the pooled label slots, so chance is 1/3. NV-Reason-CT’s 0.514 means J ≈ 0.53 derived, against about 0.37 for the frozen nnFoundation stack, on a different evaluation set.
  • GREEN is out of domain. Its official checkpoint was trained mainly on chest X-ray reports, so read its CT scores loosely. The authors say so.
  • RadBERT reads at most 512 tokens. Long structured reports get truncated (Figure 6); the paper doesn’t give the CT-RATE count.

09What the paper leaves out

  • Ablations: none for the unmerged tokens, the expert reasoning data or GRPO, so their separate contributions are unknown. The authors list this as a limitation.
  • Data: the number of expert narrations, the model that wrote the synthetic narrations and labels, per-family example counts and a breakdown of the 70,111 inputs. The release adds no training data beyond 128 illustrative CT-RATE examples, and no evaluation code.
  • Training budget: SFT epochs or steps, GRPO steps and wall-clock time. Only the hardware is given.
  • Uncertainty and a fair baseline: no confidence intervals or significance tests for its report metrics, and no strong generator re-run on the 3,002-scan set with the same labeler.
  • Faithfulness: hallucination rates, and whether the <think> text reflects what drives the answer, aren’t measured. The authors say the reasoning “is not assumed to reveal the model’s internal computation”.

10Takeaways for building a CT report generator

  1. Test whether you can afford to keep the tokens. The best clinical numbers here come with 13,824 unmerged tokens, 3D positions and end-to-end tuning. Without an ablation, whether merging costs accuracy is still your experiment to run.
  2. Reward what you measure, knowingly. GRPO on the finding list ties training to label-based evaluation, and a structure reward keeps the prose organised.
  3. Emit a machine-readable finding list. It gives a labeler-free F1 (0.601 against 0.592) and sidesteps RadBERT’s 512-token limit.
  4. Report text metrics, but don’t rank on them. After one epoch of format adaptation, the model near the bottom on CT-RATE wording led four of five text metrics on Merlin.
  5. Mask what wasn’t imaged. Coverage-aware labels stop partial scans from teaching the model false negatives.

ASources

  1. Myronenko, A., Yang, D., et al. NV-Reason-CT: 3D Visual Language Model for CT Analysis. arXiv:2609.27511 (2026).
  2. NVIDIA-Medtech. NV-Reason-CT code · model card.
  3. NVIDIA. Reasoning Visual Language Model for Chest X-Ray Analysis (NV-Reason-CXR). arXiv:2510.23968 (2025).
  4. Yu, Q., et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv:2503.14476 (2025).
  5. Qwen Team. Qwen3.5-4B.
  6. Hamamci, I. E., et al. CT-RATE: dataset; CRG score (2025).
  7. Our previous post: nnFoundation, read for report generation.