- An open, end-to-end 3D CT reporter. All 13,824 tokens of a 38.4 cm crop reach Qwen3.5-4B with their 3D positions. The whole model is fine-tuned on about 550k instructions, then GRPO, a reinforcement-learning method, rewards the report’s finding list.
- The best clinical scores in its CT-RATE table, from a mixed comparison. Its report F1 (0.592) beats each of the seven other systems with per-finding numbers, on all 18 findings. But it is the only entry scored on the corrected 3,002-scan set, and no result has a confidence interval.
- Best on findings, near the bottom on wording. BLEU-4 is 0.11, against 0.21 for 2024’s CT2Rep. The authors point to format, and one epoch of format adaptation puts it first on four of five text metrics on Merlin.
- What we can’t tell yet. No ablations separate the unmerged tokens, the reasoning data and GRPO. “50% faster reporting” comes from 2 radiologists reading 10 cases, always unaided first, with reader-reported times.
01Why this paper matters for report generation
Many CT report generators squeeze the image into far fewer tokens and train only a thin bridge plus LoRA weights, like the nnFoundation setup in our previous post. NV-Reason-CT does the opposite at every step: it keeps every visual token, tells the language model where each one sits in 3D, fine-tunes the whole model, and then applies reinforcement learning to the finding list. The weights and the training code for both stages, supervised fine-tuning (SFT) and GRPO, are public under the permissive OpenMDW-1.1 license. The paper also covers classification and interactive reasoning; these notes stay on report generation.
= a 38.4 cm cube
report, <answer> list
02The data: three sources, about eight examples per image
Training uses three CT sources with reports: public chest CT from CT-RATE, an internal NIH set with mostly whole-body coverage, and CancerVerse, whose reports cover only cancer, so it feeds cancer questions and never serves as a report target. Where the field of view allows, one volume yields both a chest and an abdominal crop.
| Source | Training volumes | Chest crops | Abdominal crops | Evaluation volumes |
|---|---|---|---|---|
| CT-RATE | 47,149 | 46,203 | 20,243 | 3,002 (corrected v2; 1,564 studies) |
| NIH internal | 15,991 | 12,742 | 12,243 | 5,092 |
| CancerVerse | 22,720 | 4,668 | 4,438 | – |
| Merlin | – | – | – | 5,125 (abdomen) |
| RAD-ChestCT | – | – | – | 3,630 (chest; labels only) |
Table 1 — Where training and evaluation data come from. Crop counts are unique regional crops per task. Sources: Table 1 and Section 4.1.
The crops add up to 100,537 (63,613 chest, 36,924 abdominal) derived, while the text reports 70,111 unique image inputs; the paper doesn’t reconcile the two. From those inputs the authors build about 550,000 instruction examples, roughly eight per input derived: structured reports, first-person reasoning walkthroughs, follow-up questions, yes/no questions per finding, section-specific reports, location and severity questions, and refusals.
Where the reasoning style comes from
Radiologists recorded themselves reading CTs in their usual workflow, narrating what they saw, suspected and were unsure of. The transcripts serve as training targets and as templates for rewriting ordinary reports into synthetic narrations in the same voice, so a small set of expert reads sets the style of a much larger corpus. The paper calls the set limited but gives no count, and doesn’t name the rewriting model. The recipe is adapted from NVIDIA’s chest X-ray model, NV-Reason-CXR, where 100 radiologist-narrated cases guided GPT-OSS-120B in rewriting 100,000 reports.
One curation detail is worth copying. A radiologist-guided ontology of 30 chest and 29 abdominal findings harmonises the sources, with rules for synonyms, negation, uncertainty and temporal status. An unmentioned finding counts as absent only if its region was fully imaged; otherwise it is masked, so a scan that never covered the pelvis can’t teach the model that the pelvis is normal.
03The model: keep every token, and tell the LLM where it is
The Primus encoder turns the crop into a 24 × 24 × 24 grid of 864-dimensional features. The projector lifts each one to Qwen3.5-4B’s 2,560-dimensional width and deliberately doesn’t merge neighbours, and each token carries its (depth, height, width) grid index into the language model’s multimodal rotary position encoding (MRoPE). That is a lot of context for a 4B model: at the 1,300-token completion cap used in GRPO, the image is about 91% of the sequence derived. One nuance from the released config: Qwen3.5-4B has rotary positions only in its 8 full-attention layers (of 32), on a quarter of each head’s dimensions, so that is where the 3D positions act our observation.
| nnFoundation stack (post 001) | NV-Reason-CT | |
|---|---|---|
| Voxel size, crop | 1 mm, 192³ = 19.2 cm cube | 2 mm, 192³ = 38.4 cm cube |
| ViT patch | 8³ voxels = 8 mm cube | 8³ voxels = 16 mm cube |
| Tokens to the LLM | 1,728 (2×2×2 merge) · 16 mm each | 13,824 (no merge) · 16 mm each |
| Encoder during report training | frozen (MAE-pretrained) | fine-tuned (COLIPRI init) |
| Language model | Qwen2.5-VL-3B, LoRA | Qwen3.5-4B, full fine-tuning |
| Report supervision | CT-RATE reports (≈12 epochs) | ~550k instructions + GRPO |
| CT-RATE F1 / CRG | 0.437 / 0.444 (split and averaging unstated) | 0.592 / 0.514 (corrected 3,002 set; macro) |
Table 2 — Same 16 mm tokens, eight times the volume. Both stacks use a Primus-style tokenizer and give the language model one token per 16 mm cube: nnFoundation by merging 8 mm patches, NV-Reason-CT with 16 mm patches and no merge. This is context, not a controlled comparison. Sources: this paper’s Sections 3–4 and Table 4; nnFoundation arXiv:2609.26924.
The bigger crop takes in the whole chest, which answers the field-of-view worry we raised about nnFoundation’s 19.2 cm crop our observation. The cost is context length, and anything outside the crop is still lost, as the authors note.
04Training: supervised fine-tuning, then GRPO on finding lists
SFT trains the encoder, projector and language model jointly on the instruction mixture, with loss only on the assistant’s reply. Reasoning answers wrap the clinical analysis in <think>…</think>, and report answers end with exactly one <answer>…</answer> block listing the findings, comma-separated. That list is the hook for the second stage.
- SFT
- All components · fused AdamW · base LR 2×10⁻⁵ (encoder 0.1×, projector 5×, LLM 1×) · cosine schedule, 3% warm-up · gradient clipping 0.3
- GRPO
- 16 samples per prompt at temperature 1.0 · at most 1,300 tokens · LR 1×10⁻⁶, 5 warm-up steps, cosine to 5% · clipping ε 0.20 / 0.28 · no KL penalty (β = 0) · advantages scaled by the batch-level reward SD
- Hardware
- 16 nodes × 8 H100 · per-device batch 1, 2 accumulation steps (256 sequences per update if data-parallel across all 128 GPUs; derived)
GRPO samples 16 reports per prompt, scores each, and pushes the policy toward those that beat their group’s mean. The release builds it on TRL’s GRPO trainer, adapted to 3D, and three settings stand out: DAPO’s asymmetric clip-higher range, which lets a token’s probability rise further in one update than it can fall; rewards scaled by the batch’s standard deviation rather than the group’s; and no KL term tying the policy to the SFT model. Here is the reward, condensed from the released vlm_rewards.py and Section 3.8:
def reward(text, ref, region, n_tokens):
"""2.0 * set-F1 + 0.5 * structure + length (Sec. 3.8)."""
r_set = 0.0
if text.count("<answer>") == 1 == text.count("</answer>"):
body = text.split("<answer>")[1].split("</answer>")[0]
pred = {s.strip() for s in body.split(",") if s.strip()}
pred = pred or {NO_FINDING[region]} # empty = normal
if not (NO_FINDING[region] in pred and len(pred) > 1):
r_set = 2 * len(pred & ref) / (len(pred) + len(ref))
r_struct = structure_score(text, region) # in [0, 1]
if n_tokens < 180:
r_len = max(-1.0, (n_tokens - 180) / 180)
elif n_tokens > 900:
r_len = max(-1.0, (900 - n_tokens) / 300)
else:
r_len = 0.0
return 2.0 * r_set + 0.5 * r_struct + r_len
Four-fifths of the best-case reward (2.0 of 2.5) is finding agreement. structure_score mostly checks that the organ headings are present and in order (weight 0.55); the rest is non-empty section text (0.20), the common report blocks (0.15) and a numbered or bulleted impression (0.10). The length term only subtracts, against bare lists and runaway reports.
Only the <answer> list is scored for content; the prose is rewarded for structure alone. The two could have drifted apart, but the scores suggest they didn’t: RadBERT reading the prose gives 0.592, and the list itself 0.601. The catch is that the reward and the benchmark both come down to agreement on finding sets, so some of the gain may be the model learning what the metric counts.
05Results on CT-RATE
The evaluation uses the corrected (v2) CT-RATE validation manifest: 3,002 chest reconstructions from 1,564 studies, with 37 non-chest volumes removed from the historical 3,039. The official RadBERT labeler turns each generated Findings + Impression into 18 labels at a 0.5 threshold, and positive-class F1 is averaged with equal weight over the labels. That is macro-F1, stated in Equation 1; for nnFoundation we had to infer it. Reconstructions of one study share its reference report, so systems scored on the 1,564 studies aren’t on quite the same footing.
Show as table
| Method | Evaluation | F1 | CRG | GREEN | BLEU-4 | ROUGE-L | METEOR |
|---|---|---|---|---|---|---|---|
| NV-Reason-CT | authors' evaluation on the corrected 3,002 set; RadBERT on Findings + Impression | 0.592 | 0.514 | 0.407 | 0.115 | 0.214 | 0.192 |
| CT-AGRG | authors' evaluation on 3,039; CT-Net variant; F1 averaging not stated | 0.501 | – | – | 0.172 | 0.280 | 0.196 |
| RMR | 1,564 test reports; reconstruction handling not stated | 0.495 | – | – | 0.203 | 0.280 | 0.235 |
| Astra | 1,564 cases; macro-F1 approximated from a bar chart | 0.484 | – | – | 0.250 | 0.441 | 0.240 |
| COLIPRI-CRM | evaluation cohort size not stated | 0.449 | – | – | – | – | – |
| VoxelFM | 3,039; GPT-OSS-120B abnormality extraction | 0.432 | – | – | – | – | – |
| CT-SSG | five-fold mean; F1 averaging not stated | 0.387 | 0.437 | – | 0.107 | 0.246 | 0.164 |
| MonteRET | 1,564 scans; macro-F1 derived from per-finding F1 | 0.365 | – | – | 0.252 | 0.379 | 0.454 |
| MS-VLM | authors' evaluation on 3,039; labeler not stated in this paper | 0.261 | – | – | 0.232 | 0.438 | 0.396 |
| BTB3D | authors' evaluation on 3,039 | 0.258 | 0.370 | – | 0.215 | – | 0.223 |
| ClinFusion-32B | 3,039; GPT-4.1 claim-level F1 | 0.239 | – | – | – | – | – |
| ClinFusion-8B | 3,039; GPT-4.1 claim-level F1 | 0.216 | – | – | – | – | – |
| CT-CHAT | CRG-paper evaluation on 3,039 | 0.184 | 0.368 | 0.436 | 0.198 | 0.326 | 0.215 |
| CT2Rep | CRG-paper evaluation on 3,039 | 0.160 | 0.359 | 0.487 | 0.213 | 0.362 | 0.197 |
| M3D-LaMed | CT-Agent evaluation; micro-F1 | 0.148 | – | – | 0.245 | 0.400 | 0.326 |
| RadFM | CRG-paper evaluation of the pretrained checkpoint | 0.059 | 0.335 | 0.018 | 0.000 | 0.042 | 0.020 |
Four things stand out:
- The lead is clear against the most comparable systems and smaller against the rest. CT2Rep, CT-CHAT, BTB3D and RadFM, scored with the same labeler and averaging on the historical 3,039 set, reach at most 0.258. The next rows use other or unstated protocols: CT-AGRG (0.501) doesn’t state its averaging, RMR (0.495) and Astra (0.484) score the 1,564 studies, and Astra’s value was read off a chart.
- It leads on every finding. Its report F1 is higher than the best of seven other systems on all 18 findings (Figure 3), though two of those systems’ values were digitised from figures and the cohorts and labelers differ.
- The report is nearly as good as asking. Macro-F1 is 0.614 from 18 yes/no questions, 0.610 from one prompt listing all findings, 0.601 from the report’s own finding list and 0.592 from RadBERT reading its prose: a spread of 0.022. The report adds location, severity and findings beyond the 18 labels.
- Best on findings, near the bottom on wording. BLEU-4 is 0.11 and ROUGE-L 0.21, below CT2Rep (0.21 and 0.36), a 2024 model with F1 0.160 (Figure 4). The authors’ explanation is format: structured reports scored against free-narrative references, which Merlin supports (section 06). GREEN, an LLM judge of clinical errors, also puts it below CT2Rep (0.41 against 0.49), though its checkpoint is out of domain for CT.
Show as table
| Finding | Report | Yes/no | Best other report | By | Positives |
|---|---|---|---|---|---|
| Pleural effusion | 0.847 | 0.859 | 0.724 | COLIPRI-CRM* | 370 |
| Arterial wall calcification | 0.766 | 0.775 | 0.700 | MonteRET | 854 |
| Coronary artery wall calcification | 0.756 | 0.777 | 0.642 | COLIPRI-CRM* | 758 |
| Lung opacity | 0.734 | 0.745 | 0.680 | COLIPRI-CRM* | 1,177 |
| Lung nodule | 0.687 | 0.659 | 0.605 | VoxelFM* | 1,351 |
| Consolidation | 0.682 | 0.712 | 0.569 | VoxelFM* | 578 |
| Medical material | 0.679 | 0.727 | 0.414 | COLIPRI-CRM* | 307 |
| Atelectasis | 0.597 | 0.604 | 0.489 | COLIPRI-CRM* | 704 |
| Cardiomegaly | 0.585 | 0.639 | 0.508 | MonteRET | 316 |
| Pericardial effusion | 0.561 | 0.578 | 0.380 | MonteRET | 217 |
| Lymphadenopathy | 0.545 | 0.575 | 0.451 | COLIPRI-CRM* | 773 |
| Hiatal hernia | 0.531 | 0.541 | 0.480 | MonteRET | 417 |
| Emphysema | 0.521 | 0.527 | 0.464 | VoxelFM* | 591 |
| Pulmonary fibrotic sequela | 0.491 | 0.541 | 0.453 | COLIPRI-CRM* | 822 |
| Bronchiectasis | 0.461 | 0.468 | 0.334 | COLIPRI-CRM* | 327 |
| Mosaic attenuation pattern | 0.457 | 0.452 | 0.425 | VoxelFM* | 246 |
| Interlobular septal thickening | 0.392 | 0.429 | 0.353 | VoxelFM* | 246 |
| Peribronchial thickening | 0.370 | 0.441 | 0.273 | VoxelFM* | 348 |
06Beyond CT-RATE: Merlin and RAD-ChestCT
Merlin (abdominal CT), after format adaptation
Merlin’s reports use short organ-specific subsections, unlike NV-Reason-CT’s native format. So the authors first fine-tuned the model for one epoch on Merlin’s 15,314 released training reports, with a Findings-only prompt listing the 15 Merlin headings; no test report was used. The comparator numbers come from the Jolia paper, whose authors ran all four models on the same test split, possibly with different evaluator versions.
Show as table
| Method | BLEU | ROUGE-L | BERTScore | RadGraph-F1 | GREEN |
|---|---|---|---|---|---|
| NV-Reason-CT† | 0.077 | 0.336 | 0.579 | 0.322 | 0.366 |
| Jolia | 0.119 | 0.323 | 0.567 | 0.317 | 0.324 |
| Merlin | 0.063 | 0.285 | 0.551 | 0.237 | 0.316 |
| MedGemma 1.5 | 0.124 | 0.106 | 0.361 | 0.032 | 0.038 |
| Med3DVLM | 0.004 | 0.079 | 0.256 | 0.010 | 0.007 |
This is the closest thing to a control in the paper: once its reports match the reference format, the model with some of the lowest wording scores on CT-RATE comes out on top, which suggests much of that gap is format. Two caveats: an epoch on 15,314 in-domain reports teaches content too, and the paper leaves out the sixth metric in Jolia’s table, CRIMSON, on which Merlin scores best.
RAD-ChestCT (external chest CT), without tuning
RAD-ChestCT publishes labels but no report text, so reports are scored only through the 16 harmonised findings extracted from them. NV-Reason-CT ran with its default prompt and no RAD-ChestCT tuning.
Show as table
| Method | Macro-F1 | Precision | Recall |
|---|---|---|---|
| NV-Reason-CT · finding list | 0.510 | 0.516 | 0.552 |
| NV-Reason-CT · RadBERT | 0.485 | 0.510 | 0.494 |
| DCP-PD | 0.395 | 0.443 | 0.418 |
| EXACT-CHAT | 0.289 | 0.469 | 0.298 |
| BTB3D | 0.266 | 0.272 | 0.329 |
| CT-CHAT | 0.182 | 0.382 | 0.171 |
| M3D-LaMed | 0.113 | 0.269 | 0.080 |
| RadFM | 0.069 | 0.283 | 0.044 |
07The reader study behind “50% faster”
Two US board-certified radiologists, each with at least 10 years of practice, read 10 cases: 5 chest and 5 abdomen, 2 normal and 8 abnormal. They reported each case unaided first, then, after a short distraction, reviewed it with the model’s reasoning and structured report, which they could copy and edit. Reported time per case fell from 26.25 to 13.13 minutes. Nine of ten agreement statements averaged at least 4.60 on a 1–5 scale; the exception was the AI finding something they had missed (2.60), and the readers said their careful first read had missed nothing.
Treat it as a feasibility signal. It is 20 case-reads, and the assisted read always came second on the same case, so memory of the first read helps; the authors flag these recall and order effects. The times are reported by the readers, and the final reports weren’t scored for accuracy. The paper calls the study preliminary and not an assessment of clinical safety.
08Reading the metrics
- CRG, re-expressed. CRG weights true positives, false negatives and false positives by prevalence. It works out to CRG = 1/(3 − 2J), with J = TPR − FPR over the pooled label slots, so chance is 1/3. NV-Reason-CT’s 0.514 means J ≈ 0.53 derived, against about 0.37 for the frozen nnFoundation stack, on a different evaluation set.
- GREEN is out of domain. Its official checkpoint was trained mainly on chest X-ray reports, so read its CT scores loosely. The authors say so.
- RadBERT reads at most 512 tokens. Long structured reports get truncated (Figure 6); the paper doesn’t give the CT-RATE count.
09What the paper leaves out
- Ablations: none for the unmerged tokens, the expert reasoning data or GRPO, so their separate contributions are unknown. The authors list this as a limitation.
- Data: the number of expert narrations, the model that wrote the synthetic narrations and labels, per-family example counts and a breakdown of the 70,111 inputs. The release adds no training data beyond 128 illustrative CT-RATE examples, and no evaluation code.
- Training budget: SFT epochs or steps, GRPO steps and wall-clock time. Only the hardware is given.
- Uncertainty and a fair baseline: no confidence intervals or significance tests for its report metrics, and no strong generator re-run on the 3,002-scan set with the same labeler.
- Faithfulness: hallucination rates, and whether the
<think>text reflects what drives the answer, aren’t measured. The authors say the reasoning “is not assumed to reveal the model’s internal computation”.
10Takeaways for building a CT report generator
- Test whether you can afford to keep the tokens. The best clinical numbers here come with 13,824 unmerged tokens, 3D positions and end-to-end tuning. Without an ablation, whether merging costs accuracy is still your experiment to run.
- Reward what you measure, knowingly. GRPO on the finding list ties training to label-based evaluation, and a structure reward keeps the prose organised.
- Emit a machine-readable finding list. It gives a labeler-free F1 (0.601 against 0.592) and sidesteps RadBERT’s 512-token limit.
- Report text metrics, but don’t rank on them. After one epoch of format adaptation, the model near the bottom on CT-RATE wording led four of five text metrics on Merlin.
- Mask what wasn’t imaged. Coverage-aware labels stop partial scans from teaching the model false negatives.
ASources
- Myronenko, A., Yang, D., et al. NV-Reason-CT: 3D Visual Language Model for CT Analysis. arXiv:2609.27511 (2026).
- NVIDIA-Medtech. NV-Reason-CT code · model card.
- NVIDIA. Reasoning Visual Language Model for Chest X-Ray Analysis (NV-Reason-CXR). arXiv:2510.23968 (2025).
- Yu, Q., et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv:2503.14476 (2025).
- Qwen Team. Qwen3.5-4B.
- Hamamci, I. E., et al. CT-RATE: dataset; CRG score (2025).
- Our previous post: nnFoundation, read for report generation.