* Our arithmetic from the paper’s shapes: 5 × (4 × 8 × 8).
- Three passes; only two see the images. A 9B language model fine-tuned on five 3D sequences writes findings and conclusion. The same design, fine-tuned separately, reasons to a FIGO stage in six tagged steps, and a text-only 8B model checks that stage against the rules.
- The findings text leads by a little. The strongest baseline is the same language model fed 2D slices, 0.033 BLEU-2 behind. The largest step is pretraining the encoder on the dataset’s own reports.
- The stage is right for 64% of cases, 26% with the substage. The text-only check supplies 10 and 8 of those points. The paper gives no stage counts, so the score of a constant answer is unknown.
- What we can’t tell. There are no intervals or repeat runs, and three of the ten staging accuracies cannot be whole counts out of 105. The data come on request only, and the released code covers the first pass.
01Why this paper matters for report generation
This is the first MRI paper in this series. Three things change for readers who build CT or chest X-ray report generators:
- Several sequences per exam. A CT report generator usually reads one volume. A pelvic MRI exam is several, each with its own contrast and often its own plane. Cervical-RG takes five: ADC, T1CA, T1CS, T2A and T2S (by their names, which the paper doesn’t expand: a diffusion map, contrast-enhanced T1 in two planes and T2 in two planes).
- A stage the report must end on. Treatment follows the FIGO stage, a rule-based label built from tumour size, spread and lymph nodes. Each exam has one clinical answer that is right or wrong.
- A single disease. Every patient has cervical cancer, so there is no list of findings to score against, as CheXpert’s 14 labels or CT-RATE’s 18 allow.
= 1,280
The sequences meet only in the language model. Each passes alone through the shared encoder and arrives as its own block of 256 tokens, the budget the design-space study gave a whole CT scan. The released prompt builders never name the sequences: five image tags follow each other.
02The data
Cervical-MD is the authors’ own dataset: 3,177 patients from seven hospitals, each with five sequences, a report and the FIGO stage that clinicians recorded. The abstract counts 3,137 MRI–report pairs, and the paper doesn’t explain the difference. The test set is 15 random cases per hospital, 105 in all. Every other case trains, and the paper mentions no validation split. The released scripts key each record by hospital and patient folder, so the split appears to be by patient.
The paper prints its example in English and never says whether the reports were written in English or translated. Every author’s institution is in China and the released code keeps a Chinese prompt in a comment, so translation is likely, not confirmed.
The data are not released. The paper page lists no dataset, and both the rebuttal and the repository’s issue tracker say to apply by email.
03The findings text
The paper compares Cervical-RG with six models, all run by the authors: two zero-shot and four fine-tuned on Cervical-MD, on inputs it doesn’t specify. A second table ablates the inputs and the training stages. Two rows appear in both, so everything fits on one axis.
Show as table, with all six metrics
| Setting | BLEU-2 | ROUGE-1 | METEOR | BERTSc. | RadGraph | RadCliQ ↓ |
|---|---|---|---|---|---|---|
| LLaVA-Med, zero-shot | 0.021 | 0.158 | 0.079 | 0.812 | 0.012 | 2.335 |
| GPT-4o, zero-shot | 0.029 | 0.151 | 0.102 | 0.804 | 0.021 | 2.391 |
| R2GenGPT† | 0.167 | 0.358 | 0.247 | 0.841 | 0.149 | 1.768 |
| M3D-LaMed† | 0.191 | 0.432 | 0.299 | 0.867 | 0.218 | 1.427 |
| miniGPT-Med† | 0.231 | 0.492 | 0.330 | 0.880 | 0.261 | 1.221 |
| 2D · T2A only | 0.244 | 0.497 | 0.350 | 0.881 | 0.272 | 1.197 |
| 2D · ADC, T1CS, T2A | 0.252 | 0.504 | 0.353 | 0.881 | 0.277 | 1.179 |
| 2D · all five; same scores as LongLLaVA-Med† | 0.253 | 0.504 | 0.358 | 0.882 | 0.283 | 1.157 |
| 3D · all five, neither retrained | 0.211 | 0.454 | 0.330 | 0.868 | 0.207 | 1.446 |
| 3D · T2A only | 0.260 | 0.507 | 0.359 | 0.882 | 0.222 | 1.196 |
| 3D · ADC, T1CS, T2A | 0.259 | 0.504 | 0.364 | 0.882 | 0.277 | 1.157 |
| 3D · all five, no projector stage | 0.270 | 0.520 | 0.376 | 0.885 | 0.295 | 1.077 |
| 3D · all five, full: Cervical-RG | 0.286 | 0.536 | 0.394 | 0.888 | 0.297 | 1.068 |
- The strongest baseline is the paper’s own 2D arm. LongLLaVA-Med† in Table 1 and the 2D row with all sequences in Table 3 share all six numbers. It is evidently the same run: Cervical-RG’s language model with its stock 2D encoder.
- Pretraining the encoder on the reports is the largest step. Without it, and without the projector stage, the 3D model scores 0.211 BLEU-2, below every 2D row. The pretraining alone takes it to 0.270. The paper doesn’t say whether the 105 test cases were kept out of it.
- Not every row agrees that 3D beats 2D, or that more sequences help. With T2A alone, 3D has a RadGraph F1 lower by 0.050 (0.222 against 0.272, as printed), a swing between neighbouring rows that exceeds the 0.014 lead in Figure 2’s title. And from one 3D sequence to three, BLEU-2 goes from 0.260 to 0.259.
- The metrics see little of the clinic. BERTScore gives a zero-shot GPT-4o 0.804 and the best model 0.888. The two clinical metrics come from chest X-ray work: RadGraph was annotated on chest X-ray reports, and RadCliQ was fitted to radiologists’ error counts on reports drawn from 50 MIMIC-CXR studies. The paper names no implementation and offers no check that either transfers to pelvic MRI.
The paper prints one report written by Cervical-RG. It finds the cervical tumour and gives it the reference’s 12 mm. It also describes two nodules in the uterine wall as probable fibroids, where the reference reports a myometrium without lesions, and it leaves out the reference’s two cysts.
The figure colours the phrases that agree. Both texts walk the pelvis organ by organ in the same order, so part of the overlap is the template. The paper reports no reader score for the findings text.
04The stage
Staging is scored by accuracy against the recorded stage: C-Acc over the four major stages and F-Acc over 19 classes that the paper doesn’t list. The label has its own limit: under FIGO 2018, pathology overrules imaging and positive lymph nodes alone make a case stage IIIC, so it can rest on information that MRI doesn’t hold, as the authors acknowledge.
Show as table, with counts and the doctors’ scores
| Configuration | Major stage | Correct of 105 | Full stage | Correct of 105 | Doctors, /10 |
|---|---|---|---|---|---|
| UNETR classifier | 0.283 0.21–0.38 | none | 0.124 0.07–0.20 | 13 | – |
| Swin-UNETR classifier | 0.333 0.25–0.43 | 35 | 0.133 0.08–0.21 | 14 | – |
| Passes 1 + 3 | 0.352 0.27–0.45 | 37 | 0.124 0.07–0.20 | 13 | – |
| Passes 1 + 2 | 0.543 0.45–0.64 | 57 | 0.181 0.12–0.27 | 19 | 4.474 |
| Passes 1 + 2 + 3 | 0.642 0.55–0.73 | none | 0.264 0.19–0.36 | none | 6.320 |
| Same, GPT-4o in pass 3 | 0.609 0.51–0.70 | 64if truncated | 0.209 0.14–0.30 | 22if truncated | – |
Under each accuracy, our 95% interval. “Correct of 105” is the whole count that prints as that accuracy; “none” means no count does. The last row is the rebuttal’s.
- The chain. A text-only LLM that stages pass 1’s report scores 0.352 and 0.124, about what two image classifiers get. The fine-tuned staging pass reaches 0.543 and 0.181, and the rule check adds 0.099 and 0.083.
- What the ablation doesn’t isolate. No row has the fine-tuned model state a stage without the tagged steps. Pass 2’s gain mixes fine-tuning on stage labels, a second look at the images and the chain of thought.
- No yardstick. The paper gives no stage counts, so the score of always answering the most common stage is unknown. It is at least 0.257 (27 of 105, with four stages), and our intervals for both classifiers include that value.
- About ±9 points. That is our 95% interval on 105 cases (Figure 3). The paper gives none, and no paired test for the rule check’s ten points.
- Half the errors are major. 36% of cases get the wrong major stage, about half of the 74% with a wrong full stage derived. The rebuttal describes errors as mostly between neighbouring substages; the paper prints no confusion matrix.
Seven of the ten accuracies in Table 2 (five configurations, two levels) are whole counts out of 105. Three are not, whether rounded or truncated: UNETR’s 0.283 and the two headline numbers, 0.642 and 0.264. All three would be whole counts out of 53, or of 106.
def whole_counts(p, n=105):
return [k for k in range(n + 1)
if round(k / n, 3) == p]
whole_counts(0.543) # [57]
whole_counts(0.642) # []
whole_counts(0.642, n=53) # [34]
whole_counts(0.642, n=106) # [68]
Typos, a 106th test case or a test set that changed would each explain it. The rebuttal says the test set was being expanded; the paper says 105.
Two doctors also scored the reports out of 10 on a staging rubric: 4.474 without the rule check and 6.320 with it. The paper gives no agreement between the two doctors, and no baseline was scored.
05What the reviews and the release add
The three reviewers scored the submission 1, 3 and 4. The lowest score came with a short review: implementation details were missing (how the sequences are combined, how image and text are aligned), the dataset description was unclear, and the submission gave too little information to reproduce it. The rebuttal described the projector, the six tags and a few dataset facts, and that reviewer moved to accept.
The other two worried about pass 3, which was GPT-4o in the submitted version: one saw a privacy issue in uploading generated reports to a proprietary model, and both a method that relied on it. The authors called it a marginal fallback, worth 2.8 points of full-stage accuracy, and replaceable. The final paper uses a locally deployable 8B model instead, and the headline row moves from the rebuttal’s 60.9% and 20.9% to 64.2% and 26.4%. Pass 3 now adds 8.3 points, 31% of the final full-stage accuracy derived. The reviewer who scored 3 stayed at reject, doubting that such changes fit a camera-ready version.
Three things the rebuttal promised are not in the final paper: full expert scores, results on other institutions and detailed distributions, of which only the ages appear.
The code came on 5 December 2025 and covers pass 1: a fine-tuning script, a generation script, data-preparation scripts and a link to a checkpoint. It has no data, no encoder pretraining, no prompts or staging rules for passes 2 and 3, no metric code and no licence file. Its generation script samples at temperature 1.0 with no fixed seed, and the paper states no decoding settings; tables made that way would be one draw each.
All four released prompt builders end the prompt with a sentence saying whether the patient has metastatic lymph nodes, read from a spreadsheet column. The paper mentions no such input and the prompt in its Figure 1 has no such sentence, so we can’t tell which prompt produced the tables. Under FIGO 2018, node status is part of the stage.
06What the paper leaves out
- Radiologists’ accuracy: how often radiologists reach the recorded stage from the same images, which a reviewer asked for.
- A baseline without images: in a one-disease cohort it says how much of each text score the template earns.
- The decoder: every variant uses the hybrid language model, fully fine-tuned. There is no Transformer-only, frozen or LoRA variant and no memory or speed figure, although efficiency is the stated reason for the hybrid.
07Takeaways for building an MRI report generator
- Pretrain the encoder on your own reports first. It moved BLEU-2 from 0.211 to 0.270, more than 3D over 2D (0.033) or five sequences over three (0.027). Keep the test cases out, and say so.
- When a report ends on a rule-based label, check the label with the rules. A text-only 8B model re-reading the output added ten points of major-stage accuracy.
- Publish the class balance with any staging accuracy, and the constant-answer score beside it. 64% can’t be read without them.
- In a one-disease cohort, add a no-image baseline and a reader for the findings. One 2D sequence already earns 85% of the best BLEU-2 derived.
ASources
- Zhang, H., Long, Y., Fan, Y., Wang, Y., Zhan, Z., Wang, S., Jiang, Y., Sun, R., Xing, Z., Li, Z., Duan, X., Zhao, W. Cervical-RG: Automated Cervical Cancer Report Generation from 3D Multi-sequence MRI via CoT-guided Hierarchical Experts. MICCAI 2025, LNCS 15964, pp. 78–88. The page holds the reviews and the rebuttal. PDF · doi. No arXiv version was found on 5 Oct 2026.
- Code: LongYu-LY/Cervical-RG (no licence file), read on 5 Oct 2026.
- Bai, F., et al. M3D: arXiv:2404.00578 (2024); code BAAI-DCAI/M3D.
- Wang, X., et al. LongLLaVA: arXiv:2409.02889 (2024); LongLLaVAMed-9B model card.
- DeepSeek-AI. DeepSeek-R1-Distill-Llama-8B model card (2025).
- FIGO 2018 staging: National Cancer Institute, Cervical Cancer Treatment (PDQ), staging tables adapted from the FIGO Committee for Gynecologic Oncology.
- Jain, S., et al. RadGraph: arXiv:2106.14463 (2021). Yu, F., et al. RadCliQ: Patterns 4(9), 100802 (2023). Delbrouck, J.-B., et al. RadGraph-F1: arXiv:2210.12186 (2022).