- A reader study, not a metric. One pathologist, whose initials match a co-author, read 50 test cases blind, the generated report next to the original: 25 common nevi and 25 other lesions. The reports get no automatic metric and no baseline model.
- On par on common nevi, within what 25 cases can show. 4.5 against 4.6, with about 0.4 factual errors per report on both sides. No test is reported; our rough interval on the gap is ±0.4 points.
- Far behind on the rest, but nothing unverifiable. On other lesions the score is 3.0 against 4.5, with 2.5 factual errors and 2.9 repeated phrases per report. None of the 50 generated reports states anything the slides can’t show, which a sibling study traces to cleaning the training reports.
- What we can’t tell. Whether the model beats a canned nevus paragraph, and whether a second reader would agree: in the sibling study another pathologist gave original reports from the same archive 3.9, not 4.6.
01Why this paper matters for report generation
Slide-level vision–language models in pathology have so far mostly written a sentence or two, as the authors point out. PRISM, the contrastive-captioner (CoCa) design this paper follows, was trained on GPT-4 rewrites of each report as one sentence of under 20 words. This paper asks the same kind of model for a full-length report, the microscopic description and the conclusion: 77 words on average for a common nevus (an ordinary mole) and 125 for other lesions. It then scores the result with a blinded pathologist instead of text-overlap metrics.
This is the blog’s first pathology post, and three things differ from the CT posts:
- The image. A case is a few glass slides scanned at 20× or 40×, far too large for an encoder. Each slide is cut into 224-pixel tiles about 0.1 mm across derived, and a frozen tile encoder turns every tile into a vector. The largest cases, fewer than 2%, exceed 100,000 tiles, about 5 gigapixels derived.
- The unit. The report covers a case, not a slide: 2.2 slides on average here derived, with all their tiles pooled.
- The report. A description of what the microscope shows, then a conclusion with the diagnosis. Reports also hold results the H&E slide can’t show, such as immunostains and molecular tests, and some conclusions rest on them.
N ≤ 100,000 in training
512 for the report,
1 for the contrastive loss
The data are a tenth of PRISM’s derived, 19,645 cases against 195,344 specimens, and come from one corner of dermatopathology.
02The data: cleaned reports, two populations
The data come from the pathology archive of the University Medical Center Utrecht: melanocytic lesions accessioned from 2013 to 2020, 42,512 H&E slides for 19,645 cases from 14,978 patients, split 80/10/10 by patient into training, validation and test sets. The dataset is not released.
The reports were written in Dutch. One fine-tuned language model translated them into English; a second cut them into sub-sentences and labelled each by content. Two kinds were kept: observations made on the H&E slides, and conclusions, which the authors judged too important to withhold. Everything else went out, immunostain and molecular results included, as did every case whose report had no H&E description. Radiology readers will recognise the move: it is the counterpart of stripping references to prior exams from chest X-ray reports. The pipeline has its own error rate: the companion paper counted 219 translation errors in 3,597 test sentences (6.1%).
| Common nevi | Other lesions | |
|---|---|---|
| Share of the 19,645 cases | 81.8% | 18.2% |
| Mean report length after cleaning | 77 words, 7 sentences | 125 words, 12 sentences |
| Cases in the test set | 1,592 | 377 |
| Read in the reader study | 25 (1.6%) | 25 (6.6%) |
Table 1 — Two populations in one dataset. The reader study samples the two groups equally; the test set doesn’t hold them equally. The 18.2% and the two sampling fractions are our arithmetic derived. Sources: Sections 2.1 and 3.1, and the paper’s Table 2.
The other lesions, from non-common nevi to melanocytomas and melanomas, share about 2,900 training cases derived.
03The reader study: one pathologist, 50 cases
A pathologist with dermatopathology experience, identified by initials that match a co-author, read 50 test cases drawn at random within the two groups. For each case the reader had the slides and two reports in random order, the generated one and the original, without knowing which was which. Each report got four counts (factual errors, unverifiable statements, omissions, repeated phrases) and a score from 1, mostly inaccurate, to 5, usable with little or no editing.
Show as table
| Reports | Score, 1–5 | Factual errors | Omissions | Repeated phrases | Unverifiable |
|---|---|---|---|---|---|
| All 50 cases · generated | 3.7 ± 1.2 | 1.4 ± 2.1 | 0.7 ± 1.0 | 1.6 ± 2.9 | 0.0 ± 0.0 |
| All 50 cases · original | 4.6 ± 0.6 | 0.5 ± 0.8 | 0.2 ± 0.5 | 0.3 ± 1.1 | – |
| Common nevi · generated | 4.5 ± 0.8 | 0.4 ± 0.6 | 0.3 ± 0.5 | 0.2 ± 0.7 | 0.0 ± 0.0 |
| Common nevi · original | 4.6 ± 0.6 | 0.4 ± 0.7 | 0.1 ± 0.3 | 0.0 ± 0.0 | – |
| Other lesions · generated | 3.0 ± 1.1 | 2.5 ± 2.5 | 1.1 ± 1.1 | 2.9 ± 3.6 | 0.0 ± 0.0 |
| Other lesions · original | 4.5 ± 0.7 | 0.5 ± 0.9 | 0.3 ± 0.6 | 0.6 ± 1.6 | – |
The paper’s “All” row weighs the two groups equally: 3.7 against 4.6. At the test set’s own mix, 81% common nevi, the model’s mean would be about 4.2 and the originals’ still 4.6 derived. Four things stand out:
- On common nevi, no difference in score is visible. 4.5 ± 0.8 against 4.6 ± 0.6. The two columns score the same 25 cases, but the paper gives no paired statistics; treated as independent samples, the 0.1 gap carries a 95% interval of roughly ±0.4 points derived, so “on par” means within about half a point.
- On other lesions the model is 1.5 points lower. It makes five times the factual errors of the original reports and omits more. In the paper’s melanoma example it gets the melanoma right and the subtype wrong.
- Nothing unverifiable, and that is the cleaning. None of the 50 generated reports holds an unverifiable statement; for a sample of 50, the rule of three still allows a true rate of about 6% of reports derived. The ablation is in a sibling study by the same group: under the same protocol, its model produced 3.0 unverifiable statements per report when trained on uncleaned reports, and none on cleaned ones.
- It loops, probably because of decoding. 2.9 repeated phrases per report on other lesions; in the paper’s example the dermal component is announced four times. The paper suggests report length and rare subtypes as causes, and doesn’t describe decoding. The released evaluation code passes no repetition penalty (the README’s inference example sets 1.1), and the documented command sets top-k to 1, which is greedy. The sibling study’s model, released with beam search and a penalty of 1.2, averaged 0.1 repeated phrases against 1.6 here likely, not confirmed.
04The same protocol, run twice
That sibling study is the closest thing to a baseline. Posted on arXiv the same day, it applies the same protocol to the same archive, with a different pathologist as reader, to judge by the initials. Its model is BLIP-2-style, uses older tile features and writes the description without the conclusion.
Show as table
| Score, 1–5 | Factual errors | Repeated phrases | ||||
|---|---|---|---|---|---|---|
| Study | Generated | Original | Gap | Generated | Original | Generated |
| All 50 cases | ||||||
| BLIP-2 model, reader A | 2.8 ± 1.4 | 3.9 ± 1.0 | −1.1 | 1.7 ± 1.7 | 0.7 ± 1.0 | 0.1 ± 0.5 |
| This paper, reader B | 3.7 ± 1.2 | 4.6 ± 0.6 | −0.9 | 1.4 ± 2.1 | 0.5 ± 0.8 | 1.6 ± 2.9 |
| Common nevi, 25 cases | ||||||
| BLIP-2 model, reader A | 3.7 ± 1.2 | 4.2 ± 0.9 | −0.5 | 0.9 ± 0.9 | 0.5 ± 0.7 | 0.1 ± 0.3 |
| This paper, reader B | 4.5 ± 0.8 | 4.6 ± 0.6 | −0.1 | 0.4 ± 0.6 | 0.4 ± 0.7 | 0.2 ± 0.7 |
| Other lesions, 25 cases | ||||||
| BLIP-2 model, reader A | 1.9 ± 1.0 | 3.6 ± 1.0 | −1.7 | 2.5 ± 1.9 | 0.9 ± 1.3 | 0.1 ± 0.6 |
| This paper, reader B | 3.0 ± 1.1 | 4.5 ± 0.7 | −1.5 | 2.5 ± 2.5 | 0.5 ± 0.9 | 2.9 ± 3.6 |
Two things follow. On other lesions little has changed: the gap went from 1.7 to 1.5 points, and both models make 2.5 factual errors per report. And the original reports are no gold standard. Across each study’s 50 cases they get 4.6 from one pathologist and 3.9 from the other, and the readers find 0.5 and 0.7 factual errors in them per report, which the sibling paper reads as disagreement between observers. The originals the paper shows are in English, so the reader presumably scored them after translation our observation. A model’s 4.5 has no meaning outside its own study.
05Why routine is easy and rare is hard
The paper’s second experiment explains the first. The model also places slides and reports in one embedding space, and the authors test whether a case’s own report can be found among the 1,969 test reports. For common nevi it is in the top 10 for 31% of cases; for other lesions, for 65%. The model writes best what it retrieves worst.
The authors’ reading: common-nevus reports are short and alike, so many fit a given case. Easy to write acceptably, hard to tell apart. That is good news for workload and a warning for evaluation, because a canned nevus paragraph might score well too.
In the released code the aggregator receives the tile vectors and a padding mask: no coordinates, no slide index. Yet the generated nevus report in the paper’s example calls the lesion symmetrical and describes how its cells change with depth. An unordered set of 0.1 mm tiles can’t measure either directly, so such sentences rest on indirect cues, such as tissue context that hints at a tile’s depth, or on habit.
Habit is right for most common nevi. Where symmetry and maturation help separate benign from malignant, it could produce the kind of factual error the reader counted. The paper doesn’t break errors down by type, so this stays a hypothesis.
06What the paper leaves out
- Baselines: no other model, and no template or nearest-report baseline. Two reviewers asked, naming PRISM, HistGen and MI-Gen. The authors replied that models trained on other hospitals’ reports would be at an unfair disadvantage, and left retraining them to future work.
- Automatic metrics: none, although the released code can compute BLEU-4, ROUGE-L, CIDEr and METEOR at validation. The other 1,919 test cases are never scored for generation.
- Readers and statistics: one reader, whose initials match a co-author from the department whose reports trained the model; the authors name the single reader and the 50 cases as limits. No test, and no breakdown of the 25 other lesions by diagnosis.
- Variants and remedies: BioGPT’s first 12 blocks are kept frozen, tuned with LoRA or fully fine-tuned, and the three are compared on retrieval only (frozen and LoRA within each other’s intervals, full fine-tuning worse). The paper doesn’t say which wrote the reports. The rebuttal names oversampling and loss weights for the rare lesions; neither is tried.
- Access and compute: the dataset is private. The code is Apache-2.0, and the three checkpoints sit behind an access request on Hugging Face. The paper omits the compute; the rebuttal gives 16 GPUs of 24 GB for about two days.
07Takeaways for building a report generator
- Clean the targets before tuning the model. Zero unverifiable statements in 50 reports after the training text was cut down to H&E observations and conclusions; 3.0 per report without the cut, in the sibling study.
- Report by case mix. 4.5 on four cases in five and 3.0 on the rest say more than any average of the two.
- Put the human reports through the same protocol. They scored 4.6 with one reader and 3.9 with another. Without that row, a model’s 4.5 can’t be read.
- Give the routine class a template baseline. With the right report in the top 10 of 1,969 for only 31% of common nevi, many reports seem to fit the same case, and a canned paragraph is the bar to clear.
- Check decoding before blaming the model. 2.9 repeated phrases per report on the harder cases is the most visible flaw here, and the cheapest to test.
ASources
- Lucassen, R. T., Moonemans, S. P. J., van de Luijtgaarden, T., Breimer, G. E., Blokx, W. A. M., Veta, M. Pathology Report Generation and Multimodal Representation Learning for Cutaneous Melanocytic Lesions. MICCAI 2025, LNCS 15965, pp. 502–511. The page holds the reviews and the rebuttal. PDF · doi · arXiv:2502.19293 (same tables).
- Code: SanderMoon/MOSAIC (Apache-2.0), read on 5 Oct 2026. Checkpoints: SaltySander/MOSAIC on Hugging Face (gated).
- The sibling study: Lucassen, R. T., van de Luijtgaarden, T., et al. On the Importance of Text Preprocessing for Multimodal Representation Learning and Pathology Report Generation. arXiv:2502.19285 (2025), Table 3. Code: nuldertien/PathBLIP-2.
- The report cleaning: Lucassen, R. T., et al. Preprocessing Pathology Reports for Vision-Language Model Development. MICCAI Workshop on Computational Pathology, PMLR 254:61–71 (2024).
- Shaikovski, G., et al. PRISM: arXiv:2405.10254 (2024). Yu, J., et al. CoCa: arXiv:2205.01917 (2022).
- Chen, R. J., et al. UNI: Nature Medicine 30, 850–862 (2024). Luo, R., et al. BioGPT: arXiv:2210.10341 (2022).
- Ramesh, V., Chi, N. A., Rajpurkar, P. Removing references to priors from radiology reports: arXiv:2210.06340 (2022).