* Our arithmetic on the paper’s ablation table: 0.024 of the 0.032 gained. The paper’s own average over seven scores gives the same share.
- A small captioner on a small benchmark. A three-layer Transformer decoder, trained from scratch on 796 LLM-written summaries of TCGA breast pathology reports, is scored mainly by word overlap on 93 test cases. No pathologist read its output.
- First in every column, without error bars. BLEU-4 is 0.135 against 0.118 for HistGen. All seven baselines are evidently the authors’ own runs, and no intervals, repeat runs or tests are reported.
- Most of the gain over a plain Transformer is not the retrieved reports. With no retrieval, the model that pools the slide into one learned vector already has 0.024 of the 0.032 BLEU-4 gained over a plain Transformer. Retrieval on its own lowers the score; the full knowledge branch adds 0.008.
- What the scores can’t see. The paper’s own example gets tumour size, grade and stage wrong and still earns a BLEU-4 close to the test set’s. The HER2 lead looks like five cases out of 35 over HistGen, and two over the best baseline.
01Why this paper matters for report generation
The blog’s first pathology post read a paper that scored 50 generated reports with a blinded pathologist and no automatic metric. This one brings the opposite evidence: a public benchmark, automatic metrics, seven baselines, and no reader.
Its idea is retrieval: as a pathologist recalls similar cases, BiGen looks up sentences from past reports for the patches it attends to most. The code and the reviews (scores of 3, 4 and 5; no rebuttal round) are public, and this post uses both.
(≤ 10,000 in the release)
M × 512 (PLIP)
Three things in it are easy to miss:
- No retrieved word reaches the decoder. A lookup returns the mean of three sentence embeddings, and all lookups are pooled into one vector.
- There is no pretrained language model. The decoder is trained from scratch: about 7.9M parameters in its three blocks, counted from the release derived, against the 347M-parameter BioGPT of the first pathology post.
- Retrieval isn’t trained. PLIP is frozen on both sides; the model learns only which patches ask. In the released code the bank is the same during training, so a training slide can pull sentences from its own report our observation.
The mechanism, condensed from the released code (helper names are ours):
def retrieve(plip, attn, bank, ratio=0.4, m=10, top=3):
# plip: M x 512 patches · bank: T x 512 sentences
order = attn.argsort(descending=True) # first pass
kept = plip[order][: int(len(order) * ratio)]
groups = pad(kept, m).view(-1, m, 512).mean(1)
sim = cosine(groups[:, None], bank[None])
nearest = sim.topk(top, dim=1).indices
return bank[nearest].mean(1) # groups x 512
v1, attn = layer(v0, patches) # visual token, pass 1
k = proj(retrieve(plip, attn, bank))
v2, _ = layer(v1, patches); t1, _ = layer(t0, k)
v3, _ = layer(v2, patches); t2, _ = layer(t1, k)
memory = cat([v1 + v2 + v3, t1 + t2]) # 2 x 512
report = decoder(memory) # beam search, width 3
02The benchmark: PathText (BRCA)
PathText comes from the MI-Gen paper (MICCAI 2024): 9,009 slide–text pairs built from TCGA. The text is not the pathologist’s report. The scanned report goes through OCR, an LLM is asked to summarise it, and a classifier trained on 88 hand-labelled pairs removes flawed summaries.
The breast subset holds 1,041 slides of 977 patients, by our count on the released split file. In that split, 12 of the 98 test patients also have a slide in the training set, and the benchmark’s loader pairs both with the same report. BiGen keeps one slide per patient, which leaves 796 training, 88 validation and 93 test cases, so a test patient’s report is in neither the training set nor, by the paper’s account, the knowledge bank. That is a real fix, and no number can be carried over from MI-Gen’s paper.
The reference summarises the report of a whole surgical case, tumour size and lymph nodes included, while the input is one slide; MI-Gen’s authors themselves suggest that a two-dimensional slide can’t give a three-dimensional size.
03Results: where the gain comes from
BiGen is first in all ten columns of the paper’s main table, seven text scores and three for HER2: BLEU-4 is 0.135 against 0.118 for HistGen, and 0.293 ROUGE-L against 0.279 for the next best. The ablation table starts from the same row, a plain Transformer (the paper’s vanilla Transformer), so both tables fit on one axis as gains over it.
Show as table
| Table 1, text scores | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | METEOR | ROUGE | Fact-ent |
|---|---|---|---|---|---|---|---|
| CNN-RNN | 0.371 | 0.185 | 0.089 | 0.043 | 0.143 | 0.239 | 0.478 |
| att-LSTM | 0.371 | 0.191 | 0.094 | 0.048 | 0.142 | 0.238 | 0.460 |
| Plain Transformer | 0.389 | 0.246 | 0.157 | 0.103 | 0.158 | 0.257 | 0.487 |
| R2Gen | 0.378 | 0.243 | 0.160 | 0.107 | 0.179 | 0.279 | 0.505 |
| R2GenCMN | 0.396 | 0.254 | 0.164 | 0.110 | 0.163 | 0.279 | 0.470 |
| MI-Gen | 0.416 | 0.267 | 0.174 | 0.115 | 0.165 | 0.270 | 0.490 |
| HistGen | 0.422 | 0.272 | 0.177 | 0.118 | 0.169 | 0.277 | 0.493 |
| BiGen | 0.450 | 0.296 | 0.196 | 0.135 | 0.180 | 0.293 | 0.521 |
| Table 2, ablation | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | METEOR | ROUGE | Fact-ent | AVG. △ |
|---|---|---|---|---|---|---|---|---|
| Plain Transformer | 0.389 | 0.246 | 0.157 | 0.103 | 0.158 | 0.257 | 0.487 | – |
| + visual token (VTCA) | 0.436 | 0.283 | 0.186 | 0.126 | 0.174 | 0.287 | 0.499 | 10.84% |
| + shared layers (WSL) | 0.437 | 0.286 | 0.190 | 0.127 | 0.174 | 0.280 | 0.507 | 11.41% |
| + retrieval (KR) | 0.416 | 0.269 | 0.176 | 0.119 | 0.169 | 0.283 | 0.505 | 7.75% |
| + textual token (TTCA) | 0.441 | 0.288 | 0.190 | 0.130 | 0.178 | 0.290 | 0.515 | 13.06% |
| + shared branches (WS) | 0.450 | 0.296 | 0.196 | 0.135 | 0.180 | 0.293 | 0.521 | 15.26% |
Three things stand out:
- One vector for the slide does most of it. Replacing self-attention over all patches by a single learned query lifts BLEU-4 from 0.103 to 0.126. With no retrieval at all, that model already tops HistGen and MI-Gen on all seven text scores. A decoder fed one vector outscores the same decoder attending to every patch: with 796 training reports the bottleneck may regularise, or the scores may reward phrasing more than detail our observation.
- Retrieval alone hurts. Adding the retrieved vectors through self-attention, with no textual token, takes BLEU-4 from 0.127 to 0.119. Pooling them into a second token instead gives 0.130; letting both tokens share weights gives 0.135. The whole knowledge branch adds 0.008, a quarter of the total gain and about half of the lead over HistGen derived. The paper’s example may hint at why: four of the nine retrieved sentences it shows describe other patients’ immunostains, receptor status, a BRCA2 mutation or a right axillary node.
- One setting moves the score as much as the lead. Retrieving for every patch instead of the top 40% gives 0.119, about HistGen’s level. Seven sentences per lookup instead of three gives 0.137, above the headline.
All seven baselines must be the authors’ own runs: the test set is new, the MI-Gen row matches none of the three in MI-Gen’s paper, and HistGen’s paper uses another dataset. The paper names UNI as the visual extractor without saying whether the baselines got it; a reviewer asked for that control, and there was no rebuttal round. The released code fixes one seed.
04What the scores can see
The paper prints one generated report beside its reference and highlights the words they share. Attribute by attribute, the pair reads differently:
| Attribute | Reference | Generated |
|---|---|---|
| Side | left breast | left breast |
| Diagnosis | poorly differentiated invasive ductal carcinoma, ICD-O M-8500/3 | the same, with the same code |
| Tumour size | up to 6 cm | 1.4 cm |
| Grade | G3 | grade 2 |
| Tumour stage | pT3 | pT2 |
| Lymph nodes | pN2a | none involved (pN0) |
Table 1 — The paper’s example, attribute by attribute. Our reading of the two texts in the paper’s Fig. 2; bold marks a disagreement.
Scored alone with the released scorer’s BLEU formula, this pair of about 100 tokens each gets a BLEU-4 of 0.127 (0.102 if full stops aren’t counted as tokens) derived, where the whole test set gets 0.135. All of its four-word matches come from the diagnosis phrase and from “there is no”.
The second headline is HER2: an F1 of 0.730 against 0.571 for HistGen, read off the generated text. All eight values in the precision column are whole thirty-fifths, while no single denominator up to the 93 test cases fits the recall column. On a fixed test set it should be the other way round. The paper doesn’t define the columns. The simplest reading: 35 reference reports state a HER2 result, and BiGen’s report matches it in 23, HistGen’s in 18 and a 2015 LSTM captioner’s in 21 (the abstract’s 19.1% matches the gain over that LSTM).
Even if BiGen is right wherever HistGen is, five extra cases out of 35 give p = 0.06 in a two-sided exact McNemar test derived. And HER2 status is set by immunostaining and in-situ hybridisation, not by the H&E slide the model sees: a generated HER2 sentence is an educated guess.
05The same models, re-run by others
Two 2026 preprints from one group, which proposes its own model, SCOUT, each re-train BiGen, HistGen and MI-Gen under one protocol. Neither reports intervals.
Show as table
| Source | Method | Features | BLEU-4 | METEOR | ROUGE-L |
|---|---|---|---|---|---|
| MI-Gen’s own paper, breast | MI-Gen | ResNet, ImageNet | 0.117 | 0.163 | 0.280 |
| MI-Gen’s own paper, breast | MI-Gen | ViT, ImageNet | 0.110 | 0.157 | 0.279 |
| MI-Gen’s own paper, breast | MI-Gen | ViT, HIPT | 0.120 | 0.171 | 0.271 |
| BiGen’s paper, breast | MI-Gen | not stated | 0.115 | 0.165 | 0.270 |
| BiGen’s paper, breast | HistGen | not stated | 0.118 | 0.169 | 0.277 |
| BiGen’s paper, breast | BiGen | UNI | 0.135 | 0.180 | 0.293 |
| SCOUT preprint, breast | MI-Gen | CONCHv1.5 | 0.1036 | 0.1521 | 0.2843 |
| SCOUT preprint, breast | HistGen | CONCHv1.5 | 0.1060 | 0.1749 | 0.2670 |
| SCOUT preprint, breast | BiGen | CONCHv1.5 | 0.1394 | 0.1581 | 0.2890 |
| PathReportEval preprint, across cancer types | MI-Gen | CONCHv1.5 | 0.1071 | 0.1481 | 0.2863 |
| PathReportEval preprint, across cancer types | MI-Gen | H-Optimus-1 | 0.1177 | 0.1542 | 0.2907 |
| PathReportEval preprint, across cancer types | MI-Gen | UNI2-h | 0.1198 | 0.1563 | 0.2881 |
| PathReportEval preprint, across cancer types | MI-Gen | mean of the three | 0.1149 | 0.1529 | 0.2884 |
| PathReportEval preprint, across cancer types | HistGen | CONCHv1.5 | 0.0971 | 0.1423 | 0.2749 |
| PathReportEval preprint, across cancer types | HistGen | H-Optimus-1 | 0.1040 | 0.1475 | 0.2840 |
| PathReportEval preprint, across cancer types | HistGen | UNI2-h | 0.1128 | 0.1545 | 0.2931 |
| PathReportEval preprint, across cancer types | HistGen | mean of the three | 0.1046 | 0.1481 | 0.2840 |
| PathReportEval preprint, across cancer types | BiGen | CONCHv1.5 | 0.1070 | 0.1472 | 0.2966 |
| PathReportEval preprint, across cancer types | BiGen | H-Optimus-1 | 0.1198 | 0.1583 | 0.2984 |
| PathReportEval preprint, across cancer types | BiGen | UNI2-h | 0.1214 | 0.1549 | 0.2986 |
| PathReportEval preprint, across cancer types | BiGen | mean of the three | 0.1161 | 0.1535 | 0.2979 |
On the breast subset, with other features and another split, BiGen’s BLEU-4 and its lead reproduce: 0.139 against 0.106 for HistGen. On PathText across cancer types, with about nine times the training data, BiGen and MI-Gen are 0.001 apart, and 0.0021 at most with any one encoder.
06What the paper leaves out
- A floor: no row copies the nearest training report or prints one fixed report, in this paper or in the benchmark’s, so what a template would score is unknown.
- Scoring code: the release computes BLEU, METEOR and ROUGE-L. Factent, an entity-overlap score borrowed from radiology, and the HER2 columns aren’t in it; two requests have been open since December 2025.
- The bank’s recipe: the bank is shared as a file, without the scripts that build it and extract PLIP features (a third open request). That it holds training reports only rests on the paper’s word.
- Paper against release: the paper sets groups of 20 patches and the release defaults to 10; the paper motivates groups as neighbouring tissue, and the release groups patches by attention rank.
- Readers and other organs: no pathologist read a generated report, and only breast cancer is tested. A reviewer pointed at the single dataset and at findings absent from the bank.
07Takeaways for building a pathology report generator
- With under a thousand reports, try a bottleneck first. One learned query over the patches gave +0.023 BLEU-4; everything added after it gave +0.009.
- Ablate retrieval against the same model without it. Against a plain Transformer BiGen gains 0.032; against its own no-retrieval variant, 0.008. Keep each case’s own report out of the bank while training.
- Score what a slide can show. Build references from slide-level findings, as the melanocytic paper did by cleaning its reports. Whole-case summaries reward a model for guessing size and stage.
- Report intervals, and counts for small subsets. On 93 test cases a 0.017 lead is untested, and the HER2 gap to HistGen looks like five cases.
ASources
- Zhang, L., Yun, B., Li, Q., Wang, Y. Historical Report Guided Bi-modal Concurrent Learning for Pathology Report Generation. MICCAI 2025, LNCS 15965, pp. 343–352. The page holds the reviews. PDF · doi · arXiv:2506.18658 (one version, same tables).
- Code: DeepMed-Lab-ECNU/BiGen (MIT), read on 5 Oct 2026; open requests 2, 3 and 4.
- Chen, P., et al. WsiCaption, the MI-Gen and PathText paper. MICCAI 2024; arXiv:2311.16480. Code and split file: cpystan/Wsi-Caption.
- Guo, Z., et al. HistGen. MICCAI 2024; arXiv:2403.05396.
- Singh, S., et al. SCOUT: arXiv:2605.01144 (2026). Singh, S., et al. PathReportEval: arXiv:2607.18448 (2026).
- Chen, R. J., et al. UNI. Nature Medicine 30 (2024); repository note on its pretraining data. Huang, Z., et al. PLIP. Nature Medicine 29 (2023); repository.
- Wolff, A. C., et al. HER2 testing in breast cancer: ASCO–CAP guideline update. J Clin Oncol (2023).
- Miura, Y., et al. Factent: arXiv:2010.10042 (NAACL 2021).
- Our post on melanocytic lesion reports.