Jean-Benoit Delbrouck / Research blog

Research blog · Pathology report generation

BiGen, read for pathology report generation

BiGen, a MICCAI 2025 model, writes a breast pathology report from two vectors, one for the slide and one for sentences retrieved from past reports, and tops seven baselines on a public benchmark. This post reads how much the retrieved reports explain, and what the scores can see.

Paper Ling Zhang, Boxiang Yun, Qingli Li, Yan Wang. Historical Report Guided Bi-modal Concurrent Learning for Pathology Report Generation. MICCAI 2025, LNCS 15965, pp. 343–352 By Jean-Benoit Delbrouck Published 6 Oct 2026 Reading time ~8 min Tags MICCAI 2025 · Main
93
test reports, one slide per patient; 796 reports to train on
2
vectors are all the decoder sees of a slide and its retrieved text
0.135
BLEU-4, against 0.118 for the best of seven baselines; no intervals
75%*
of the BLEU-4 gained over a plain Transformer comes before any report is retrieved

* Our arithmetic on the paper’s ablation table: 0.024 of the 0.032 gained. The paper’s own average over seven scores gives the same share.

TL;DR
  • A small captioner on a small benchmark. A three-layer Transformer decoder, trained from scratch on 796 LLM-written summaries of TCGA breast pathology reports, is scored mainly by word overlap on 93 test cases. No pathologist read its output.
  • First in every column, without error bars. BLEU-4 is 0.135 against 0.118 for HistGen. All seven baselines are evidently the authors’ own runs, and no intervals, repeat runs or tests are reported.
  • Most of the gain over a plain Transformer is not the retrieved reports. With no retrieval, the model that pools the slide into one learned vector already has 0.024 of the 0.032 BLEU-4 gained over a plain Transformer. Retrieval on its own lowers the score; the full knowledge branch adds 0.008.
  • What the scores can’t see. The paper’s own example gets tumour size, grade and stage wrong and still earns a BLEU-4 close to the test set’s. The HER2 lead looks like five cases out of 35 over HistGen, and two over the best baseline.

01Why this paper matters for report generation

The blog’s first pathology post read a paper that scored 50 generated reports with a blinded pathologist and no automatic metric. This one brings the opposite evidence: a public benchmark, automatic metrics, seven baselines, and no reader.

Its idea is retrieval: as a pathologist recalls similar cases, BiGen looks up sentences from past reports for the patches it attends to most. The code and the reviews (scores of 3, 4 and 5; no rebuttal round) are public, and this post uses both.

Figure 1 — The decoder sees a slide as two vectors. With 40% of the patches in groups of 20, there is one retrieved vector per 50 patches derived. The encoders’ output sizes and the patch cap come from the released code. Sources: Sections 2.2–2.4 and 3.1, Fig. 1.

Three things in it are easy to miss:

  • No retrieved word reaches the decoder. A lookup returns the mean of three sentence embeddings, and all lookups are pooled into one vector.
  • There is no pretrained language model. The decoder is trained from scratch: about 7.9M parameters in its three blocks, counted from the release derived, against the 347M-parameter BioGPT of the first pathology post.
  • Retrieval isn’t trained. PLIP is frozen on both sides; the model learns only which patches ask. In the released code the bank is the same during training, so a training slide can pull sentences from its own report our observation.

The mechanism, condensed from the released code (helper names are ours):

def retrieve(plip, attn, bank, ratio=0.4, m=10, top=3):
    # plip: M x 512 patches · bank: T x 512 sentences
    order = attn.argsort(descending=True)  # first pass
    kept = plip[order][: int(len(order) * ratio)]
    groups = pad(kept, m).view(-1, m, 512).mean(1)
    sim = cosine(groups[:, None], bank[None])
    nearest = sim.topk(top, dim=1).indices
    return bank[nearest].mean(1)           # groups x 512

v1, attn = layer(v0, patches)      # visual token, pass 1
k = proj(retrieve(plip, attn, bank))
v2, _ = layer(v1, patches); t1, _ = layer(t0, k)
v3, _ = layer(v2, patches); t2, _ = layer(t1, k)
memory = cat([v1 + v2 + v3, t1 + t2])      # 2 x 512
report = decoder(memory)           # beam search, width 3

02The benchmark: PathText (BRCA)

PathText comes from the MI-Gen paper (MICCAI 2024): 9,009 slide–text pairs built from TCGA. The text is not the pathologist’s report. The scanned report goes through OCR, an LLM is asked to summarise it, and a classifier trained on 88 hand-labelled pairs removes flawed summaries.

The breast subset holds 1,041 slides of 977 patients, by our count on the released split file. In that split, 12 of the 98 test patients also have a slide in the training set, and the benchmark’s loader pairs both with the same report. BiGen keeps one slide per patient, which leaves 796 training, 88 validation and 93 test cases, so a test patient’s report is in neither the training set nor, by the paper’s account, the knowledge bank. That is a real fix, and no number can be carried over from MI-Gen’s paper.

The reference summarises the report of a whole surgical case, tumour size and lymph nodes included, while the input is one slide; MI-Gen’s authors themselves suggest that a two-dimensional slide can’t give a three-dimensional size.

03Results: where the gain comes from

BiGen is first in all ten columns of the paper’s main table, seven text scores and three for HER2: BLEU-4 is 0.135 against 0.118 for HistGen, and 0.293 ROUGE-L against 0.279 for the next best. The ablation table starts from the same row, a plain Transformer (the paper’s vanilla Transformer), so both tables fit on one axis as gains over it.

Other methodsBiGen without retrievalBiGen with retrievalBLEU-4Gain0+0.01+0.02+0.03BLEU-4 gained over the plain Transformer (0.103)Other methodsre-run by the authorsR2Gen0.107+0.004R2GenCMN0.110+0.007MI-Gen0.115+0.012HistGen0.118+0.015BiGen, one part at a timethe paper’s ablationOne learned token reads the slide0.126+0.023+ one set of weights for all passes0.127+0.024+ retrieved report vectors, self-attended0.119+0.016+ a second token pools them instead0.130+0.027+ both tokens share weights = BiGen0.135+0.032HistGen’s level
Figure 2 — One token for the slide brings +0.023 BLEU-4; the knowledge branch then adds 0.008. The dashed line marks HistGen, the best baseline; the two LSTM baselines (−0.060 and −0.055) are in the table only, and Table 1’s HER2 columns are in Figure 3. 93 test reports; no intervals or repeat runs are reported. Sources: Tables 1 and 2.
Show as table
Table 1, text scoresBLEU-1BLEU-2BLEU-3BLEU-4METEORROUGEFact-ent
CNN-RNN0.3710.1850.0890.0430.1430.2390.478
att-LSTM0.3710.1910.0940.0480.1420.2380.460
Plain Transformer0.3890.2460.1570.1030.1580.2570.487
R2Gen0.3780.2430.1600.1070.1790.2790.505
R2GenCMN0.3960.2540.1640.1100.1630.2790.470
MI-Gen0.4160.2670.1740.1150.1650.2700.490
HistGen0.4220.2720.1770.1180.1690.2770.493
BiGen0.4500.2960.1960.1350.1800.2930.521
Table 2, ablationBLEU-1BLEU-2BLEU-3BLEU-4METEORROUGEFact-entAVG. △
Plain Transformer0.3890.2460.1570.1030.1580.2570.487–
+ visual token (VTCA)0.4360.2830.1860.1260.1740.2870.49910.84%
+ shared layers (WSL)0.4370.2860.1900.1270.1740.2800.50711.41%
+ retrieval (KR)0.4160.2690.1760.1190.1690.2830.5057.75%
+ textual token (TTCA)0.4410.2880.1900.1300.1780.2900.51513.06%
+ shared branches (WS)0.4500.2960.1960.1350.1800.2930.52115.26%

Three things stand out:

  1. One vector for the slide does most of it. Replacing self-attention over all patches by a single learned query lifts BLEU-4 from 0.103 to 0.126. With no retrieval at all, that model already tops HistGen and MI-Gen on all seven text scores. A decoder fed one vector outscores the same decoder attending to every patch: with 796 training reports the bottleneck may regularise, or the scores may reward phrasing more than detail our observation.
  2. Retrieval alone hurts. Adding the retrieved vectors through self-attention, with no textual token, takes BLEU-4 from 0.127 to 0.119. Pooling them into a second token instead gives 0.130; letting both tokens share weights gives 0.135. The whole knowledge branch adds 0.008, a quarter of the total gain and about half of the lead over HistGen derived. The paper’s example may hint at why: four of the nine retrieved sentences it shows describe other patients’ immunostains, receptor status, a BRCA2 mutation or a right axillary node.
  3. One setting moves the score as much as the lead. Retrieving for every patch instead of the top 40% gives 0.119, about HistGen’s level. Seven sentences per lookup instead of three gives 0.137, above the headline.

All seven baselines must be the authors’ own runs: the test set is new, the MI-Gen row matches none of the three in MI-Gen’s paper, and HistGen’s paper uses another dataset. The paper names UNI as the visual extractor without saying whether the baselines got it; a reviewer asked for that control, and there was no rebuttal round. The released code fixes one seed.

04What the scores can see

The paper prints one generated report beside its reference and highlights the words they share. Attribute by attribute, the pair reads differently:

AttributeReferenceGenerated
Sideleft breastleft breast
Diagnosispoorly differentiated invasive ductal carcinoma, ICD-O M-8500/3the same, with the same code
Tumour sizeup to 6 cm1.4 cm
GradeG3grade 2
Tumour stagepT3pT2
Lymph nodespN2anone involved (pN0)

Table 1 — The paper’s example, attribute by attribute. Our reading of the two texts in the paper’s Fig. 2; bold marks a disagreement.

Scored alone with the released scorer’s BLEU formula, this pair of about 100 tokens each gets a BLEU-4 of 0.127 (0.102 if full stops aren’t counted as tokens) derived, where the whole test set gets 0.135. All of its four-word matches come from the diagnosis phrase and from “there is no”.

Precision in thirty-fifths our observation

The second headline is HER2: an F1 of 0.730 against 0.571 for HistGen, read off the generated text. All eight values in the precision column are whole thirty-fifths, while no single denominator up to the 93 test cases fits the recall column. On a fixed test set it should be the other way round. The paper doesn’t define the columns. The simplest reading: 35 reference reports state a HER2 result, and BiGen’s report matches it in 23, HistGen’s in 18 and a 2015 LSTM captioner’s in 21 (the abstract’s 19.1% matches the gain over that LSTM).

Even if BiGen is right wherever HistGen is, five extra cases out of 35 give p = 0.06 in a two-sided exact McNemar test derived. And HER2 status is set by immunostaining and in-situ hybridisation, not by the H&E slide the model sees: a generated HER2 sentence is an educated guess.

“Precision”as printed“Recall”as printedF1as printed0714212835our reading: generated reports with the right HER2 result, of 35BiGen2323/35 = 0.65723/28 = 0.8210.730att-LSTM2121/35 = 0.60021/33 = 0.6360.618HistGen1818/35 = 0.51418/28 = 0.6430.571R2Gen1717/35 = 0.48617/27 = 0.6300.548Plain Transformer1616/35 = 0.45716/20 = 0.8000.582CNN-RNN1515/35 = 0.42915/20 = 0.7500.546†R2GenCMN1515/35 = 0.42915/22 = 0.6820.526MI-Gen99/35 = 0.2579/17 = 0.5290.346
Figure 3 — The HER2 scores read as small counts: 23, 21 and 18 right out of 35. Each bar is the count that reproduces the printed precision; the columns show the fractions that give the printed values to three decimals. † The counts give 0.545. Our reconstruction, not confirmed by the authors. Source: Table 1.

05The same models, re-run by others

Two 2026 preprints from one group, which proposes its own model, SCOUT, each re-train BiGen, HistGen and MI-Gen under one protocol. Neither reports intervals.

00.040.080.120.16BLEU-4MI-Gen’s own paper, 2024breast · 98 test slides · released split: 12 of their patients also in trainingMI-Gen0.120BiGen’s paper, 2025breast · 93 test patientsMI-Gen0.115HistGen0.118BiGen0.135SCOUT preprint, 2026breast · 90 test cases · CONCHv1.5 featuresMI-Gen0.104HistGen0.106BiGen0.139PathReportEval preprint, 2026PathText, across cancer types · 885 test cases · mean of three encodersMI-Gen0.115HistGen0.105BiGen0.116
Figure 4 — On breast cancer the lead reproduces; across cancer types BiGen and MI-Gen tie. BLEU-4 for the three methods these tables share, grouped by who trained them. MI-Gen’s own row uses its best features; both preprints call MI-Gen WSI-Caption. Sources: MI-Gen paper, Table 1; this paper, Table 1; SCOUT, Table 3; PathReportEval, Tables 2 and 4.
Show as table
SourceMethodFeaturesBLEU-4METEORROUGE-L
MI-Gen’s own paper, breastMI-GenResNet, ImageNet0.1170.1630.280
MI-Gen’s own paper, breastMI-GenViT, ImageNet0.1100.1570.279
MI-Gen’s own paper, breastMI-GenViT, HIPT0.1200.1710.271
BiGen’s paper, breastMI-Gennot stated0.1150.1650.270
BiGen’s paper, breastHistGennot stated0.1180.1690.277
BiGen’s paper, breastBiGenUNI0.1350.1800.293
SCOUT preprint, breastMI-GenCONCHv1.50.10360.15210.2843
SCOUT preprint, breastHistGenCONCHv1.50.10600.17490.2670
SCOUT preprint, breastBiGenCONCHv1.50.13940.15810.2890
PathReportEval preprint, across cancer typesMI-GenCONCHv1.50.10710.14810.2863
PathReportEval preprint, across cancer typesMI-GenH-Optimus-10.11770.15420.2907
PathReportEval preprint, across cancer typesMI-GenUNI2-h0.11980.15630.2881
PathReportEval preprint, across cancer typesMI-Genmean of the three0.11490.15290.2884
PathReportEval preprint, across cancer typesHistGenCONCHv1.50.09710.14230.2749
PathReportEval preprint, across cancer typesHistGenH-Optimus-10.10400.14750.2840
PathReportEval preprint, across cancer typesHistGenUNI2-h0.11280.15450.2931
PathReportEval preprint, across cancer typesHistGenmean of the three0.10460.14810.2840
PathReportEval preprint, across cancer typesBiGenCONCHv1.50.10700.14720.2966
PathReportEval preprint, across cancer typesBiGenH-Optimus-10.11980.15830.2984
PathReportEval preprint, across cancer typesBiGenUNI2-h0.12140.15490.2986
PathReportEval preprint, across cancer typesBiGenmean of the three0.11610.15350.2979

On the breast subset, with other features and another split, BiGen’s BLEU-4 and its lead reproduce: 0.139 against 0.106 for HistGen. On PathText across cancer types, with about nine times the training data, BiGen and MI-Gen are 0.001 apart, and 0.0021 at most with any one encoder.

06What the paper leaves out

  • A floor: no row copies the nearest training report or prints one fixed report, in this paper or in the benchmark’s, so what a template would score is unknown.
  • Scoring code: the release computes BLEU, METEOR and ROUGE-L. Factent, an entity-overlap score borrowed from radiology, and the HER2 columns aren’t in it; two requests have been open since December 2025.
  • The bank’s recipe: the bank is shared as a file, without the scripts that build it and extract PLIP features (a third open request). That it holds training reports only rests on the paper’s word.
  • Paper against release: the paper sets groups of 20 patches and the release defaults to 10; the paper motivates groups as neighbouring tissue, and the release groups patches by attention rank.
  • Readers and other organs: no pathologist read a generated report, and only breast cancer is tested. A reviewer pointed at the single dataset and at findings absent from the bank.

07Takeaways for building a pathology report generator

  1. With under a thousand reports, try a bottleneck first. One learned query over the patches gave +0.023 BLEU-4; everything added after it gave +0.009.
  2. Ablate retrieval against the same model without it. Against a plain Transformer BiGen gains 0.032; against its own no-retrieval variant, 0.008. Keep each case’s own report out of the bank while training.
  3. Score what a slide can show. Build references from slide-level findings, as the melanocytic paper did by cleaning its reports. Whole-case summaries reward a model for guessing size and stage.
  4. Report intervals, and counts for small subsets. On 93 test cases a 0.017 lead is untested, and the HER2 gap to HistGen looks like five cases.

ASources

  1. Zhang, L., Yun, B., Li, Q., Wang, Y. Historical Report Guided Bi-modal Concurrent Learning for Pathology Report Generation. MICCAI 2025, LNCS 15965, pp. 343–352. The page holds the reviews. PDF · doi · arXiv:2506.18658 (one version, same tables).
  2. Code: DeepMed-Lab-ECNU/BiGen (MIT), read on 5 Oct 2026; open requests 2, 3 and 4.
  3. Chen, P., et al. WsiCaption, the MI-Gen and PathText paper. MICCAI 2024; arXiv:2311.16480. Code and split file: cpystan/Wsi-Caption.
  4. Guo, Z., et al. HistGen. MICCAI 2024; arXiv:2403.05396.
  5. Singh, S., et al. SCOUT: arXiv:2605.01144 (2026). Singh, S., et al. PathReportEval: arXiv:2607.18448 (2026).
  6. Chen, R. J., et al. UNI. Nature Medicine 30 (2024); repository note on its pretraining data. Huang, Z., et al. PLIP. Nature Medicine 29 (2023); repository.
  7. Wolff, A. C., et al. HER2 testing in breast cancer: ASCO–CAP guideline update. J Clin Oncol (2023).
  8. Miura, Y., et al. Factent: arXiv:2010.10042 (NAACL 2021).
  9. Our post on melanocytic lesion reports.