* Our reading of the sources the rows trace to, and our measurement of the bars in Extended Data Fig. 1.
- The checkpoint decides the score. On the same 912 MIMIC-CXR test images, the pretrained MedGemma 4B in Table 4 scores 29.5 RadGraph F1 and the instruction-tuned 4B, the most downloaded checkpoint, scores 21.9, per Google’s model card. The 30.3 comes from a 4B fine-tuned on MIMIC-CXR that is not released.
- The comparison rows were measured elsewhere. Only one of the thirteen copied rows, Med-Gemini-2D, is stated to use the same 912 test images. Two don’t match their source: M²Transformer’s 22.0 is its BLEU score, and CXR-RePaiR writes impressions, not findings.
- The state-of-the-art claim rests on a 0.3-point margin. The tuned model’s 30.3 tops MedVersa’s 30.0, measured on another test set, and sits 0.8 above the pretrained checkpoint, from a single evaluation run with no interval. The journal calls the tuning RL; the technical report gives the same 30.3 as supervised fine-tuning.
- The clinical read is one radiologist who knew which report was the AI’s. Over 306 cases, 81% of generated reports would have led to the same or better management. Among the 207 abnormal studies, the AI report missed key findings in about one in five.
01What the paper offers for report generation
MedGemma is a family of Gemma 3 models tuned on medical text and images: a 4B and a 27B that read images, and a 27B for text alone. Report generation is one row of the paper’s evaluation table, and its evidence sits in four places: Table 4, one paragraph on reinforcement learning, a reader study in Extended Data Fig. 1, and a paragraph quoting a public leaderboard. All of it is chest X-ray, and all of it but the leaderboard is MIMIC-CXR.
to 896 × 896
(64 × 64 patches of 14 px)
one per 56 × 56 px
the original report
Two details in it shape the results:
- MIMIC-CXR in every stage. Its 231,483 examples are listed for encoder tuning, decoder pretraining and the reinforcement-learning stage of post-training. RadGraph F1 on MIMIC-CXR tuning data was one of the scores used to set the pretraining data mix, and the pretraining checkpoint was picked partly on chest X-ray report generation.
- The main scored model is the pretrained one. The technical report explains why: RadGraph F1 is sensitive to reporting style, and the instruction-tuned model writes more like Gemma 3 than like MIMIC-CXR.
02Which MedGemma writes the report
Three checkpoints of MedGemma 4B write reports, and on the same images they score far apart.
Show as table
| Checkpoint | RadGraph F1 | Where the number is |
|---|---|---|
| MedGemma 4B + RL for CXR | 30.3 | Table 4 · not released |
| MedGemma 4B pretrained | 29.5 | Table 4 · released (-pt) |
| MedGemma 1.5 4B | 27.2 | model card · released Jan 2026 |
| Med-Gemini-2D | 24.4 | Table 4 · earlier Google model |
| MedGemma 4B instruction-tuned | 21.9 | model card · released (-it) |
| Gemma 3 4B instruction-tuned | 12 | model card · the base model |
The pretrained checkpoint, released as medgemma-4b-pt, scores 29.5, the number Google also reports for PaliGemma 2 10B fine-tuned for this one task. The instruction-tuned checkpoint, medgemma-4b-it, scores 21.9: 7.6 points lower, which the model card puts down to reporting style derived. It is the most downloaded checkpoint: Hugging Face counted 891,316 downloads of it over the last 30 days, against 1,452 for the pretrained checkpoint (6 October 2026). Gemma 3 4B, the model MedGemma starts from, scores 12 in its instruction-tuned version, and MedGemma 1.5, released in January, 27.2.
For a team that wants MIMIC-style reports without fine-tuning, the released pretrained checkpoint is the better choice; for any other style, the 29.5 says little about what either checkpoint will score.
03Table 4, row by row
Table 4 sets MedGemma’s two rows against thirteen copied from other papers; its caption warns that inclusion criteria differ. Ten of the thirteen already sit, with the same values, in Med-Gemini’s table of 2024, and seven of those in Flamingo-CXR’s of 2023, which says it took them from the original publications. That can’t hold for R2Gen and M²Transformer, which predate RadGraph: their numbers come from a 2023 study by Yu et al. that re-scored older models on the impressions, the findings, or both, of the MIMIC-CXR test set.
Show as table
| Row (sections in Table 4) | Table 4 | Source of the number | What the source scored |
|---|---|---|---|
| CXR-RePaiR (F) | 9.1 | Flamingo-CXR Table 1 (0.091); Yu et al. Table 1 and X-REM print 0.090 | impressions · 2,191 |
| M²Transformer (F) | 22.0 | Yu et al. 2023, Table 2 | findings · 1,597 · RadGraph F1 24.4, BLEU 22.0 |
| Med-PaLM M 12B (F) | 25.2 | Med-PaLM M Table A.7 | findings · 4,834 · + indication |
| Med-PaLM M 84B (F) | 26.7 | Med-PaLM M Table A.7 | findings · 4,834 · + indication |
| MAIRA-1 (F) | 24.3 | MAIRA-1 Table 5 | findings · 2,461 · + indication |
| MAIRA-2 (F) | 34.6 | MAIRA-2 Fig. 3B | findings · 2,461 · + lateral view, prior study, indication |
| R2Gen (F + I) | 13.4 | Yu et al. 2023, Table 3 | findings + impression · 2,192 |
| WCT (F + I) | 14.3 | Yu et al. 2023, Table 3 (“WCL”) | findings + impression · 2,192 |
| CvT2DistilGPT2 (F + I) | 15.4 | Yu et al. 2023, Table 3 | findings + impression · 2,192 |
| Flamingo-CXR (F + I) | 20.5 | Flamingo-CXR Table 1 | findings + impression · 1,931 |
| Med-Gemini-2D (F + I) | 24.4 | Med-Gemini Table 8 | findings + impression · 912 |
| PaliGemma 2 10B (F + I) | 29.5 | PaliGemma 2 Table 8 | findings + impression · Flamingo-CXR’s split · + indication; tuned for this task |
| MedVersa (F + I) | 30.0 | MedVersa Table 1, row “All” | findings and impression, joined · 1,437 |
Two rows don’t survive the trip our observation:
- M²Transformer, 22.0. In Yu et al.’s findings table, 0.220 is its BLEU score; its RadGraph F1 is 0.244. The X-REM paper prints the same pair.
- CXR-RePaiR, 9.1, “F only”. CXR-RePaiR retrieves impressions: its code extracts the impression section, and both Yu et al. and X-REM score it on impressions (0.090).
The other rows match their sources, but not each other’s terms, as the second lines of Figure 3 show. MAIRA-2’s 34.6, the highest in the table, is findings only, from a model that also reads the lateral view and the prior study.
The caption adds that writing both sections is generally considered harder. On this metric it doesn’t look harder: in MedVersa’s table the three models trained on MIMIC-CXR score higher on the joined text than on findings alone, by 1.7 to 2.4 points derived.
The article cites the RadGraph dataset paper and names no scorer. Scorers differ by more than the gaps between Table 4’s top rows: on the same 2,461 MAIRA-1 reports, Yu et al.’s RadGraph F1 gives 24.3 and RG_ER, from the radgraph package and my RadGraph-F1 paper, gives 29.6. Med-Gemini, from the same group, used Yu et al.’s package on lower-cased text; MedGemma likely did too, not confirmed.
# Yu et al. 2023: exact entity and relation matches
04The RL row
Table 4 adds a row that the technical report’s Table 10 lacks. In the journal’s account, the authors took the instruction-tuned 4B, fine-tuned it with reinforcement learning on MIMIC-CXR, gave it the image and the indication, and trained it to write findings and impression. It scores 30.3, which the paper calls the state of the art “for this specific metric and dataset”.
The main text and Methods don’t give the reward, the number of training reports or steps, or a second metric, and the model is in none of Google’s Hugging Face listings. The technical report gives the same 30.3, from the same starting checkpoint and inputs, as supervised fine-tuning for one epoch, and says this task “strictly required the usage of SFT instead of RL”. The model card says only “tuned for CXR”, and the repository’s RL notebook trains on MedQA multiple-choice answers. If the journal’s reward is RadGraph F1 or close to it, the model is scored on what it was trained to raise.
The margin over the pretrained checkpoint is 0.8, from one inference run per report. For scale, MAIRA-1’s bootstrap interval on 2,461 reports is ±0.55; on 912 images the same spread would be about ±0.9 our arithmetic. A paired test of the two checkpoints could be tighter, and the paper reports none. The margin over MedVersa’s 30.0 is 0.3, on a different test set.
05The reader study
The one clinical check is in Extended Data Fig. 1. A US board-certified thoracic radiologist compared the pretrained model’s report with the original one for 306 MIMIC-CXR test images, 99 normal and 207 abnormal, on the original DICOMs, and rated each pair on a six-level rubric about patient management. The radiologist knew which report was the AI’s.
Show as table
| Category (rubric) | Normal · n = 99 | Abnormal · n = 207 | All · n = 306 |
|---|---|---|---|
| AI superior, original missed key findings | 2% | 3% | 2.5% |
| Both correct, AI more detailed | 6% | 8% | 7.5% |
| Similar findings | 59% | 38% | 44% |
| Both correct, original more detailed | 21% | 29% | 27% |
| Original superior, AI missed key findings | 10% | 21% | 17.5% |
| Both missed key findings | 2% | 1% | 1.5% |
The headline holds: 81% of reports would have led to the same or better management, the first four rubric levels of the bar for all cases. The split shows where the other 19% sits. In abnormal studies, the original report caught key findings the AI report missed in 21% of cases, and both missed them in 1% our measurement. In normal studies the first figure is 10%.
The paper sets its 81% against 73% for Med-Gemini, on the same 306 cases per the technical report. Med-Gemini’s own report prints 72%, and its design differs: five readers kept after two were dropped for low agreement, reports shown in random order with their origin masked, and ratings where neither report was correct left out of the percentages. One reader who knew the source, against five who didn’t, is not a head-to-head.
06The leaderboard, then and now
The paper also quotes ReXrank, a public leaderboard that scores submitted models on four datasets. As of 9 February 2026, MedGemma 4B had the highest GREEN on ReXGradient (0.566) and IU-Xray (0.724) and sat mid-table on MIMIC-CXR and CheXpert Plus (the paper writes “MIMIC-III”). The leaderboard’s history on GitHub lets us check that, and see what changed since.
| Dataset | 9 Feb rank, 1/RadCliQ-v1 | 6 Oct rank, CRIMSON | GREEN, both dates |
|---|---|---|---|
| ReXGradient | 7 | 22 | 0.566 (highest) |
| MIMIC-CXR | 16 | 14 | 0.293 |
| IU-Xray | 8 | 25 | 0.724 (highest) |
| CheXpert Plus | 15 | 13 | 0.246 |
Table 1 — Same reports, a new ruler. MedGemma 4B’s rank among 30 models in ReXrank’s overview, on the page the paper accessed (version of 6 Feb 2026) and on 6 Oct 2026, by the metric each version ranked on, with its GREEN score in the findings view. Source: rajpurkarlab/ReXrank, gh-pages history; the paper’s Results.
The paper’s account checks out for February, and the GREEN values haven’t moved: they are still the highest on those two datasets in the findings view. The ranking has. In February the leaderboard ranked by 1/RadCliQ-v1, and MedGemma was 7th of 30 on ReXGradient and 8th on IU-Xray. Since 7 August it ranks by CRIMSON, an LLM-judged metric posted in March, and MedGemma is 22nd and 25th; by the old metric it would still be 7th and 9th today derived. CRIMSON’s default judge is itself a fine-tuned MedGemma 4B; the leaderboard doesn’t say which judge produced its column. MedGemma 4B’s rows there link to the instruction-tuned checkpoint and cover findings only.
07What the paper leaves out
The article sends evaluation details to its Supplementary Note, which this post did not read; these are gaps in the article and its technical report.
- The test set: the article’s Table 1 lists 306 report-generation examples, the reader-study size. The 912 images behind Table 4 appear only in the technical reports.
- The scorer: no RadGraph F1 implementation is named.
- The RL recipe: reward, data, steps, a second metric and the weights; the technical report describes the same 30.3 as supervised fine-tuning.
- Uncertainty: one inference run per example and no interval on any report-generation number.
- A blinded, multi-reader study: one radiologist who knew the source, and no agreement statistic.
08Takeaways for building a report generator on MedGemma
- Pick the checkpoint for the job. For MIMIC-style reports the released pretrained 4B scores 29.5 where the instruction-tuned one scores 21.9. Plan to fine-tune either for your own house style.
- Re-score baselines on your own split. Two of the thirteen copied rows here don’t match their source, and only one is stated to share MedGemma’s test set.
- Name the scorer and the sections. The same reports score 24.3 or 29.6 with two RadGraph-based scorers, and joined sections can score higher than findings alone.
- Count the misses, with blinded readers. The one human read here found key findings missing in about a fifth of abnormal studies. That is the number to track.
- Read a leaderboard rank with its metric. The same MedGemma reports went from 7th to 22nd on ReXGradient when the ranking metric changed.
ASources
- Sellergren, A., Kazemzadeh, S., … Golden, D., Yang, L. An open vision-language model for diverse medical applications. Nature Medicine (2026), published online 6 Oct 2026. doi:10.1038/s41591-026-04626-w.
- MedGemma Technical Report: arXiv:2507.05201, v1 (7 Jul 2025) to v4 (6 Apr 2026); report generation in §3.5, §5.1, Tables 10 and 13 and Appendix A, the same in all four. MedGemma 1.5 Technical Report: arXiv:2604.05081.
- Model cards: medgemma-4b-it, medgemma-4b-pt, medgemma-1.5-4b-it; Health AI Developer Foundations. Code: google-health/medgemma, read on 6 Oct 2026. Gemma 3 report: arXiv:2503.19786.
- Sources of the copied rows: Yu, F. et al. Evaluating progress in automatic chest X-ray radiology report generation, Patterns (2023); X-REM, arXiv:2303.17579; Flamingo-CXR, arXiv:2311.18260; Med-Gemini, arXiv:2405.03162; PaliGemma 2, arXiv:2412.03555; MedVersa, arXiv:2405.07988; MAIRA-1, arXiv:2311.13668; MAIRA-2, arXiv:2406.04449; Med-PaLM M, arXiv:2307.14334; CvT2DistilGPT2, arXiv:2201.09405; CXR-RePaiR code.
- ReXrank: rexrank.ai and its repository (gh-pages, commits of 6 Feb and 3 Oct 2026). CRIMSON: arXiv:2603.06183 and its README, which names the default judge.
- Delbrouck, J.-B. et al. RadGraph-F1 and the RadGraph rewards: arXiv:2210.12186 (2022). Ostmeier, S. et al. GREEN: arXiv:2405.03595 (2024).
- Our posts on chest X-ray reports: SPEC-CXR, RadAlign, Diff-RRG and phrase-grounded fact-checking.