What Happened
Biomedicine · Machine Learning · Benchmarking · Research BriefNature Biotechnology · April 2026
A survey of 220 biomedical foundation models
Biomedical Foundation Models Are Moving From One Modality to Many
A survey of 220 biomedicine-specific foundation models developed over four years. More than half — 53.6% — already combine two or more data modalities, and natural language is emerging as the connective interface between otherwise unrelated kinds of biomedical data. General-purpose large language models such as ChatGPT, Gemini and Claude were deliberately excluded.
- 220models curated, over four years
- 53.6%are multimodal — 118 combine two or more modalities
- 17distinct data modalities in total
- 4domains: language, imaging, omics, sequences
| Standing | What happened | Detail |
|---|---|---|
| Confirmed | The survey is published | Chang, Cheng, Modi, Wang, Xu and Ma, in Nature Biotechnology — online 30 April 2026, in print that August at 44(8):1263–1266. DOI 10.1038/s41587-026-03135-y. |
| Confirmed | 220 models, four fields, more than half multimodal | Every count and share on this page — the 220, the 118 (53.6%), the architecture split and the evaluation split — is stated in the paper's own running text and was read there. |
| Confirmed, not verifiable here | Which 220 models they are | The roster is Supplementary Table 1. The repository serves that file behind a challenge this project could not satisfy, so the list was never read model by model — the count is the paper's, not a recount. |
| Confirmed, not verifiable here | The citation counts | They come from Semantic Scholar and are printed only inside the paper's figure, which does not say when the snapshot was taken. They were read off that figure. Citation counts move, so treat them as one undated moment rather than a current ranking. |
What Counts as a Biomedical Foundation Model the survey's boundaries
- Large-scale machine learning models trained on varied biomedical and clinical data.
- That data spans genetic sequences, molecular profiling, biomedical imaging and electronic health records.
- They learn generalisable representations to support downstream discovery and clinical work.
- The survey covers four domains: natural language, imaging and signals, omics, and molecular sequences.
- General-purpose LLMs are excluded; the focus is 220 biomedicine-specific models.
Timeline
From BioBERT to Evo 2 Every model the survey names, dated to its own first publication
- 15 Feb 2020BioBERT — BERT re-trained on PubMed abstracts and PMC full text — is published; it becomes the survey's most-cited natural-language model.
- 13 Apr 2021ESM-1b shows that a language model trained on 250 million protein sequences, with no labels at all, learns real structural information; it is the ancestor of the survey's molecular-sequence leader.
- 26 Feb 2024scGPT, pretrained on more than 33 million individual cells, is published; it becomes the survey's omics leader.
- 19 Mar 2024UNI, pretrained on tissue images from over 100,000 whole-slide pathology scans, is published; it becomes the survey's imaging-and-signals leader.
- 16 Jan 2025ESM3 — a generative model reasoning jointly over protein sequence, structure and function — is published; the survey later names it a post-2025 riser.
- 4 Mar 2026Evo 2, trained on 9 trillion DNA base pairs spanning every domain of life, is published; the survey names it a second post-2025 riser.
- 30 Apr 2026Chang, Cheng, Modi, Wang, Xu and Ma publish the survey itself, online, in Nature Biotechnology.
- Aug 2026The survey appears in Nature Biotechnology's print issue, volume 44, number 8, pages 1263–1266.
The Argument
The key observation
Natural language is the modality most often paired with others, and is becoming the connective interface between data types. This runs through the whole paper: the emerging citation leaders and the underdeveloped directions both point back to the same structure.
Encoders Dominate architecture breakdown
| Architecture | Count | Share |
|---|---|---|
| Encoder-based | 114 | 51.8% |
| Mixed | 67 | 30.5% |
| Decoder-based | 31 | 14.1% |
| Non-Transformer (e.g. Mamba) | — | 3.6% |
Decoders are scarce not out of architectural preference but because of data: text-generation tasks need high-quality paired image–text annotation, and that is hard to obtain in biomedicine.
Evaluation Is Still Narrow downstream tasks used
| Task | Share |
|---|---|
| Classification | 50.9% |
| Report generation | 10.9% |
| Question answering | 8.6% |
| Segmentation | 7.3% |
The evaluation gap
There are no multitask benchmarks, little human-in-the-loop assessment, and no combined treatment of data quality, interpretability and translational relevance. Half of these models have been tested on classification alone.
Citation Leaders by Domain * marks a multimodal model · counts are a Semantic Scholar snapshot the paper does not date
| Domain | Leading models and citations |
|---|---|
| Natural language | BioBERT 6,925 · MultiMedQA 3,707 · PubMedBERT 2,278 |
| Imaging and signals | UNI 1,141 · ConVIRT* 1,003 · MedCLIP* 794 |
| Omics | scGPT 846 · Geneformer 836 · scBERT 509 |
| Molecular sequences | ESM-1b 2,807 · ProtTrans 1,207 · DNABERT 1,051 |
The Risers Since 2025 growth is coming from multimodal work
| Domain | Risers since 2025, and citations |
|---|---|
| Natural language | MedGemma-27B 165 · MediPhi-Instruct 13 |
| Imaging and signals | BiomedCLIP* 498 · Quilt-1M* 216 · MedGemma* 165 |
| Omics | Nicheformer* 81 · OmiCLIP* 58 · scGPT-spatial 34 |
| Molecular sequences | ESM3 201 · Evo 2 197 · Borzoi 190 |
Citation growth is being driven by models that integrate heterogeneous modalities rather than by specialists in a single data type. Five of the eleven risers are multimodal, and natural language is the only field with fewer than three of them.
What Others Add
What the Leading Models Are Actually Doing Independent papers behind the domain leaders
The names behind the survey's citation counts have their own literature, and it explains why each leads rather than merely how often it is cited. BioBERT is BERT re-trained on PubMed abstracts and PMC full text, which is what lets it outperform general-purpose BERT at biomedical named-entity recognition, relation extraction and question answering. ESM-1b learned its representations from 250 million unlabelled protein sequences; its direct descendant, ESM3, reasons jointly over a protein's sequence, structure and function and is itself one of the survey's post-2025 risers. UNI, the imaging-and-signals leader, is pretrained on tissue images from over 100,000 whole-slide pathology scans across 20 tissue types and was tested on 34 separate diagnostic tasks. scGPT, the omics leader, is pretrained on more than 33 million individual cells. Evo 2, a second riser, is trained on 9 trillion DNA base pairs spanning every domain of life and predicts the effect of a genetic variant — including clinically significant BRCA1 mutations — without being fine-tuned for that task at all.
- 250MProtein sequences ESM-1b learned from, with no labels at all (Rives et al., 2021).
- 100k+Whole-slide pathology scans behind UNI's pretraining, across 20 tissue types (Chen et al., 2024).
- 33MIndividual cells behind scGPT's pretraining (Cui et al., 2024).
- 9TDNA base pairs behind Evo 2, spanning every domain of life (Brixi et al., 2026).
The Integrations Nobody Has Built Yet three openings, one shared bottleneck
Mechanistic interpretability
Language ↔ sequence and omics
- A language interface could assist scRNA-seq annotation, pathway enrichment and spatial transcriptomics.
Therapeutic targets
Imaging ↔ omics
- Integration of H&E pathology with omics remains weak.
- The potential to link tissue morphology to molecular mechanism is considerable.
Regulatory layers
DNA, RNA, protein, epigenome
- The layers are only loosely connected, protein to the others most of all.
The bottleneck is not ideas
What blocks this is not a shortage of concepts but the state of the data: paired datasets are scarce, preprocessing and annotation are not standardised, and there is no standard multimodal benchmark. It is the same cause that explains the scarcity of decoders earlier.
Conclusion
What Gets a Model Adopted and what still stands in the way
- It addresses a high-impact application: clinical language processing, protein modelling, digital pathology.
- It is rigorously evaluated: testing across varied downstream tasks builds trust and reuse.
- It is accessible: open pretrained weights, open code, a usable interface.
- And newly decisive: it fits into agent-based AI workflows.
Still unsolved
Data quality, cohort diversity, batch effects, privacy, hallucination and adversarial robustness all still limit real-world deployment. The authors recommend cross-disciplinary collaboration — biologists, clinicians and model developers together — and efficient adaptation of strong pretrained backbones in place of training from scratch.