Tratopedia
Side N繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.177.0

The release this page was built from. It is what the service worker caches under.

Biomedicine · Machine Learning · Benchmarking · Research BriefNature Biotechnology · April 2026

A survey of 220 biomedical foundation models

Biomedical Foundation Models Are Moving From One Modality to Many

A survey of 220 biomedicine-specific foundation models developed over four years. More than half — 53.6% — already combine two or more data modalities, and natural language is emerging as the connective interface between otherwise unrelated kinds of biomedical data. General-purpose large language models such as ChatGPT, Gemini and Claude were deliberately excluded.

  • 220models curated, over four years
  • 53.6%are multimodal — 118 combine two or more modalities
  • 17distinct data modalities in total
  • 4domains: language, imaging, omics, sequences
StandingWhat happenedDetail
ConfirmedThe survey is publishedChang, Cheng, Modi, Wang, Xu and Ma, in Nature Biotechnology — online 30 April 2026, in print that August at 44(8):1263–1266. DOI 10.1038/s41587-026-03135-y.
Confirmed220 models, four fields, more than half multimodalEvery count and share on this page — the 220, the 118 (53.6%), the architecture split and the evaluation split — is stated in the paper's own running text and was read there.
Confirmed, not verifiable hereWhich 220 models they areThe roster is Supplementary Table 1. The repository serves that file behind a challenge this project could not satisfy, so the list was never read model by model — the count is the paper's, not a recount.
Confirmed, not verifiable hereThe citation countsThey come from Semantic Scholar and are printed only inside the paper's figure, which does not say when the snapshot was taken. They were read off that figure. Citation counts move, so treat them as one undated moment rather than a current ranking.

What Counts as a Biomedical Foundation Model the survey's boundaries

  • Large-scale machine learning models trained on varied biomedical and clinical data.
  • That data spans genetic sequences, molecular profiling, biomedical imaging and electronic health records.
  • They learn generalisable representations to support downstream discovery and clinical work.
  • The survey covers four domains: natural language, imaging and signals, omics, and molecular sequences.
  • General-purpose LLMs are excluded; the focus is 220 biomedicine-specific models.

From BioBERT to Evo 2 Every model the survey names, dated to its own first publication

  1. 15 Feb 2020BioBERT — BERT re-trained on PubMed abstracts and PMC full text — is published; it becomes the survey's most-cited natural-language model.
  2. 13 Apr 2021ESM-1b shows that a language model trained on 250 million protein sequences, with no labels at all, learns real structural information; it is the ancestor of the survey's molecular-sequence leader.
  3. 26 Feb 2024scGPT, pretrained on more than 33 million individual cells, is published; it becomes the survey's omics leader.
  4. 19 Mar 2024UNI, pretrained on tissue images from over 100,000 whole-slide pathology scans, is published; it becomes the survey's imaging-and-signals leader.
  5. 16 Jan 2025ESM3 — a generative model reasoning jointly over protein sequence, structure and function — is published; the survey later names it a post-2025 riser.
  6. 4 Mar 2026Evo 2, trained on 9 trillion DNA base pairs spanning every domain of life, is published; the survey names it a second post-2025 riser.
  7. 30 Apr 2026Chang, Cheng, Modi, Wang, Xu and Ma publish the survey itself, online, in Nature Biotechnology.
  8. Aug 2026The survey appears in Nature Biotechnology's print issue, volume 44, number 8, pages 1263–1266.

The key observation

Natural language is the modality most often paired with others, and is becoming the connective interface between data types. This runs through the whole paper: the emerging citation leaders and the underdeveloped directions both point back to the same structure.

Encoders Dominate architecture breakdown

ArchitectureCountShare
Encoder-based11451.8%
Mixed6730.5%
Decoder-based3114.1%
Non-Transformer (e.g. Mamba)—3.6%

Decoders are scarce not out of architectural preference but because of data: text-generation tasks need high-quality paired image–text annotation, and that is hard to obtain in biomedicine.

Evaluation Is Still Narrow downstream tasks used

TaskShare
Classification50.9%
Report generation10.9%
Question answering8.6%
Segmentation7.3%

The evaluation gap

There are no multitask benchmarks, little human-in-the-loop assessment, and no combined treatment of data quality, interpretability and translational relevance. Half of these models have been tested on classification alone.

Citation Leaders by Domain * marks a multimodal model · counts are a Semantic Scholar snapshot the paper does not date

DomainLeading models and citations
Natural languageBioBERT 6,925 · MultiMedQA 3,707 · PubMedBERT 2,278
Imaging and signalsUNI 1,141 · ConVIRT* 1,003 · MedCLIP* 794
OmicsscGPT 846 · Geneformer 836 · scBERT 509
Molecular sequencesESM-1b 2,807 · ProtTrans 1,207 · DNABERT 1,051

The Risers Since 2025 growth is coming from multimodal work

DomainRisers since 2025, and citations
Natural languageMedGemma-27B 165 · MediPhi-Instruct 13
Imaging and signalsBiomedCLIP* 498 · Quilt-1M* 216 · MedGemma* 165
OmicsNicheformer* 81 · OmiCLIP* 58 · scGPT-spatial 34
Molecular sequencesESM3 201 · Evo 2 197 · Borzoi 190

Citation growth is being driven by models that integrate heterogeneous modalities rather than by specialists in a single data type. Five of the eleven risers are multimodal, and natural language is the only field with fewer than three of them.

What the Leading Models Are Actually Doing Independent papers behind the domain leaders

The names behind the survey's citation counts have their own literature, and it explains why each leads rather than merely how often it is cited. BioBERT is BERT re-trained on PubMed abstracts and PMC full text, which is what lets it outperform general-purpose BERT at biomedical named-entity recognition, relation extraction and question answering. ESM-1b learned its representations from 250 million unlabelled protein sequences; its direct descendant, ESM3, reasons jointly over a protein's sequence, structure and function and is itself one of the survey's post-2025 risers. UNI, the imaging-and-signals leader, is pretrained on tissue images from over 100,000 whole-slide pathology scans across 20 tissue types and was tested on 34 separate diagnostic tasks. scGPT, the omics leader, is pretrained on more than 33 million individual cells. Evo 2, a second riser, is trained on 9 trillion DNA base pairs spanning every domain of life and predicts the effect of a genetic variant — including clinically significant BRCA1 mutations — without being fine-tuned for that task at all.

  • 250MProtein sequences ESM-1b learned from, with no labels at all (Rives et al., 2021).
  • 100k+Whole-slide pathology scans behind UNI's pretraining, across 20 tissue types (Chen et al., 2024).
  • 33MIndividual cells behind scGPT's pretraining (Cui et al., 2024).
  • 9TDNA base pairs behind Evo 2, spanning every domain of life (Brixi et al., 2026).

The Integrations Nobody Has Built Yet three openings, one shared bottleneck

  • Mechanistic interpretability

    Language ↔ sequence and omics

    • A language interface could assist scRNA-seq annotation, pathway enrichment and spatial transcriptomics.
  • Therapeutic targets

    Imaging ↔ omics

    • Integration of H&E pathology with omics remains weak.
    • The potential to link tissue morphology to molecular mechanism is considerable.
  • Regulatory layers

    DNA, RNA, protein, epigenome

    • The layers are only loosely connected, protein to the others most of all.

The bottleneck is not ideas

What blocks this is not a shortage of concepts but the state of the data: paired datasets are scarce, preprocessing and annotation are not standardised, and there is no standard multimodal benchmark. It is the same cause that explains the scarcity of decoders earlier.

What Gets a Model Adopted and what still stands in the way

  • It addresses a high-impact application: clinical language processing, protein modelling, digital pathology.
  • It is rigorously evaluated: testing across varied downstream tasks builds trust and reuse.
  • It is accessible: open pretrained weights, open code, a usable interface.
  • And newly decisive: it fits into agent-based AI workflows.

Still unsolved

Data quality, cohort diversity, batch effects, privacy, hallucination and adversarial robustness all still limit real-world deployment. The authors recommend cross-disciplinary collaboration — biologists, clinicians and model developers together — and efficient adaptation of strong pretrained backbones in place of training from scratch.

Chang, Cheng, Modi, Wang, Xu & Ma, Nature Biotechnology, published online 30 Apr 2026, print vol. 44 no. 8 pp. 1263–1266 (Aug 2026). DOI 10.1038/s41587-026-03135-y (PMID 42062613) · Lee et al., Bioinformatics, 2020 (PMID 31501885) · Rives et al., PNAS, 2021 (PMID 33876751) · Chen et al., Nature Medicine, 2024 (PMID 38504018) · Cui et al., Nature Methods, 2024 (PMID 38409223) · Hayes et al., Science, 2025 (PMID 39818825) · Brixi et al., Nature, 2026 (PMID 41781614). Citation counts in this article's tables are as reported in the survey's own Fig. 1, current as at its April 2026 publication.

Versions

This document is rewritten when what it says has to change. Every version stays published at its own address.

  1. v0001 current

The current version is also at latest/.