Tratopedia
Side N繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.177.0

The release this page was built from. It is what the service worker caches under.

Single-cell genomics · Benchmarking · Research briefNature Biotechnology, 22 Apr 2026

Stanford · Silesian University of Technology · Yale

A Single-Cell Method That Needs No Reference Genome at All

A team led by Julia Salzman at Stanford and Sebastian Deorowicz at the Silesian University of Technology has published sc-SPLASH, a statistical method that finds regulated sequence variation in barcoded single-cell and spatial sequencing data directly from the raw reads, without aligning to a reference genome. The paper is peer-reviewed, published as a Brief Communication in Nature Biotechnology, and its own figures show it consistently faster and lighter on memory than the field's standard tools, Cell Ranger and STARsolo, across every one of the eleven datasets it is benchmarked on.

  • ~50×faster than UMI-tools, for the barcode and UMI preprocessing step (BKC)
  • 11benchmark datasets on which sc-SPLASH ran faster and used less memory than both Cell Ranger and STARsolo
  • 5new secreted-protein genes found in a sponge with no matching reference genome
  • 47homologous YYD-repeat genes found across five tunicate genome assemblies

How Firmly Each Part Stands Read directly off the published figures, and read only in the predecessor preprint.

How firmWhatWho says so
ConfirmedDehghannasiri, Kokot and nine co-authors describe sc-SPLASH, its BKC preprocessing module, and discoveries in human, sponge, tunicate and electric-eel dataNature Biotechnology, peer-reviewed and PubMed-indexed (PMID 42020848)
ConfirmedAcross 11 benchmark datasets, sc-SPLASH used less runtime and less memory than both Cell Ranger v8.0.1 and STARsolo v2.7.10b in every caseFigure 1c of the published paper, read directly
ConfirmedBKC, the barcode and UMI preprocessing step, ran about 50-fold faster than UMI-toolsThe paper's own abstract
ConfirmedOf every Pfam protein domain found in the Tabula Sapiens data, V-set (the immunoglobulin variable-region fold) has the highest average entropy, the highest average effect size, and the largest number of unaligned extendorsFigure 1d of the published paper, read directly
Reported, and supersededV-set's effect size of about 0.90. Its entropy was given as 2.16 in the earlier preprint, but the published figure's own axis stops short of that and places V-set near 1.8 — so the precise entropy is not carried here.The effect size is consistent with Figure 1d, read directly; the entropy figure comes from the predecessor bioRxiv preprint, which the published figure does not bear out
Confirmed, but not verifiable hereGranrep expression in the sponge rises over development, from gemmule (day 0) to day 12; a preliminary test found no evident difference in expression after immune challenge with lipopolysaccharide or a cyclic dinucleotideStated in the predecessor bioRxiv preprint's own text; not independently re-read in the paywalled final text

What sc-SPLASH Actually Found Five results, each read from the published figures or their own supplementary note

  • Immune-receptor diversity, human — 5,212 in-frame V(D)J sequences (4,098 from plasma cells, 1,114 from B cells) recovered across the Tabula Sapiens tissues, concentrated in spleen (1,632), lymph node (1,484), salivary gland (409), bone marrow (378) and blood (332).
  • A cancer mutation, spatially resolved — in a human squamous-cell-carcinoma Visium sample, the strongest signal (effect size 0.757) is a mitochondrial MT-ND4 double mutation (ChrM 11,413–11,466), expressed mainly, but not only, inside the tumour region.
  • The same splice choice, two phyla apart — an RPS24 exon included in human fetal-intestine epithelium and excluded in stromal cells is echoed by the electric eel's own RPS24: included in the electrocytes of its electric organ, excluded in the insulating septa around them.
  • Five genes missing from a sponge's own reference genome — a 27-nucleotide anchor absent from both the Spongilla lacustris assembly and NCBI (entropy 6.2, 667 distinct targets, the highest of any anchor in the dataset) led to five newly named genes, Granrep1–5, expressed mainly in granulocytes (47 of 74 profiled express it) and amoebocytes (12 of 49); RNA-FISH confirms 30 of 34 cells co-express it with the granulocyte marker Acp5.
  • A parallel repeat in a second phylum — the tunicate Ciona's own high-diversity anchor (entropy 4.80–5.97, the only repeat anchor above 4 among nine invertebrates surveyed) points to 47 homologous “YYD” genes across five tunicate genome assemblies, expressed in circulating haemocytes and peaking at metamorphosis, twelve hours after hatching.

Four Years, One Method Growing Up From the first SPLASH preprint to this paper, and what came after it

  1. 24 Jun 2022Chaung, Baharav, Henderson, Zheludev, Wang and Salzman first post the original SPLASH algorithm to bioRxiv.
  2. 17 Mar 2023Kokot, Dehghannasiri, Baharav, Salzman and Deorowicz post SPLASH2 — the fast, scalable engine sc-SPLASH's later stages reuse — to bioRxiv.
  3. 7 Dec 2023SPLASH is published, peer-reviewed, in Cell.
  4. 23 Sep 2024SPLASH2 is published, peer-reviewed, in Nature Biotechnology.
  5. 24 Dec 2024Dehghannasiri, Kokot and colleagues post the sc-SPLASH preprint to bioRxiv, extending SPLASH2 to barcoded single-cell and spatial data for the first time.
  6. 22 Apr 2026The peer-reviewed, condensed version is published as a Brief Communication in Nature Biotechnology.
  7. 25 Aug 2026Europe PMC records no independent citation, response or replication of the paper.

What the Statistics Buys, and What It Gives Up The core claim, attributed to the paper itself

The paper's argument is that a reference genome is not required to find real, regulated biology in sequencing data — only statistics is. sc-SPLASH takes a fixed k-mer from each read, the anchor, and the varying k-mer that follows it, the target, and asks whether the distribution of targets for a given anchor depends on which sample (which cell, in barcoded data) it came from. The OASIS test gives a closed-form p-value for that question and an effect size from 0 (identical distributions) to 1 (completely separable ones); anchors are called significant at an effect size above 0.2 and a Benjamini–Yekutieli-corrected p-value below 0.05. Nothing in this process ever looks at a reference genome. The price, the paper's own methods make explicit, is that an anchor alone carries no gene-level meaning: turning a significant anchor back into biology (a gene name, a splice site, a Pfam domain) is a separate, optional step — building an “extendor” and aligning or searching it — that the reference-free part of the pipeline does not need and does not do.

  • Effect size, 0 to 1How separable the two target distributions are.
  • A closed-form p-valueThe OASIS test keeps multiple testing under control without simulation.

Two Places the Headline Numbers Need a Second Look What complicates the paper's own framing

The benchmark compares two different jobs

Cell Ranger and STARsolo only align reads; a separate tool such as Seurat or Scanpy is still needed afterwards for differential expression. sc-SPLASH's reported runtime already includes calling significant anchors — a further step. So the gap Figure 1c shows is, if anything, an understatement of the full pipeline-to-pipeline gap, not an overstatement — but the two are also not finishing the same job.

A repeated protein in immune cells is not the same as a proven immune role

The paper's biological argument — that repeat proteins turning up in immune-like cells of both a sponge and a tunicate is evidence of a shared strategy for generating diversity under immune pressure — rests on where the genes are expressed, not on a demonstrated immune trigger. In the one preliminary sponge experiment that tested a trigger directly, expression rose over ordinary development but showed no evident change after immune challenge. The parallel across two phyla is real; a causal immune function is not yet shown for either.

Two Papers It Builds On, Three Tools It Is Measured Against None of this is new to sc-SPLASH itself

Paper or toolWhat it established
SPLASH
Chaung et al., Cell, 2023
The anchor/target statistic itself: a reference-free test for sample-dependent sequence variation, first shown on SARS-CoV-2, plate-based single-cell data and non-model organisms such as eelgrass and octopus.
SPLASH2
Kokot et al., Nature Biotechnology, 2024
The fast, scalable k-mer-counting engine. sc-SPLASH's stages 2 and 3 (merging counts, computing p-values, multiple-testing correction) reuse it directly; only stage 1 (BKC) is new.
UMI-tools
Smith, Heger & Sudbery, Genome Research, 2017
The standard network-based method for correcting UMI sequencing errors and deduplicating reads — the job BKC is benchmarked against and, on this paper's own figures, replaces about 50 times faster.
10x Chromium chemistry
Zheng et al., Nature Communications, 2017
The droplet-barcoding chemistry that Cell Ranger processes, and that every barcoded dataset sc-SPLASH is benchmarked or demonstrated on was generated with.
STARsolo
Kaminow, Yunusov & Dobin, bioRxiv, 2021
The second alignment-based comparator in Figure 1c. Never itself published beyond this preprint — a reminder that a widely used tool is not always a peer-reviewed one.

Somebody Else's Atlases, and No Response Yet Where the demonstration datasets came from, and what has been said about the paper itself

None of the three human datasets sc-SPLASH is demonstrated on are its own: the >400,000-cell Tabula Sapiens atlas is the Tabula Sapiens Consortium's (Science, 2022); the squamous-cell-carcinoma Visium sample is Ji and colleagues' (Cell, 2020); the fetal-intestine Visium sample is Fawkner-Corbett and colleagues' (Cell, 2021). sc-SPLASH re-analyses each without touching its original conclusions. As for the paper's own reception: Europe PMC records no independent citation, response or replication as of 25 August 2026, four months after publication.

What to Take Away A solid engineering result, a real but not-yet-causal biological pattern, and no outside check yet

The speed and memory result is the firmest part

Across all 11 benchmark datasets, sc-SPLASH used less runtime and less memory than both Cell Ranger and STARsolo, and BKC ran about 50 times faster than UMI-tools — read directly off the paper's own figures and abstract. The V(D)J and Ciona-homolog counts on this page are the exact figures printed in the paper's own Figure 1e and Supplementary Note, not a secondary retelling of them.

The biology is real; the immune story is not yet closed

Five new genes in a sponge with no matching reference, and a parallel repeat family in a tunicate, are genuine discoveries the reference-based tools this paper is compared against could not have made without them already being annotated. That both families sit in immune-like cells, in two disparate animals, is a real pattern. That they respond to immune challenge is not yet shown — the one preliminary test of that ran negative.

Four months old, and not yet independently checked

The paper is peer-reviewed and its code is public, but as of August 2026 no independent group has cited, replicated or disputed it. That is ordinary for a four-month-old Brief Communication, not a mark against it — it simply means the claims above rest on one paper and its own predecessor preprint, not yet on anyone else's confirmation.

Dehghannasiri, Kokot, Starr et al., Nature Biotechnology, 22 Apr 2026, DOI 10.1038/s41587-026-03084-6 (PMID 42020848) · predecessor preprint, bioRxiv, 24 Dec 2024, DOI 10.1101/2024.12.24.630263 (PMID 39763839) · Chaung et al., Cell, 7 Dec 2023, DOI 10.1016/j.cell.2023.10.028 (PMID 38065078) · Kokot et al., Nature Biotechnology, 23 Sep 2024, DOI 10.1038/s41587-024-02381-2 (PMID 39313645) · Smith, Heger & Sudbery, Genome Research, 18 Jan 2017, DOI 10.1101/gr.209601.116 (PMID 28100584) · Kaminow, Yunusov & Dobin, bioRxiv, 5 May 2021, DOI 10.1101/2021.05.05.442755 · Zheng et al., Nature Communications, 16 Jan 2017, DOI 10.1038/ncomms14049 (PMID 28091601) · Tabula Sapiens Consortium et al., Science, 13 May 2022, DOI 10.1126/science.abl4896 (PMID 35549404) · Ji et al., Cell, 23 Jun 2020, DOI 10.1016/j.cell.2020.05.039 (PMID 32579974) · Fawkner-Corbett et al., Cell, 5 Jan 2021, DOI 10.1016/j.cell.2020.12.016 (PMID 33406409) · Baharav, Tse & Salzman, PNAS, 2 Apr 2024, DOI 10.1073/pnas.2304671121 (PMID 38564640)

Versions

This document is rewritten when what it says has to change. Every version stays published at its own address.

  1. v0001 current

The current version is also at latest/.