Summary · Nature Biotechnology, 2026
Single-cell genomics · Method briefNature Biotechnology · Brief Communication
Stanford · Silesian University of Technology · Yale
Skip Alignment — Infer Straight From the Raw Reads
sc-SPLASH extends reference-free, statistics-first inference to 10x Chromium and Visium data. Rather than aligning to a reference genome, it extracts k-mer pairs from the raw reads and tests whether their distribution depends on the sample. The price is giving up direct gene-level interpretation; what it buys is the ability to see what a reference-based pipeline erases — alleles, paralogues, repeats, and genes that are not in the annotation at all.
- ~20×faster than Cell Ranger, and that includes the statistics
- 50×speed-up of BKC preprocessing over UMI-tools
- 8 GBmemory used, against 70 GB for Cell Ranger
- 400K+single cells analysed in Tabula Sapiens
Context · Motivation and method · Content 1 / 3
Why scRNA-seq Still Mostly Measures Gene Expression four limits of the existing tools
- Alignment bias is hard to avoid, and worst for species without a good reference.
- Tools do not talk to each other — splicing, V(D)J and editing each have their own, and the compute adds up.
- Barcoded droplet data is enormous — millions of cells share one library, and barcodes and UMIs are error-prone.
- The blind spot: alleles, paralogues and repeats — real biological diversity — are erased by a reference-based pipeline.
The Intuition: Anchors and Targets no reference required
The method takes k-mer pairs straight from the raw reads: a fixed k-mer — the anchor — followed by a varying one, the target. For each anchor it builds a table of target counts across samples, then tests whether that distribution depends on the sample. SPLASH had already shown this works on bulk and Smart-seq2 data; this study extends it to barcoded platforms.
- Effect size, 0 to 1How separable the two target distributions are.
- A closed-form p-valueThe OASIS test gives a closed-form p-value, which keeps multiple testing under control.
A Three-Stage Pipeline from barcoded FASTQ to significant anchors
- 01BKCbarcode extraction, UMI deduplication
- 02Contingency tablesparse matrices merge the counts
- 03OASISp-values with BY correction
The significance threshold is an effect size above 0.2 with p below 0.05, Benjamini–Yekutieli corrected. Post-processing is optional: build an extendor and compare against a reference, Pfam, BLAST or IgBLAST, all automatable from a JSON configuration.
Pivot · Performance and human data · Content 2 / 3
Performance: BKC Preprocessing a parallel C++ rewrite replacing single-threaded Python
| Measure | BKC | UMI-tools |
|---|---|---|
| Average runtime | 165 s | 9,272 s, for whitelist and extract only — deduplication not included |
| Memory | 7 GB | — |
| Extras | Single-base barcode correction, 3-to-1 base packing, SATC output | — |
Performance: Against Cell Ranger and STARsolo Tabula Sapiens muscle, donor 1 · 16 threads
| Tool | Runtime | Memory | Note |
|---|---|---|---|
| Cell Ranger v8.0.1 | 2,540 s | 70 GB | Alignment only |
| STARsolo v2.7.10b | 491 s | 35 GB | Alignment only |
| sc-SPLASH | 106 s | 8 GB | Includes the statistical inference |
Read that comparison carefully
Cell Ranger and STARsolo only align; Seurat or Scanpy is still needed afterwards for differential expression. sc-SPLASH runs one pipeline through to significant anchors. So the gap in speed is in fact wider than the table shows — but the two are also not producing the same kind of output.
Tabula Sapiens: 400,000 Human Cells across 16 tissues and two donors
Supervised inference with an L1-regularised GLM over cell-type metadata found 555 cell-type-specific genes, including known splicing targets such as RPS24 and MYL6. The same anchors recurred across donors and tissues — RPS24 37 times, MYL6 28. SPLASH is robust to batch effects because target variation is tested as a conditional probability given the anchor: anchor clusters overlapped between donors in the same tissue far more than chance would give, at p below 2.2 × 10⁻¹⁶ for both lung and muscle.
V(D)J: The Highest-Entropy Pfam Domain diversity is exactly what alignment cannot see
| Measure | Value | What it means |
|---|---|---|
| Entropy | 2.16 | V-set has the highest entropy of any Pfam domain, far above keratin, histone and the rest |
| Effect size | 0.90 | The two target distributions are almost entirely separable — cell specificity is unambiguous |
| Unalignable | 35 / 121 | 28.9% could not be aligned to the genome — high diversity being precisely the blind spot of reference-based tools |
| IgBLAST | 60,697 | in-frame V(D)J sequences, from plasma and B cells across 16 tissues |
Resolution · Spatial data, other species, new genes · Content 3 / 3
Visium: A Mitochondrial Double Mutation in Squamous Cell Carcinoma the same pipeline, unmodified
In Visium spatial data from a human cutaneous squamous cell carcinoma, the strongest signal was a CC → TT double mutation in MT-ND4 (effect size 0.757, ChrM 11,413–11,466), expressed mainly but not exclusively in the tumour region — consistent with an early spontaneous mutation in the carcinoma lineage. The second strongest distinguished the keratin paralogues KRT16 and KRT17 (effect size 0.348): the tumour region tends to express KRT17 and normal epithelium KRT16, both being commonly upregulated in this cancer.
- Spot resolutionVisium works at spot resolution, but the barcoded framework applies without modification.
A Splicing Difference Conserved Across Species: RPS24 Exon 6 human fetal tissue and the electric eel
| Species | Observation |
|---|---|
| Human | Fetal small intestine, Visium, highest effect size: epithelial cells include the 3-nt microexon (exon 5), stromal cells exclude it. |
| Electric eel | Main electric organ: electrocytes include exon 6, while the stromal cells of the insulating septa exclude it (Chr11 7,503,567–7,506,064). |
Why the correspondence holds
Electrocytes derive from the skeletal muscle lineage — one of the few human cell types that also includes exon 6. The RPS24 exons are homologous between the two, so this is not coincidence but a cell-type distinction conserved across species.
A Non-Model Organism: The Sponge's granny Anchor a sequence that is in no reference
| Measure | Value | Note |
|---|---|---|
| Entropy | 6.2 | The highest in the entire Spongilla lacustris dataset |
| Distinct targets | 667 | The number of different sequences following one anchor |
| BLASTn and reference hits | 0 | None in odSpoLacu1.1 or NCBI — the sequence is in no annotation at all |
| Cell-type specificity | 73 | Of the cells expressing granny, 47 (64%) are granulocytes and 12 (16%) amoebocytes |
HCR RNA-FISH confirmed that 88% of Acp5⁺ cells are also granny⁺. Manual assembly from PacBio HiFi long reads revealed five Granrep genes — secreted repeat proteins sharing one structure: signal peptide, then 30-bp granny repeats, then a lysine-rich region, then 18-bp C-terminal repeats, with the granny repeats likely O-glycosylated. Twelve days after treatment with LPS or cGAMP, total Granrep expression rose about 1.4-fold, in step with granulocyte markers such as Acp5 — which hints at a role in immune defence.
A Parallel in Another Phylum: the Sea Squirt's YYD Repeat Ciona robusta haemocytes
Ciona's 24-bp tandem repeat anchor is the only repeat anchor with entropy above 4 among aquatic invertebrates, ranging 4.80–5.97 across subsets. The gene consists largely of that 24-bp repeat, carries only a signal peptide, and encodes a secreted protein; most homologues share a conserved YYD motif, presumed to be tyrosine sulfation as in cionin — a clear parallel with the sponge's Granrep. RNA-FISH shows expression mainly in circulating haemocytes in juveniles, peaking at metamorphosis, twelve hours after hatching.
- 2 + 43Two homologues in the HT genome, plus 43 TBLASTn hits across C. intestinalis and C. savignyi.
What sc-SPLASH Delivers five conclusions
- New genes without a reference — Granrep is absent from the Spongilla annotation, and YYD is highly polymorphic across sea squirt species.
- One pipeline for droplet and spatial data — 10x Chromium and Visium alike, with no cell labels required.
- BKC stands alone — UMI deduplication 50 times faster, droppable into an existing pipeline.
- JSON-driven post-processing — extendor through STAR, Bowtie2, Pfam and IgBLAST, fully automated.
- Suited to non-model organisms — repeat-rich, highly polymorphic, reference-poor work on immunity and diversity.