What the Filters Remove
Single-cell genomics · Bioinformatics · Data Quality · Methods Briefpeer-reviewed record to Jan 2026 · vendor material captured 26 Aug 2026
Methods brief · how a cell becomes a column
The Filters That Decide What a Single-Cell Experiment Saw
Between a sequencer and a matrix of genes by cells sits a chain of filters, each discarding a different thing that is not a healthy single cell: a barcode that was only background, a cell that was already dying, two cells sharing one label, and RNA that leaked out of one cell and was counted in another. Every filter is a threshold, and a threshold is a claim about what could not have been a cell. The most widely used of them — discard anything above 5% mitochondrial content — has been measured against 5,530,106 cells, and it does not survive a change of tissue.
By the Numbers the measured ones
- 13 / 44human tissues where the 5% mitochondrial default fails to separate healthy from low-quality cells
- 5,530,106cells, across 1,349 annotated datasets, behind that finding
- 1.7%cross-species doublet rate published for one combinatorial kit — a floor, not a total
- 3peer-reviewed comparisons of the two chemistries — and three different verdicts
What Each Filter Removes, and How Firmly We Know standing first, then the claim
| How firm | What | Who says so |
|---|---|---|
| Confirmed | Each filter in the chain has a published method behind it: EmptyDrops separates cells from background; Scrublet, DoubletFinder and scDblFinder find two cells under one label; SoupX, DecontX and CellBender remove RNA counted in the wrong cell | the method papers themselves, 2019 to 2023 |
| Confirmed | The 5% mitochondrial default fails to discriminate healthy from low-quality cells in 13 of the 44 human tissues examined, or 29.5%. Mean mitochondrial content is significantly higher in human tissue than in mouse, and the sequencing platform does not explain the difference. For mouse tissues the same threshold generally performs well | Osorio and Cai, Bioinformatics, 2021, over 5,530,106 cells in 1,349 datasets |
| Confirmed | Of nine doublet-detection methods benchmarked over 16 real datasets with experimentally annotated doublets and 112 synthetic ones, DoubletFinder had the best detection accuracy and cxds the highest computational efficiency | Xi and Li, Cell Systems, 2020 |
| Confirmed | On cerebellar organoids, combinatorial barcoding and droplet capture gave comparable cellular diversity and technical reproducibility — but the droplet arm carried more stressed cells, and the combinatorial arm lower mitochondrial and ribosomal protein-coding transcript fractions and higher gene-biotype diversity | Sarieva and colleagues, iScience, January 2026 |
| Vendor-reported | Doublet rates of less than 3% even at 100,000 cells, and 1.7% for one kit version against 1.3% for its predecessor, measured by mixing human and mouse cell lines in equal numbers and counting the transcriptomes that carry both genomes | Parse Biosciences support suite and dataset page, captured 26 Aug 2026 |
| Vendor-reported | Free-floating RNA would have to follow a cell's exact path through four rounds of split-pool barcoding to be misassigned — put at 1 in 3,538,944 — and a wash step removes free molecules before lysis. Ambient RNA is worse than dropout, because dropout loses a gene while ambient RNA gives one to the wrong cell: a CD8+ T cell reading as though it carried CD4 | Parse Biosciences support suite, captured 26 Aug 2026; no independent measurement of the figure was found |
| Our inference | A species-mixing plot cannot see a doublet made of two cells of the same species. With an equal two-way mix roughly half of all doublets are therefore invisible, and any published cross-species figure is a floor on the true rate rather than an estimate of it. This holds for every vendor's species-mixing number, not one company's | this project's reading, on the arithmetic in Bloom, PeerJ, 2018 |
In What Order
How the Toolkit Was Built publications only — the vendor support pages carry three dates each and are not placed here
- 2016Ilicic and colleagues set out the general problem of classifying low-quality cells, reporting over 30% better accuracy than traditional methods across more than 5,000 cells — Genome Biology
- Mar 2018SPLiT-seq: split-pool ligation labels each cell's RNA by combinatorial barcoding, with no microfluidics and no custom equipment. 156,049 nuclei, over 100 cell types — Science. It is the method behind the combinatorial platform in every comparison below
- Sep 2018Bloom derives how an observed rate of mixed transcriptomes relates to the true multiplet frequency, for cell types mixed in any proportion — PeerJ
- Mar 2019EmptyDrops calls cells by testing each barcode against the ambient expression profile, keeping distinct cell types that earlier methods discarded — Genome Biology
- Apr 2019Scrublet and DoubletFinder appear in the same issue of Cell Systems. Both build artificial doublets out of the data itself and ask which real cells sit nearest them
- Jun 2019Luecken and Theis publish the field's first end-to-end best-practice tutorial, from quality control to downstream analysis — Molecular Systems Biology
- 2020The ambient-RNA problem gets its tools: DecontX in March, estimating per-cell contamination as a Bayesian mixture, and SoupX in December, quantifying the contamination and correcting the counts. Between them, in December, Xi and Li benchmark nine doublet callers
- 2021Osorio and Cai test the 5% mitochondrial default across 5,530,106 cells and publish reference values for 121 mouse and 44 human tissues instead — Bioinformatics, May. scDblFinder follows in September
- 2023Heumos and colleagues extend best practice across modalities in Nature Reviews Genetics, March. CellBender follows in August, modelling droplet background noise as a generative process and reporting operation near the theoretically optimal denoising limit on simulated data — Nature Methods
- 2024The two chemistries are compared head to head, twice, and disagree. Xie and colleagues on human PBMCs in March — International Journal of Molecular Sciences; Filippov and colleagues on mouse thymus in November — BMC Genomics
- Jan 2026Sarieva and colleagues compare them a third time, on cerebellar organoids, and find the stress signal itself differs between the two workflows — iScience. This is the most recent technical finding on the page
- Aug 2026Commercial context, and one party's account only: Parse Biosciences announces that the US Court of Appeals for the Federal Circuit rejected 10x Genomics' appeal, affirming the invalidation of patents asserted against it. The court's opinion was not read here, and the release's own dateline (19 August) and byline (20 August) disagree. It bears on who may sell what; it bears on none of the measurements above
The Argument
A Threshold That Does Not Travel what the measurements say, and what they do not
The mitochondrial fraction is the workhorse of single-cell quality control, and the logic behind it is sound: a stressed cell leaks cytoplasmic transcripts through a failing membrane while its mitochondrial transcripts, enclosed in their own organelle, stay put. So the mitochondrial share of a cell's counts rises as the cell dies, and a cell above some share is discarded. The share almost everybody uses is 5%. Osorio and Cai report that it was set by early publications, adopted as a default in several analysis packages, and taken up as a de facto standard — and that its validity across species, technologies, tissues and cell types had not been adequately assessed. So they assessed it, over 5,530,106 cells in 1,349 annotated datasets. In mouse tissue the threshold generally holds. In human tissue it fails to discriminate healthy from low-quality cells in 13 of the 44 tissues they examined, and mean mitochondrial content runs significantly higher in human than in mouse for reasons the sequencing platform does not explain. Their answer is not a better constant but a table: reference values for 121 mouse and 44 human tissues. Five years later Sarieva and colleagues found the same metric moving again, this time with the chemistry — the droplet arm of their organoid comparison carried more stressed cells and higher mitochondrial and ribosomal transcript fractions than the combinatorial arm, on the same biology. A threshold that shifts with the species, the tissue and the capture method is not a property of cells. It is a property of an experiment, and it is being typed in as though it were a property of cells.
- 29.5%of the human tissues tested — 13 of 44 — where the 5% default does not separate healthy cells from low-quality ones
- 121 + 44tissues, mouse and human, for which reference values are published instead of a single number
- 1,349annotated datasets behind the finding, holding 5,530,106 cells
What this does not say
It does not say the 5% figure is wrong. On mouse tissues its authors report that it generally performs well, and on many human tissues it works too. The finding is about transfer, not about the number: a default carried unchanged from the tissue it was derived on to a tissue nobody checked.
What it does say
A filter is not neutral machinery. Every cell it removes is a cell the experiment is no longer allowed to have seen, and in nearly a third of the human tissues tested the commonest filter cannot tell a dying cell from a mitochondria-rich one. The authors put the consequence plainly: omitting the filter, or adopting a suboptimal threshold, may lead to erroneous biological interpretations.
What Others Add
Four Things the Rest of the Literature Adds chemistry, doublets, ambient RNA, scale
Chemistry
Three comparisons, three verdicts
- Human PBMCs, two donors — comparable cell-type frequencies. Combinatorial barcoding captured fewer cells but detected rare types, plasmablasts and dendritic cells among them, on higher gene-detection sensitivity. Gene length and GC content distributed differently between platforms (Xie, 2024)
- Mouse thymus — combinatorial barcoding detected nearly twice the genes, and each platform detected a distinct set. But the droplet data had lower technical variability and more precise annotation of biological states (Filippov, 2024)
- Cerebellar organoids — comparable diversity and reproducibility, but more stressed cells in the droplet arm and lower mitochondrial and ribosomal fractions in the combinatorial one (Sarieva, 2026)
Doublets
Nobody agrees which caller to use
- Xi and Li benchmarked nine methods and put DoubletFinder first on detection accuracy and cxds first on computational efficiency (2020)
- The scDblFinder paper states that scDblFinder was already found by an independent benchmark to outcompete alternatives (2021). It is a different benchmark, and it is not a re-run of the first
- Whatever the caller, the input parameter it usually wants is an expected doublet rate — and the figures available for that are species-mixing figures, which cannot see a same-species doublet at all
Ambient RNA
The toolkit is droplet-shaped
- SoupX, DecontX and CellBender each describe themselves as droplet methods in their own abstracts. CellBender's model reasons explicitly about the difference between cell-containing and cell-free droplets
- Ambient RNA is worse than dropout, and the reason is worth holding on to: dropout loses a gene, while ambient RNA hands one to the wrong cell — a CD8+ T cell that reads as though it also carried CD4
- A workflow with no droplets has no cell-free droplets to model. Parse Biosciences puts its own misassignment odds at 1 in 3,538,944 across four barcoding rounds, plus a wash before lysis — the vendor's figure, arithmetically consistent, and repeated by nobody else
Scale
What the published guidance actually says
- Published minimum sequencing depth for one combinatorial platform: 20,000 reads per cell for kit versions 2 and 3, and 10,000 for version 4 on improved library efficiency. Deeper for rare cell types in a heterogeneous population; shallower is defensible for a homogeneous one
- A published one-million-cell walkthrough, on peripheral blood from 24 donors, specifies a cloud instance of at least 8 threads and 160 GB of RAM
- Automatic cell-type annotation, where a commercial platform offers it, is ScType, Decoupler and CellTypist; clustering is Leiden or Louvain. None of that is proprietary, and the choice of which is a judgement the platform hands back to you
So What
Five Things to Carry Out of This and one question nobody has settled
- Check the mitochondrial threshold against your tissue, not against the default. Reference values now exist for 121 mouse and 44 human tissues, and 13 of those 44 are tissues where the 5% figure does not do the job.
- Read a species-mixing doublet figure as a floor. Whatever the plot shows, the same-species half of the doublets was never in it — and that figure is usually what your doublet caller asks you to type in.
- Do not read the platform comparisons as a scoreboard. Three peer-reviewed studies, three tissues, three different answers. Pick the one whose tissue and endpoint resemble yours, and treat the other two as evidence that the question is open.
- Ask whether your ambient-RNA tool was built for your chemistry. All three of the common ones say droplet in their own abstracts. That is not proof they fail elsewhere; it is a reason to find out before trusting the corrected counts.
- Treat a vendor figure as a vendor figure. Every one on this page may well be right. None of them has been independently repeated, and each is a measurement a company made of its own product.
Still open
Which doublet caller to use. One benchmark of nine methods put DoubletFinder first on accuracy; the scDblFinder paper cites a different independent benchmark putting scDblFinder ahead. Both are published, neither has been re-run against the other here, and this page does not choose between them.
Bottom line
Every step in this chain is a threshold, and a threshold is a claim about what could not have been a cell. The most widely used one in the field is measurably tissue-specific and platform-sensitive, and it is typed in as a constant. That is not a failure of the metric. It is a failure to carry its conditions along with it — and the conditions have been published since 2021, tissue by tissue.