Tratopedia
繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.81.0

The release this page was built from. It is what the service worker caches under.

Sequencing technology · AI · Benchmarking · Genomics · ComparisonTratopedia · 17 Aug 2026

Four generations of Oxford Nanopore basecalling models, read as a history of changes rather than a league table

Dorado v3 to v6 — What Changed, How Much, and Whether It Was Good

Between 2023 and May 2026 Oxford Nanopore shipped four generations of basecalling model. Each one was better than the last on the measurements that exist. But the gains shrank at every step, the compute bill stepped up sharply once, and the newest generation dropped the most accurate tier entirely — so the most accurate model Oxford Nanopore ships is still sup@v5.2.0, released in May 2025. This is an account of what changed, how large each change was, and which of them were worth having.

  • 0DNA sup@v6 models released — v6 is a HAC-only generation
  • 18.7%fewer read errors from HAC v5.2.0 to v6.0.0, recomputed from the published Q values
  • 43.8%fewer read errors sup@v5.2.0 still has than hac@v6.0.0
  • 2.7–5.7×how much slower SUP is than HAC when both run in the same run

What changed, and how firmly each change is established standing first, because a grade placed after the claim reads as a footnote; rows are grouped by standing and never interleaved

StandingThe changeSize
Confirmed — re-derivableSUP became a transformer at v5.0.0; HAC stayed recurrent throughout and got a new custom recurrent design at v6.0.0sup 78,718,162 parameters against hac@v5.2.0 8,790,768 — 8.95×
Confirmed — re-derivableOxford Nanopore released no sup@v6.0.0 and no fast@v6.0.0v6.0.0 is HAC only
Confirmed — re-derivablesup@v5.0.0 and sup@v5.2.0 are the same size78,718,162 parameters both — weights and training changed, architecture did not
Published measurementHAC median read accuracy rose at every step, v5.0.0 → v5.2.0 → v6.0.0Q16.3 → Q17.2 → Q18.1, i.e. 18.7% fewer errors at each step
Published measurementThe largest SUP gain in the history was v4.2.0 → v4.3.0, before the transformermedian Q18.5 → Q20.5, 36.9% fewer read errors; assembly errors across nine genomes 173 → 37
Published measurementSUP read accuracy did not move from v5.0.0 to v5.2.0median ~Q20.6 both; assembly mean Q60.4 → Q61.6
Vendor claimv5 SUP “achieves Q26” — 99.75% modal raw-read accuracymodal, on data Oxford Nanopore selected
Vendor claimHAC v6.0.0 reaches single-molecule accuracy of Q23.5against an independently observed median of Q18.1 and a mode near Q19.7
Vendor claimThe SUP tier is “no longer needed”stated as the reason for shipping no SUP v6
One user reportsup@v5.2.0 scored lower Q than sup@v5.0.0 on one matched dataset27,751 identical reads; 62.5% fewer passed Q20
Not publicAny per-read comparison of sup@v4.3.0 against sup@v5.0.0 on identical raw datathe architecture change is the biggest event here and the worst evidenced
SupersededAn earlier research pass tabulated accuracy, yield, memory and throughput for a fourth SUP generationthe model it described does not exist

The fabricated column is a finding, not the subject

This article was first written from a research pass that tabulated accuracy figures for a DNA sup@v6. No such model was released, and Oxford Nanopore said publicly why it was not shipping one. That is worth knowing — a fluent table is not evidence that its columns exist — but it is one row in the table above. Building a whole article around a tool’s mistake describes the tool, not the subject.

Every percentage here was recomputed

Read accuracy is quoted as a Q score, and a Q score is a logarithm: error rate is 10^(−Q/10). Differences in Q therefore cannot be subtracted, and the error reductions on this page were each derived from the Q values printed beside them rather than carried across from the source. Three of them came out differently — see “What others add”.

The order it happened in model versions are dated by release; every entry here is documented, and the one thing that cannot be placed in the sequence is named at the end

  1. 2023v4.2.0 brings 5 kHz sampling — about 12.5 samples per nucleotide against roughly 10. A chemistry-side change, not a model-quality one, and it should not be folded into the accuracy trend.
  2. 2024v4.3.0 SUP folds in bacterial-methylation training that had previously needed a separate research model. Still LSTM, and still the largest single accuracy jump in this history.
  3. May 2024v5.0.0 at London Calling, with Dorado 0.7: SUP switches from LSTM to a transformer; HAC stays recurrent. Oxford Nanopore headlines Q26. Clive Brown says the highest-accuracy Q30 mode is “four times slower” with about 15% yield loss.
  4. 2025Dorado 1.0.0 removes R9.4.1 support — last carried in v0.9.6 — along with 4 kHz R10.4.1 and RNA002. Models are chemistry-specific, so this closes the door on re-basecalling old data with new models.
  5. May 2025v5.2.0 at London Calling, with Dorado 1.0.0: HAC grows from five to seven LSTM layers; SUP keeps its architecture and changes only weights and training.
  6. May 2026v6.0.0 at London Calling, with Dorado 2.0.0: a new custom recurrent architecture for HAC only. Mike Vella, VP of Machine Learning: “Importantly, we’re not releasing a sup model this time… we believe that this particular model is good enough that a sup tier is no longer needed.”
  7. 9 Jul 2026The explanation for HAC v6’s roughly 100 errors on one Klebsiella genome changes: 7-deazaguanine from the dpd system, per a Biggel preprint, superseding an earlier suggestion of phosphorothioate backbone modifications. Both are the same shape of cause — a modification absent from training.

The two orderings a reader might use — by date, and by how well evidenced each step is — disagree in one place, and it is the place that matters most. The transformer switch at v5.0.0 is the largest change in this history and the one with the weakest direct evidence. There is no public per-read comparison of sup@v4.3.0 against sup@v5.0.0 on identical raw data, and no clean yield comparison for that transition either. What exists is consensus-level: peer-reviewed cgMLST work found v5.0 significantly better than v4.3 on the same isolates, with average allele distance to an Illumina reference improving from about 1.78 to about 0.04 once polished.

That is real evidence and it is not the same evidence. Consensus accuracy is what most laboratories care about; per-read accuracy is what a model version is claimed to change. The step everyone cites as the architectural turning point has to be taken partly on the vendor’s word.

  • 1.78 → 0.04average cgMLST allele distance to an Illumina reference, v4.3 against v5.0 with polishing — peer-reviewed, and consensus-level rather than per-read
  • nonepublic per-read v4.3.0-against-v5.0.0 comparisons on identical raw data

The comparison’s own case what the attached report claims, and where its own figures complicate it

The report’s argument is a shape rather than a verdict: within a tier, newer generations have generally improved accuracy, but the gains have decelerated sharply and are not always monotonic — while performance moves in the opposite direction, because the transformer SUP models are markedly slower and hungrier than the LSTM ones they replaced.

It is careful in the two places this subject invites carelessness. It separates vendor figures from independent ones instead of averaging them, and it states plainly that the SUP tier remains the accuracy choice despite Oxford Nanopore’s messaging that it is no longer needed. Its own numbers then complicate its trend line twice: the SUP step from v5.0.0 to v5.2.0 is flat and on one dataset slightly negative, and modified-base calling does not follow the version order at all — for routine CpG 5mC the v4-era models scored highest.

  • deceleratingthe report’s central claim about accuracy, and the reason a version number is a poor proxy for a decision
  • oppositethe direction performance moves as accuracy improves
  • Claim 1

    Accuracy improved, then stopped improving much

    • The biggest SUP jump was v4.2.0 to v4.3.0 — before the architecture changed at all.
    • v5.0.0 to v5.2.0 moved SUP read accuracy essentially not at all.
    • HAC kept gaining in real but modest steps: 18.7% fewer read errors per generation.
  • Claim 2

    The cost went the other way

    • SUP carries 8.95× HAC’s parameters and is the compute bottleneck in every measured run.
    • Oxford Nanopore’s own Q30 mode costs about four times the speed and roughly 15% of the yield.
    • HAC v6.0.0 is the exception: more parameters than v5.2.0 and slightly faster, through optimisation.
  • Claim 3

    The tier that was dropped is still the best one

    • No sup@v6.0.0 exists; the reason given is that HAC v6 makes the tier unnecessary.
    • Independent testing has sup@v5.2.0 ahead of hac@v6.0.0 by 43.8% on read errors.
    • Median assembly errors: 4 for SUP v5.2.0 against 11 for HAC v6.0.0.

The measurements, and three places our arithmetic disagrees independent testing is the backbone of every figure here; the three disagreements below are this article’s own, derived from the source’s own numbers

TransitionMedian read accuracyFewer read errorsMedian assembly errors
HAC v5.0.0 → v5.2.0Q16.3 → Q17.218.7%37.5 → 13
HAC v5.2.0 → v6.0.0Q17.2 → Q18.118.7%25.5 → 11
SUP v4.2.0 → v4.3.0Q18.5 → Q20.536.9%173 → 37 total, nine genomes
SUP v5.0.0 → v5.2.0~Q20.6 → ~Q20.63 → 3
sup@v5.2.0 against hac@v6.0.0Q20.6 against Q18.143.8%4 against 11
Same run, same dataHACSUPRatio
Dorado 1.0.0, no modification callinghac@v5.2.0 3h18msup@v5.2.0 18h45m5.7×
Dorado 2.0.0, modification calling onhac@v5.2.0 8h04msup@v5.2.0 21h33m2.7×
  • Our arithmetic, 1

    A percentage that travelled without its number

    HAC v5.0.0 to v5.2.0

    • The source reports about 15% fewer read errors, and tabulates Q16.3 to Q17.2.
    • Q16.3 to Q17.2 is 18.7%. The 15% figure belongs to Q17.0 — 14.9%.
    • The source records that the earlier post said Q17.0 and the later one Q17.2, so the Q value was updated and the percentage was not. Minor, and a working example of why a percentage should be recomputed from the figures beside it.
  • Our arithmetic, 2

    Not an order of magnitude

    SUP against HAC throughput

    • Compared only within a single run, SUP costs 5.7× HAC’s time with modification calling off and 2.7× with it on.
    • Throughput ratio and wall-clock ratio agree to two decimal places in both runs, so this is not a misreading.
    • The 9× figure is the parameter ratio, which is correct. The slide from parameters to throughput is the step to resist.
  • Our arithmetic, 3

    One generation’s only real gain was software

    SUP v5.0.0 to v5.2.0

    • Identical parameter counts, and read accuracy that did not move.
    • Throughput rose 34% — 25h06m to 18h45m on the same data.
    • The source prints both numbers without joining them. On this transition, what improved was the engineering rather than the model.

Each change, sized and judged the three questions a version history has to answer: what changed, was it large, was it worth having

ChangeHow largeWas it good
SUP v4.2.0 → v4.3.0, LSTM weightsLarge — 36.9% fewer read errors and assembly errors cut to 21.4% of beforeYes, unambiguously. The best value in the whole history, and it came from training rather than from architecture
SUP v4.3.0 → v5.0.0, LSTM to transformerClaimed large; not measurable per-read from anything publicProbably, on consensus evidence. The peer-reviewed cgMLST improvement is real, and the per-read case rests on the vendor’s figures
SUP v5.0.0 → v5.2.0, weights onlyNone on accuracy; 34% on throughputGood but not as advertised. A speed release wearing a model version, and on one dataset the Q went down
HAC v5.0.0 → v5.2.0, five to seven layersModest on reads, 18.7%; large on assemblies, 65% fewer errorsYes — and it cost about 20% throughput, which is the honest price
HAC v5.2.0 → v6.0.0, new recurrent designModest on reads, 18.7%; 57% fewer assembly errors, at a different subsample depthYes, and cheaply — slightly faster despite more parameters. The caveat is a genome it fails on
No SUP v6The accuracy ceiling did not move in 2026Not on the evidence. The tier called unnecessary is still 43.8% better on read errors than the model said to replace it
  • For accuracy-critical work, use sup@v5.2.0. As of August 2026 it is the most accurate model released, and the newer HAC generation does not match it. Do not read the absence of a SUP v6 as a recommendation to move.
  • If compute-bound, hac@v6.0.0 is the right HAC choice — more accurate than v5.2.0 and no slower. Know its one documented failure: roughly 100 errors on a genome carrying base modifications its training did not include.
  • Do not treat a newer version as automatically higher quality on your data. One SUP transition was flat, and negative on one dataset. Verify on a subset before adopting.
  • Re-examine any Q-score filter after changing model. On identical reads, a newer model put 62.5% fewer of them over Q20 — recalibration, not lost sequence. Check aligned accuracy rather than Q.
  • Match the model to the modification, not to the calendar. For routine CpG 5mC the v4-era models scored highest; v5-era and later are better for non-CpG 5mC and 4mC, and for 6mA.
  • Re-basecall old data with the newest model for its own chemistry before comparing studies — but only its own: R10 models cannot be applied to R9 data, and support for R9.4.1 ended after Dorado v0.9.6.

The pattern, in one line

Every generation was an improvement and each improvement was smaller than the last, while the compute cost stepped up once, sharply, and stayed there. That is the ordinary shape of a maturing technology, and it is why a version number is a poor proxy for a decision: the largest accuracy gain in this history came from retraining an old architecture, and the most recent generation’s clearest win was that it did not get slower.

What to hold loosely

Nearly every cross-version figure here comes from bacterial isolates on one chemistry, measured by one independent tester; human, metagenomic and RNA results may differ. Vendor and independent numbers are not like-for-like — modal against median, on differently chosen data — and the gap between them is not by itself evidence that either is wrong. The timing figures come from shared machines and their author calls them loosely controlled, which is why every ratio on this page is taken within a single run.

What would change the answer

A released sup@v6.0.0; or a HAC model independently shown to reach SUP’s roughly Q20.6 read accuracy and four-error assemblies; or a Dorado release that closes the throughput gap. Until one of those, the accuracy ceiling is a model from May 2025 and the tier that reached it has been discontinued.

Attached technical comparison, Oxford Nanopore Dorado Basecalling Model Versions (v3 → v6), 17 Aug 2026 — a synthesis which attributes its figures to: Ryan Wick’s independent testing on bacterial isolates (R10.4.1/E8.2, 5 kHz); Oxford Nanopore at London Calling (Clive Brown, 2024; Mike Vella, 2026); Diallo et al., J. Clin. Microbiol. 2025, on cgMLST accuracy; Dorado GitHub issue #1565, for the matched 27,751-read dataset; and a preprint by Michael Biggel, University of Zurich, on 7-deazaguanine modifications. Percentages and ratios on this page were recomputed from the Q values and timings the source tabulates; where the result differs from the source’s own figure, both are shown and the difference is attributed to this article.