Tratopedia
Side N繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.177.0

The release this page was built from. It is what the service worker caches under.

DNS · Rust · Performance Engineering · Software Engineering · Technical Reportpublished 27 Aug 2026 · rollout 18 May to 6 Jul 2026

Representation, not algorithm · and the rare case where smaller is also faster

533 Bytes an Entry, 100 Terabytes a Fleet

Big Pineapple, the platform behind 1.1.1.1, holds over 250 billion DNS cache entries at any moment. Five changes to how one of those entries is laid out in memory took it from 953 bytes to 420, freeing roughly 100 terabytes across Cloudflare's fleet. Not one of the five is an algorithm: they are changes to which Rust type holds what. The layout claims below were checked by compiling them, not by taking them on trust.

The figures Every one re-derived from the post's own before and after

  • −56%per-entry cache footprint, 953 bytes to 420, in the benchmark
  • ~100 TBaggregate working set freed across the fleet, measured in production
  • 250bnDNS cache entries held at any given time
  • +43%cache insert throughput, 625,000 to 893,000 entries a second
  • −19%cache lookup latency, 828 ns to 670 ns

The record, graded Some of this a reader can check without Cloudflare's help

StandingWhatHow it was checked
ConfirmedA Vec<T> is 24 bytes and a Box<[T]> is 16, so dropping the unused capacity field saves 8 bytes per field — 64 bytes across the eight such fields in a cache entry. The same holds for String against Box<str>.Compiled here and the sizes printed. Matches the post exactly.
ConfirmedReplacing three record lists with one list and two u16 section offsets saves 28 bytes an entry: three fat pointers cost 48 bytes, one plus two offsets costs 20.Compiled here. 48 − 20 = 28, as stated.
ConfirmedA Rust enum is as large as its largest variant, so an A record holding 4 bytes of address sat in a 144-byte slot sized for NAPTR. Boxing the large variants brings the enum to 24 bytes, saving 120 on every A and AAAA record — and those are 81% of the traffic.Compiled here: the boxed enum is 24 bytes, and NAPTR rebuilt from the post's own prose is 136. The enum's 144 depends on their exact field types.
ConfirmedDNS repeats owner names on the wire and compresses them with a two-octet pointer, so the cache stores the full owner instead — trading memory for not chasing pointers on the hot path. Where the owner equals the queried name, it is now dropped and inferred from the cache key.The two-octet compression pointer is RFC 1035 §4.1.4, read directly. Its first two bits are 11, leaving a 14-bit offset.
Confirmed, not verifiable hereThe production result: p99 resident memory per instance fell from 9.3 GB to 5.3 GB and p90 from 6.5 GB to 3.8 GB, over a rollout from 18 May to 6 July 2026, for an aggregate of about 100 TB.Both percentages re-derive correctly. Only Cloudflare can measure its own fleet; no breakdown is published.
UnconfirmedThat 130 Gen 13 servers hold 100 TB of RAM between them, which the post gives as the equivalence. That works out at about 769 GB a server; no published Gen 13 memory specification was found to check it against.Arithmetic from the post's own two numbers; the specification is not public.

Rolled out, then written up The work was finished and in production seven weeks before it was described

  1. 18 May 2026The rollout begins. Each release carries one or more of the five changes, so memory falls in steps rather than at once.
  2. 18 May – 6 Jul 2026Restarted instances begin with empty caches and use less memory until those caches fill, so the dips in the graph are not the result. The plateaus are.
  3. 6 Jul 2026The rollout completes across all services — 49 days end to end.
  4. 27 Aug 2026Cloudflare publishes the account, by Sebastiaan Neuteboom, with the benchmark method, the production graph and nine diagrams.

What has no date Two of them, and one is the point of the whole exercise

The post says each release carried one or more of the five changes but does not say which carried which, so the steps in the graph cannot be matched to the optimisations that caused them. And the freed memory has not been spent: Cloudflare says it plans to reinvest it into a larger cache, which would raise hit rates and cut upstream query volume. That is written in the future tense, and no date is given. Until it happens, the result of this work is 100 terabytes of headroom rather than a better cache — which is a perfectly good result, and a different one from the one the plan describes.

Change the representation, not the algorithm And the premise that lets you: the entry is never modified again

The claim the post makes about its own work is modest and worth taking seriously: no cache policy, eviction rule or data structure changed. What changed is which type holds what. The premise that unlocks all five changes is one sentence in the second section — once a DNS response is in the cache, it is never modified again. A Vec carries a capacity field and over-allocated heap space so that it can grow; a value that will never grow is paying for a capability it cannot use. The same reasoning runs through the rest: three separate record lists become one list with two u16 offsets, because section counts fit in 16 bits; a record's owner name is dropped when it equals the queried name, because the cache key is already in hand at every lookup; and the record-data enum stops being sized for its rarest variant. The most interesting pair is the last two, because they argue with each other. Boxing the large enum variants fixes the padding but introduces allocator round-up and scatters data across the heap; storing record data as one contiguous byte buffer removes both of the costs that boxing had just created. A reader who stopped at the fourth change would build something worse than one who read on — and the post is unusually honest in laying the intermediate step out rather than presenting the final design as though it arrived whole.

  • 5changes, none of them to an algorithm
  • 1premise underneath all five: the entry is never mutated
  • 2of the five undo costs the others introduced

The five changes In the order they are described, which is also the order they build on each other

ChangeWhy it is freeSaving
1. Vec<T> becomes Box<[T]>, String becomes Box<str>A cached entry never grows, so the capacity field and the reserved heap tail are both dead weight.64 bytes an entry, plus the wasted heap tail
2. Three record lists become one, with two u16 section offsetsRecord counts per section fit in 16 bits, so an offset costs 2 bytes where a list costs a 16-byte fat pointer.28 bytes an entry
3. Drop a record's owner name when it matches the queryThe cache key is present at every lookup, so the name can be restored rather than stored. Most records match; those behind a CNAME do not, and keep theirs.One heap allocation per record, in the common case
4. Box the large enum variantsIt is not free. It fixes 120 bytes of padding on the common records but adds allocator round-up and scatters data across the heap.120 bytes per A or AAAA record
5. Store record data as one contiguous byte bufferRemoves both costs change 4 introduced, and lets most record types be copied straight into an outgoing response without being re-serialised. The cost: records can no longer be randomly indexed.Lookup latency −5%, insert throughput +13%, on their own

Three ways to check this without Cloudflare A rare property for an operator's post about its own fleet

  • A compiler

    The layout claims

    rustc, optimised, edition 2021

    • Vec 24 bytes against Box<[T]> 16, and String 24 against Box<str> 16 — so 8 a field, 64 across eight fields. Exactly the post's figure.
    • Three fat pointers 48 bytes, one plus two u16 offsets 20 — a 28-byte saving, as stated.
    • Rebuilding NAPTR from the post's own prose lands on 136 bytes exactly. The enum came out 136 rather than 144 — the tag fitting in spare padding, which depends on field types the post does not publish. The mechanism is confirmed; the last 8 bytes are theirs.
  • An RFC

    The DNS claims

    RFC 1035, November 1987

    • §4.1.4 gives name compression as “a two octet sequence” whose first two bits are 11, leaving a 14-bit offset. The post's description is correct.
    • Why the cache does not use it is Cloudflare's own judgement, not the RFC's: chasing compression pointers on the hot lookup path is expensive, so it spends memory instead.
  • Arithmetic

    The published results

    fourteen figures, re-derived

    • Every percentage in the post re-derives from its own before and after: −55.9%, −58.1%, +42.9%, −19.1%, −43.0%, −41.5%.
    • The traffic mix sums to 100%, and A plus AAAA is 81% — the post's “over 80%”.
    • “Over 15 terabytes” for the first change is 16.0 TB decimal but 14.6 TiB binary. True as decimal, and conservative, since the entry count is a floor.

The guarantee the post relies on and does not name This report's own finding, compiled and confirmed

The third change stores the owner as Option<Box<Name>> and sets it to None when the owner equals the queried name. Compiled here, Option<Box<T>> is the same size as Box<T> — 16 bytes either way. That is Rust's null-pointer optimisation: because a Box can never be null, the compiler uses the null pointer itself as the None discriminant, so the Option wrapper costs nothing. Without that guarantee the change would be a much worse deal. Every record would pay a discriminant byte plus alignment padding for the privilege of sometimes omitting an owner — buying back a heap allocation in the common case at the price of bytes in every case, on a struct where the whole exercise is counting bytes. The post presents the change as a straight trade of storage for a lookup-time inference and does not mention the language rule that makes the trade one-sided. It is the sort of thing that is invisible when it works and expensive when it does not, and it is worth knowing which of the two you are relying on.

What the benchmark is, and what it is not The post draws this line itself, which is worth carrying over

The 56% is a benchmark figure: entries randomly generated to roughly match the production record mix, with TXT standing in for every non-A/AAAA type at a random 64 to 224 bytes, and memory counted by a custom allocator wrapping Rust's system allocator. The post says in its own words that these inputs “approximate production rather than reproduce it exactly”, and that process memory also depends on traffic mix, cache occupancy, allocator state and everything outside the cache — which is why resident memory was measured across production instances as well. The two numbers are different sizes and the gap is instructive. Taking the benchmark saving straight to the fleet gives 533 bytes times 250 billion entries, or about 133 terabytes. The reported figure is 100 — roughly three quarters of the product. Cloudflare could have published the larger number with a footnote and did not. It is also why the production graph's plateaus rather than its dips are the result: restarted instances start with empty caches, and an empty cache is not an efficient one.

What generalises, and what does not Three findings, each with its limit

The checkable part is the part that matters

The number in the title, about 100 terabytes, is the least checkable figure in the post: an aggregate over a fleet only its operator can measure, with no breakdown. The mechanism beneath it is the most checkable, because anyone with a compiler can reproduce it in a minute. That is the right way round. A reader who confirms that Vec is 24 bytes and Box<[T]> is 16 has earned a reason to believe the part they cannot see, and the post supplies enough detail for that check to be possible — which is not true of most engineering write-ups that lead with a headline number.

Smaller and faster together is not a law

Space and time usually trade against each other, and here they did not: the footprint fell 56% while inserts got 43% faster and lookups 19% quicker. The reason is specific rather than general. Fewer allocations means less allocator work on the insert path, and contiguous data means fewer cache-line fetches on the lookup path, so both metrics happened to be improved by the same changes. Change three is the counterexample within the same post: it explicitly spends memory — storing full owner names instead of following compression pointers — to keep the hot path fast. The lesson is to measure both, not to expect both.

The saving is headroom, not yet a better cache

Cloudflare says it plans to spend the freed memory on a larger cache, which would raise hit rates and cut upstream query volume. That is written in the future tense with no date. Until it happens, what the work produced is 100 terabytes of unused capacity and a faster cache of the same size — a good outcome, and a different one from the plan. Whether the reinvestment lands, and what it does to hit rates, is the thing to watch and the thing nobody can yet report.

What would settle the rest None of it is contested; all of it is simply unpublished

  • Whether the freed memory becomes cache capacity. Stated as a plan, in the future tense, with no date. A follow-up post with hit-rate figures would settle it.
  • Which release carried which change. The post says each carried one or more, so the steps in the production graph cannot be attributed to individual optimisations.
  • The Gen 13 memory specification. The 130-server equivalence implies about 769 GB a server; no published specification was found to check it.
  • How much of the fleet is ECS-heavy. The post says the gains are largest where EDNS Client Subnet multiplies cached variants of the same query, but gives no share.

Checked on 30 Aug 2026 against: Sebastiaan Neuteboom, “How we saved 100 terabytes of memory by optimizing 1.1.1.1's DNS cache”, The Cloudflare Blog, published 27 Aug 2026 17:02 UTC, fetched and read in full with its nine diagrams · RFC 1035 §4.1.4, P. Mockapetris, November 1987, read at rfc-editor.org for the two-octet compression pointer · this session's own rustc, optimised, edition 2021, used to compile the type shapes the post describes and print each size — the layout figures attributed here to compiling are this report's own measurements, not citations. Every percentage in the post was re-derived from its own before-and-after values and all fourteen agree. Not established: the composition of the ~100 TB; Cloudflare's Gen 13 memory specification; the share of the fleet using ECS; which release carried which change; the exact NAPTR field types that make the enum 144 rather than 136 bytes.

Versions

This document is rewritten when what it says has to change. Every version stays published at its own address.

  1. v0001 current

The current version is also at latest/.