Tratopedia
繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.177.0

The release this page was built from. It is what the service worker caches under.

AI Security · Cryptography · Information Security · AI Agents · Research BriefarXiv:2608.09867, 10 Aug 2026 · record current to September 2026

One key, every model

AI Reasoning Traces: Encrypted, Not Hidden

Anthropic, OpenAI and Google now send each model’s step-by-step reasoning back to the client as an encrypted block rather than as plain text, to keep it from competitors and from the user reading it directly. A paper published in August 2026 shows the encryption is not bound to the session, the user or the model that produced it: a captured block can be replayed into a smaller, less-guarded sibling model, which decodes and repeats it verbatim.

A ten-week gap separates a hobby project from a working attack The headline figures, before the grading.

  • 315,320reasoning blocks decoded from 6,708 scraped public agent traces
  • 704privacy artefacts recovered from genuine, non-benchmark sessions alone
  • 64of those found only inside the encrypted reasoning, never in the visible chat
  • $720estimated cost to decode ten thousand reasoning traces at one provider’s cheapest model

What is established, and how firmly Sorted by standing; the grades are never mixed within a row.

StandingWhat the record showsSource
ConfirmedResearchers demonstrated the extraction attack against Anthropic, OpenAI and Google’s APIs: a reasoning block captured from a heavily safeguarded model was replayed into a weaker, less-guarded sibling from the same provider, which transcribed it verbatim.Panfilov et al., Aug 2026
ConfirmedDecoding 315,320 reasoning blocks scraped from 6,708 public agent traces recovered hundreds of credentials and personal-data items from genuine sessions alone, including 62 API keys, 33 passwords, 24 access tokens and 7 private keys.Panfilov et al., Aug 2026
ConfirmedThe paper’s authors disclosed the flaw to the affected providers, Microsoft and Hugging Face before publication; every provider acknowledged receipt. Matthew Green, whose earlier post the paper builds on, separately reported the underlying replay behaviour to OpenAI and Anthropic in May 2026.Panfilov et al.; M. Green
Confirmed but not verifiable hereThe paper’s own account is that, as of August 2026, its central extraction results no longer reproduce against the providers’ APIs, following mitigations the providers put in place after disclosure.Panfilov et al., Aug 2026 (self-reported)
Confirmed but not verifiable hereThe Hacker News reports that Anthropic’s current documentation ties a thinking block to the model that produced it and advises stripping it out when switching models, since other models ignore it.The Hacker News, 12 Aug 2026
Confirmed but not verifiable hereAs of its own publication date, The Hacker News found no public acknowledgement from Anthropic, OpenAI or Google naming this specific paper.The Hacker News, 12 Aug 2026
UnconfirmedWhether reasoning blocks already published online before the providers’ fixes remain decodable today, as distinct from whether new extraction attempts still succeed, is not established either way.The Hacker News, 12 Aug 2026
Not publicAs of the paper’s own testing period, no provider had published a detailed description of the cryptographic mechanism securing its reasoning blocks.Panfilov et al., Aug 2026

From a weekend experiment to a scalable extraction method Dated as each source dates itself.

  1. 2026Matthew Green, a cryptographer at Johns Hopkins University, publishes a hobby write-up finding that encrypted reasoning blocks from OpenAI, Anthropic and Google APIs can be replayed across sessions, across accounts, and — for OpenAI — across models, without a valid decryption key.
  2. 29 May 2026Green reports the replay behaviour to OpenAI and Anthropic through their bug-bounty programmes. By his own account, OpenAI calls the report unreproducible and Anthropic says it sees no security implication in replay or side-channel behaviour.
  3. Jul 2026The paper’s authors run their API evaluation across Anthropic, OpenAI and Google, developing the cross-model extraction method and the public-trace scan.
  4. 10 Aug 2026“Stealing Reasoning Traces from Proprietary LLM APIs” is submitted to arXiv, describing the cross-model extraction method and the public-trace scan’s findings.
  5. 11 Aug 2026Green appends an update to his May post pointing readers to the new paper, describing it as turning his own finding into “a real working attack.”
  6. 12 Aug 2026The Hacker News publishes its own coverage, corroborating the paper’s figures and adding that the encryption itself was never broken — the attack relies on intact blocks being accepted and processed by the provider.
  7. Aug 2026By the paper’s own account, all three providers deploy mitigations following disclosure, and its central extraction results stop reproducing.
  8. Sep 2026Current documentation shows the three providers have not converged on one posture: Google still describes backend-managed cross-model compatibility, while OpenAI now restricts reasoning reuse to models in the same family and automatically omits the rest.

The cryptography was never the weak point — the boundary around it was Attributed to the paper and to Green’s own account, by name.

Panfilov and colleagues describe an architectural choice, not a broken cipher: to avoid storing every reasoning trace on their own servers, Anthropic, OpenAI and Google each package it into an encrypted block and hand it to the client to hold and return on the next turn. Their tests indicate each provider authenticates and encrypts every such block with one key, shared across the whole provider, rather than a key tied to a single session, a single user or a single model. Nothing in the cryptography itself needed breaking — the block simply carries no marker of where it was made or who it belongs to.

Green’s May post is the origin of that observation, and the paper says so directly. He found that a captured block would replay across sessions, across accounts, and — for OpenAI — across different models, and that a replayed block is “semantically active” rather than simply ignored. He also demonstrated, on data he planted himself, that the length and timing of a reasoning block can leak a secret bit the model was told never to state outright. What he did not manage, by his own account, was to turn either finding into a reliable way of pulling a real secret out of someone else’s session. The paper’s own contribution is the missing step: routing a captured trace through a smaller, less-guarded sibling model from the same provider, which decodes and transcribes it on request — without the more capable model that produced the reasoning ever being asked to reveal it directly.

That routing step works because of a security asymmetry the paper names directly: a frontier model is heavily trained to refuse requests to reveal its own chain-of-thought, but its cheaper, faster siblings — built for cost and speed rather than as the flagship product — are not held to the same standard. The attacker never needs to jailbreak the capable model at all; it only needs a compatible sibling willing to decode what it is handed.

The same portability the paper exploits deliberately, a third party can exploit incidentally. Developers routinely publish raw agent session logs for reproducibility, carrying their reasoning blocks along with the visible transcript; a third party who can read those blocks with any compatible model recovers whatever the model reasoned over, whether or not it ever appeared in the visible text. The paper frames this, and a second consequence — that a malicious instruction can be planted entirely inside a reasoning block and later replayed into an unrelated task, invisible to anyone reading only the visible conversation — as the reason opacity cuts both ways: it can hide harmful content a model reasoned about without saying, and it can just as easily hide something a user or a monitor needed to see.

The paper draws its own conclusion from this design tension: as long as some model must decrypt and act on prior reasoning to continue a conversation, an encrypted reasoning block “can never be more than semi-hidden”, whatever the transport-level cryptography looks like — the content stays reachable through whichever model implicitly holds the key. That is the precise sense in which the reasoning is encrypted, not hidden: it was encrypted against the user reading it and against a rival scraping it in bulk; it was never, and could not be, hidden from a model in the same family asked nicely enough to say it back.

Three providers, three different postures a month on Read directly from each provider’s current developer documentation.

  • Google

    Compatibility stays, and the backend manages it

    Gemini API docs, “Thought signatures”

    • Its own documentation still calls thought signatures “encrypted representations of the model’s internal reasoning”, required for continuity across turns.
    • On switching models within a session, the guidance is to keep resending the previous model’s thought blocks — “the backend manages compatibility.”
  • OpenAI

    A boundary has appeared since the paper’s testing window

    OpenAI Platform docs, reasoning models guide

    • Stateless responses still return an encrypted_content reasoning item for the client to hold and pass back.
    • But persisted reasoning “can be reused only within the same model family”, and the API now omits incompatible reasoning automatically — narrower than the July 2026 behaviour the paper measured, where reasoning carried across every earlier GPT generation tested.
  • Anthropic

    A claim not found in Anthropic’s own current documentation

    The Hacker News, citing Anthropic’s documentation

    • The Hacker News reports that Anthropic’s documentation ties a thinking block to its producing model and advises stripping it on a switch, since other models ignore it.
    • Anthropic’s own public Extended Thinking documentation, as of September 2026, does not contain that language and does not use the word “signature” at all.

What the disclosure record and the mitigation proposals add From the paper’s own discussion section and Green’s own recommendations.

The paper’s own list of fixes ranges from the drastic to the incremental: moving reasoning storage back onto the provider’s own servers and handing the client only an opaque lookup identifier would remove the replayable asset entirely, at a real infrastructure cost; short of that, it proposes binding each encrypted envelope to a specific user and conversation, and rejecting a block presented outside the context that produced it. Green’s own recommendation from May, aimed at the providers directly, was narrower and arrived first: improve key management, and stop treating a side channel as beneath a security team’s attention just because it is slow.

Encryption protected the wrong boundary For anyone who publishes agent traces, holds session state, or reads a provider’s documentation to plan around it.

  • Strip reasoning or thinking blocks from any session log or agent trace before publishing it, whenever that session ever touched anything sensitive — scrubbing the visible text alone leaves whatever the model reasoned over sitting in the encrypted block.
  • Do not treat an opaque, unreadable-looking block as safe to hand to another party just because it cannot be read without the key — a compatible model, not the key, is what turned out to be able to read it.
  • Check the specific reasoning-continuity guidance for whichever provider’s API is in use rather than assuming one policy applies everywhere — as of September 2026, the three named providers describe three different postures.
  • For any tool that stores or forwards conversation state on a user’s behalf, avoid retaining reasoning blocks any longer than the conversation itself needs them.
  • Expect the cross-model boundary to keep moving: it has already narrowed once, for one provider, since the paper’s own testing window, and no single description of an API’s behaviour should be trusted to hold for long.

What to hold loosely

The claim that the demonstrated attacks no longer work rests on the paper’s own account of its follow-up testing, not on a public statement from Anthropic, OpenAI or Google. The recovered-secret counts are counts of well-formed secrets: the paper grades a value by its shape and says explicitly that one claimed to be revoked, expired or “just for testing” still counts, so they are not a tally of keys checked to be working. Whether reasoning blocks already scraped and published before the fix remain decodable today remains an open question as of September 2026.

The bottom line

Encrypting a reasoning block stopped a user from reading their own model’s thoughts, and stopped a rival from scraping them in bulk. It never stopped, and structurally could not stop, a compatible model asked to read them back — because some model always has to. That gap between encrypted and hidden, not a broken cipher, is what let a captured trace travel.

Sources: Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping & Maksym Andriushchenko, “Stealing Reasoning Traces from Proprietary LLM APIs,” arXiv:2608.09867v1, 10 Aug 2026. Matthew Green, “Let’s talk about encrypted reasoning,” A Few Thoughts on Cryptographic Engineering, 29 May 2026. Swati Khandelwal, “OpenAI, Anthropic, Google API Flaw Let Weaker AI Models Decode Stronger Models’ Reasoning,” The Hacker News, 12 Aug 2026. Google, “Gemini thinking: thought signatures,” Gemini API docs. OpenAI, reasoning models guide, OpenAI Platform docs. Anthropic, “Extended thinking,” Claude Platform Docs.

Versions

This document is rewritten when what it says has to change. Every version stays published at its own address.

  1. v0001 current

The current version is also at latest/.