Tratopedia
Side N繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.177.0

The release this page was built from. It is what the service worker caches under.

LLM Engineering · Language Models · Software Engineering · AI Agents · Engineering PostmortemAnthropic Engineering Blog, Apr 2026 · record to May 2026

Resolved in v2.1.116 · the API itself was unaffected

Three Unrelated Changes That Looked Like One Regression

A lowered default reasoning effort; a caching bug that kept clearing reasoning history; a system-prompt rule capping output length. The three touched different slices of traffic at different times, and together they presented as a broad, inconsistent drop in quality. Claude Code, the Agent SDK and Cowork were affected; the API itself was not. An AMD engineer's own data audit reached similar conclusions three weeks before Anthropic said anything in public.

  • 3separate problems, all fixed by 20 April
  • 34days for the longest-running of them
  • −3%drop on the one evaluation that caught the third
  • v2.1.116the release in which all three were resolved
StandingEventDetail
ConfirmedThree unrelated product changes, 4 March – 20 AprilA lowered reasoning-effort default, a caching bug that erased thinking history, and a verbosity-capping system prompt — each hit a different slice of traffic on its own schedule. All three resolved in v2.1.116; usage limits reset for all subscribers on 23 April.
ConfirmedAn AMD engineer's public data audit, 3 AprilStella Laurenzo, director of AMD's AI group, analysed 6,852 Claude Code sessions and filed a GitHub issue — three weeks before Anthropic's own account.
Confirmed but not verifiable hereThe audit's own theory of the causeLaurenzo's data correlates the decline with a mid-February setting that hides reasoning from the interface. Anthropic's three-cause account does not name this setting, and Anthropic's first public response (7 April) said it does not reduce reasoning at all.
UnconfirmedThat the real driver was compute rationing, not the stated tradeoffsRaised by some users and, separately, by a rival executive. Anthropic's postmortem does not address compute at all, though a separate statement to Fortune acknowledged demand had strained infrastructure.

Timeline thirteen dated events across three months

  1. 02Opus 4.6 ships in Claude Code with reasoning effort defaulted to high.
  2. 02-12The setting named redact-thinking-2026-02-12 hides reasoning from the interface. An outside audit later names it as a possible cause; Anthropic's own account never does.
  3. 03-04The default reasoning effort drops to medium (Problem One begins).
  4. 03-26The caching bug ships (Problem Two begins).
  5. 04-03Stella Laurenzo, AMD's AI director, files a GitHub issue analysing 6,852 sessions.
  6. 04-06The Register reports the issue.
  7. 04-07Problem One is fixed; the default returns to high (xhigh for Opus 4.7). Anthropic's Boris Cherny publicly disputes Laurenzo's theory.
  8. 04-10Problem Two is fixed (v2.1.101).
  9. 04-16A verbosity-capping system prompt ships alongside Opus 4.7 (Problem Three begins).
  10. 04-20Problem Three is fixed; all three resolved (v2.1.116).
  11. 04-23Anthropic publishes the postmortem and resets usage limits for all subscribers; Boris Cherny elaborates the caching mechanism on Hacker News the same day.
  12. 04-24Fortune reports user backlash, cancellations, and an Anthropic statement citing infrastructure strain.
  13. 05-14InfoQ publishes a synthesis, including a Reddit-reported concern the postmortem does not address.

One Explanation, Two Accounts attributed by name, and left open where they disagree

Anthropic's own diagnosis: each of the three changes passed its own review individually — human review, automated tests, evaluations. The failure sat between them: no process was responsible for asking what several simultaneous changes would do together, and each touched a different slice of traffic on a different schedule, so in aggregate they looked like one broad, inconsistent decline rather than three separate causes.
What complicates it: an independent, data-driven complaint reached the public three weeks before Anthropic's own account did, and named a different change as the correlate — one Anthropic's three-cause account does not mention, and one Anthropic's own first public response had already argued against.

  • 3 AprilStella Laurenzo files a data-backed complaint naming a mid-February change.
  • 23 AprilAnthropic's own postmortem names three different changes, none from February.
SourceWhat it names as the cause
Anthropic, 23 April postmortemThree product-layer changes: a lowered reasoning-effort default, a caching bug, a verbosity-capping prompt — no mention of compute or of the mid-February setting.
Anthropic, statement to FortuneInfrastructure “stretched” by demand growth, in general terms — not tied to any one of the three named bugs.
Stella Laurenzo, GitHub audit (3 Apr)The setting redact-thinking-2026-02-12, which hides reasoning from the interface, correlated with a measured shift toward shallower, edit-first behaviour. The two reports of the audit differ on when that setting reached Claude Code: TechRadar dates it by the setting's own name, The Register to an early-March deployment in v2.1.69.
Anthropic, first public response (7 Apr)That same setting hides reasoning from the UI only and does not reduce reasoning; credits Opus 4.6's adaptive thinking instead.

What Others Add depth from outside Anthropic's own account

  • On Hacker News, Boris Cherny of the Claude Code team filled in the mechanism Anthropic's blog only summarised: ordinarily all but a conversation's latest message hit the prompt cache, but after an hour idle the next message is a full cache miss, and in an extreme case a 900,000-token session left idle for an hour could write over 900,000 tokens to cache at once — “a significant % of your rate limits, especially for Pro users”. Eliding old thinking after an idle period was the fix Anthropic tried for that cost, and its implementation is what introduced Problem Two.
  • Anthropic's first public response to the AMD audit, reported by TechRadar on 7 April, disputed its cause rather than confirming it: Boris Cherny said the mid-February setting only withholds reasoning from the interface and does not reduce it, and pointed to Opus 4.6's adaptive thinking as a separate improvement. Anthropic's 23 April postmortem does not return to this exchange or name the mid-February setting at all.
  • Fortune reported that some users described feeling “gaslit” by Anthropic's communications in the weeks before 23 April, which they read as denying any problem existed, and that some say they have cancelled their subscriptions. In a statement to Fortune, Anthropic cited demand having “stretched” its infrastructure “particularly at peak hours” — a different register of explanation to the blog's three specific bugs, answering a question the postmortem itself does not raise.
  • In an internal memo first reported by CNBC and cited by Fortune, OpenAI's revenue chief called Anthropic's failure to secure enough compute a “strategic misstep”. That is a competitor's characterisation in a leaked memo, not a finding, and Anthropic's own postmortem does not attribute any of its three named bugs to compute at all.
  • InfoQ's coverage relayed a concern from Reddit that the postmortem does not address: Claude Code can delegate work to the cheaper Haiku model without making that visible except in verbose logging — a risk InfoQ noted is hardest to catch in unattended pipelines, where a quality drop surfaces only steps later rather than at the moment it happens.

What Changes Now five follow-up measures, from Anthropic's own account

  • Have a larger share of internal staff run the exact public build of Claude Code, rather than a separate test build.
  • Run a broad suite of per-model evaluations for every system-prompt change, continuing the ablation analysis that eventually caught Problem Three.
  • Build tooling that makes system-prompt changes easier to review and audit, and add guidance to CLAUDE.md so model-specific changes are gated to the model they target.
  • Give any change that could trade off against intelligence a soak period, a broader evaluation suite, and a gradual rollout, so issues are caught earlier.
  • Widen its Code Review tool's repository context — the change that let Opus 4.7 catch Problem Two's offending pull request where Opus 4.6 had not — and use @ClaudeDevs on X and a centralised GitHub thread to explain product decisions as they are made.

The lesson underneath, and what's still open

Anthropic's own account of the failure is structural: each change passed its own review, and nothing was responsible for testing what several changes would do at once — which is exactly what the most substantive follow-up measure now targets. Two things the postmortem does not settle: whether the mid-February change an outside audit pointed to played any real part in the decline, and whether Claude Code's quiet delegation to a smaller model is a comparable risk. Both are attributed to named sources and left open here, not folded into the three causes Anthropic named.

Sources: Anthropic, “An update on recent Claude Code quality reports”, 23 Apr 2026 · Beatrice Nolan, Fortune, 24 Apr 2026 · Steef-Jan Wiggers, InfoQ, 14 May 2026 · Brandon Vigliarolo, The Register, 6 Apr 2026 · Craig Hale, TechRadar, 7 Apr 2026 · Boris Cherny, Hacker News, 23 Apr 2026. Dates follow each source as published.

Versions

This document is rewritten when what it says has to change. Every version stays published at its own address.

  1. v0001 current

The current version is also at latest/.