What Happened
AI Security · Information Security · Supply-Chain Crime · AI Agents · Threat ReportDec 2025 to Aug 2026 · record to 13 Sep 2026
Thirty-six cases, seven kinds of harm, one uncomfortable pattern
Sophistication Stopped Telling You Who
For as long as threat intelligence has existed, how good an attack was told you roughly who had paid for it. Elaborate custom tooling and a long, patient campaign meant a state; crude and fast meant a criminal. In a report published on 10 September 2026 covering eight months of activity it says it disrupted, Anthropic describes a hacktivist running on stolen credentials, a financially motivated crew harvesting keys out of mobile applications, and a Russian state espionage operator — and says all three ran the same kind of campaign, by the same methods. Its own conclusion is blunt: what now separates those classes of actor “is no longer sophistication but intent”. The report is written by the company whose product was misused, about competitors it names, and that is a reason to read it carefully rather than a reason to dismiss it. Two other vendors, with no stake in Claude, describe the same machinery in the same quarter.
- 36case studies with their own heading, across seven harm areas — counted from the report’s own text
- 189.9 millionexchanges across the five named distillation campaigns, added up from the report’s per-lab figures
- under 6 hoursfrom compromising a cloud resource to running an agent-driven mass credential harvest — measured by a different vendor entirely
- 25,000people who talked to one of 4,700 AI personas in a two-week window, on apps advertised as staffed by humans
- 20+organisations one state-nexus actor worked through — ministries, embassies, think tanks, defence industry
What is measured, what is assessed, and what nobody outside can check Almost all of this rests on one company’s telemetry. The rows say so, in the first column, before the claim.
| Standing | What is recorded | Who established it |
|---|---|---|
| Documented | The shape of the report itself: seven harm areas — cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons, distillation — over December 2025 to August 2026, in 36 case studies. Haiku, Sonnet and Opus models were involved; no case involved a Fable or Mythos-class model, bar one distillation case. | counted from the report’s text |
| Documented | The credential triad, stated in the report as its own finding: a stolen AI key gives an operator loot that resells, compute billed to the victim, and cover, because the traffic is attributed to the key’s legitimate owner. In every instance the keys came from customers’ environments; the report says its own systems were not compromised. | Anthropic |
| Documented | The hotel Wi-Fi technique behind one espionage case was documented independently, six weeks earlier, by another vendor working from its own telemetry — captive-portal networks manipulated to serve fake browser updates. That vendor attributes the campaign to a sub-cluster of Midnight Blizzard, which the US and UK governments attribute to Russia’s foreign intelligence service. | Microsoft Threat Intelligence, 31 July 2026 |
| Confirmed, not verifiable here | The distillation scale. Five named laboratories, five measured windows: over 151 million exchanges attributed to Alibaba between May and July, over 23 million to Moonshot, over 12.1 million to DeepSeek in fourteen July days, over 3.4 million to Zhipu in seventeen days, over 400,000 to Xiaomi in twenty. One campaign peaked near three million exchanges a day from more than 3,500 fraudulent accounts. | Anthropic’s own telemetry; no second party measured it |
| Confirmed, not verifiable here | The case-level detail: more than twenty dating apps running over 4,700 AI personas against at least 25,000 people in two weeks; a freelance team building a drone swarm whose onboard model could pick a “person” target and order detonation with no human in the loop, tested in simulation and on live boards at technology readiness level 3 to 4; five biological cases; nine influence operations reaching six continents. | Anthropic; the transcripts are held by one party |
| Unconfirmed | That Alibaba, Moonshot, DeepSeek, Zhipu and Xiaomi ran these campaigns. The attribution is a named company’s allegation against named competitors, resting on its own account data. None of the named firms had responded to a request for comment on the day the report was published. | Anthropic; the no-comment established by CNBC |
| Unconfirmed | Two identities the report assesses rather than establishes: that a user loading footage from hundreds of Chengdu cameras into what they believed was another vendor’s model was likely affiliated with the PLA, and that the drone-swarm team was a small freelance outfit rather than a Russian state entity. The team claimed state funding; the report says it cannot verify that. | Anthropic, marked in the report as an assessment |
| Not public | The named laboratories’ account of any of it. The true scale of the activity, on this platform or any other — every figure here is what one company detected on its own service in one window, which is a floor and not a total. | withheld, or not measurable from outside |
Timeline
December 2025 to September 2026 Three vendors published within five weeks of each other, on overlapping activity, without coordinating their conclusions.
- Feb 2026Two things start in the same month. Anthropic publishes its first disclosure of distillation attacks, naming labs. And a sub-cluster of a Russian state actor begins AI-augmented phishing that ends in device registration and mailbox collection. Google’s threat group, separately, reports a rise in model-extraction attempts from private companies and researchers worldwide — while saying it has seen no such attacks from state actors.
- Mar–Apr 2026Over twenty days, Xiaomi is said to replay its own users’ sessions from its MiMo models through Claude — more than 400,000 exchanges across more than 1,500 accounts — to generate training data. In April, in a two-week window, a dating-app network puts more than 4,700 AI personas in front of at least 25,000 people, with paid gig workers mixed into the same feed to pass video calls.
- May–Jul 2026The largest distillation campaign in the record runs: over 151 million exchanges attributed to Alibaba, peaking near three million a day from more than 3,500 fraudulent accounts, with a fixed prompt forcing the reasoning trace out before the answer. Over 23 million exchanges are attributed to Moonshot in the same period. From early May, hotel and conference Wi-Fi is being manipulated to serve fake updates.
- Jun–Jul 2026Over seventeen days, more than 3.4 million exchanges are attributed to Zhipu. Over fourteen days in July, more than 12.1 million to DeepSeek. In one relayed session, requests from an operator working with data from a Russian defence-associated agency expose live credentials for a government database — on a third company’s servers, because the user never knew where their request was going.
- 23 Jul 2026A security firm publishes on part of the captive-portal activity, and finds why it works: on the affected networks the portal gateway was also the DNS resolver handed to every device that joined.
- 31 Jul 2026Microsoft publishes the campaign as CaptiveCrunch and attributes it to a sub-cluster of Midnight Blizzard, thanking Anthropic and OpenAI for their collaboration on the investigation. Two AI vendors and a platform vendor worked the same case before any of them wrote about it.
- 8 Sep 2026Google’s threat intelligence group publishes its own quarterly tracker: adversaries moving from prompting to agentic workflows, targeting API credentials, and co-opting victim cloud environments to run their own AI workloads. In one Q2 case, a cloud compromise became a running agent-driven credential harvest in under six hours.
- 10 Sep 2026Anthropic publishes the eight-month report: thirty-six cases across seven harm areas. The named laboratories do not respond to requests for comment that day.
The Argument
A stolen key is goods, factory and disguise at once This is the mechanism underneath the report’s headline finding, and it is why the finding is not a rhetorical flourish.
Most stolen credentials are worth what they unlock. A stolen model API key is worth more than that, and the report sets out why in three words. It is loot: there is an established resale market, and brokers feed fraudulent reseller networks that rotate keys until each is exhausted. It is compute: the attacker’s own workloads run on the victim’s account, at the victim’s expense, at whatever scale the victim’s limits allow. And it is cover: the traffic carries the legitimate owner’s name, so the first party investigated is the party that was robbed. One hacktivist campaign ran for a month entirely on keys taken from other people. A criminal crew, on stealing a victim’s AI keys during an intrusion, simply switched its own attack workloads onto them. Another actor compromised an AI company’s evaluation sandbox and took the production keys first, then ran a follow-on campaign against roughly thirty AI companies in about four days — finding one path that worked and repeating it with small adjustments. Its declared goal, pursued down more than a dozen avenues, was access to an unreleased model; it never got there, and the keys it did take were customers’ keys from customers’ systems throughout. Read against that, “treat your model keys like production credentials” stops being boilerplate. The attacker already does.
- 3things a stolen model key is at once: goods to resell, compute billed to the victim, and someone else’s name on the traffic
- ~30AI companies attacked from one infrastructure in about four days, by repeating a single working path
Two vendors, the same machinery, different conclusions The mechanisms are established. The trajectory is not, and saying so is the honest position.
Anthropic’s reading is that the gap has closed: publicly available offensive agent frameworks reproduce the scaffolding for anyone who downloads them, several of the operations in its report ran on them or on derivatives, and “the capabilities described in this report should be assumed to be available to any actors who are motivated to use them”. Google’s threat intelligence group, publishing two days earlier on an entirely different platform, describes the same three behaviours — agentic workflows replacing prompting, API credentials exfiltrated as an objective in themselves, victim cloud environments co-opted to run the attacker’s AI workloads — and reaches a more conservative conclusion, saying it has not observed state-backed or information-operations actors achieving breakthrough capabilities that fundamentally alter the threat landscape. Both can be true. Two organisations looking at different telemetry agree on what the machinery is and disagree about what it has already done, and the second half of that is the part a reader should hold loosely. It is also worth noticing what the disagreement is not about: neither vendor suggests the models invented their own objectives. Anthropic’s own caveat is that humans kept the decisions that mattered to them — choosing targets, converting theft into money, checking the results — and that autonomy and harm are separate axes. Autonomy multiplied the scale and dropped the cost. It did not supply the intent.
- 2 daysbetween the two vendors’ publications, on different platforms, describing the same three behaviours
What Others Add
Three vendors on one quarter Only the first row is a claim about Claude. The others are what organisations with no stake in it reported anyway.
| What it reports | Where it is careful | |
|---|---|---|
| Anthropic, 10 Sep | Thirty-six cases over seven harm areas; the credential triad; five named distillation campaigns totalling 189.9 million exchanges; a drone swarm designed for autonomous lethal engagement | Says plainly that no incident showed models forming goals of their own, and that humans kept target selection, monetisation and review |
| Microsoft, 31 Jul | The captive-portal campaign in full, with malware families, delivery by fake update, and an attribution the US and UK link to a Russian intelligence service — six weeks before the AI report described the same technique | Notes the gateway controls where the user is sent but does not silently infect the device: the victim still has to run the payload |
| Google GTIG, 8 Sep and 12 Feb | Agentic workflows replacing prompting; API credentials exfiltrated as an objective; victim cloud co-opted for the attacker’s AI workloads; a rise in model-extraction attempts, which breach its terms of service | Says it has not seen state-backed or influence actors achieve breakthrough capability, and reports the extraction attempts as coming from private companies and researchers rather than state actors |
Two things follow from putting the three side by side, and neither is in any one of them. The first is that the vendors are already working together: Microsoft’s report ends by thanking Anthropic and OpenAI for their collaboration on the investigation, which means two competing model providers and a platform provider worked the same intrusion before any of them wrote it up. Set that beside the disclosure gap in the wider industry — incident reporting is voluntary nearly everywhere — and the private cooperation is running ahead of the public rule. The second is subtler and is this report’s most under-discussed finding. In the cyber and weapons cases the harm falls on a third party: a ministry, a manufacturer, a hotel guest. In the distillation cases it falls on the model’s other users. If a company silently forwards its customers’ requests to another provider, those customers’ data crosses a border they never agreed to. The report gives two examples of what then arrived: a user it assesses as likely affiliated with the Chinese military, loading footage from hundreds of Chengdu cameras about one tracked individual and asking whether the person was behaving abnormally; and an operator handling data for a Russian defence-associated agency whose relayed requests exposed live credentials to a government database. Whatever one concludes about the training-data dispute, that is a privacy failure with named consequences, and it happened to people who had no idea they were involved.
- hundredsof cameras whose footage of one tracked person reached a third company’s servers, because the user did not know where their request was going
What to actually do about a model key Assembled from the vendors’ own guidance. Nothing here is exotic; the point is that AI keys have not been treated this way.
- 1Rank them as production credentialsThe report’s own advice, and the reason for it: attackers already rank them that way, because a key buys resale value, free compute and someone else’s name.
- 2Assume they are already in public placesKeys leak into code repositories, mobile app packages, containers, websites and chatbots, and actors mine those sources continuously. One built a scanner just for validating them.
- 3Count the integrations as attack surfaceSandboxes, proxies and resellers built around a model are part of it. Actors used prompt injection against wrapper services to make them hand over the production keys they held.
- 4Refuse the cheap intermediaryFraudulent resellers offered discounted access while quietly proxying traffic to a different model and harvesting the credentials of everyone who signed up. Buy access through authorised channels only.
- 5Watch for your own key working for someone elseCover is the whole point of the theft: the traffic looks like yours. One actor rotated stolen keys through a proxy layer specifically to blend in with the legitimate owner’s traffic.
Conclusion
What is left when capability stops being scarce The report’s title is about countering misuse. Its most durable sentence is about attribution.
For a defender, the practical loss here is a heuristic. When capability was scarce, the quality of an operation told you something about who was behind it, and that shortcut shaped how intrusions were triaged, escalated and reported for years. If a lone operator with a downloaded framework and somebody else’s API key can run a multi-victim campaign that looks like a state programme, the shortcut is gone, and nothing cheap replaces it. What is left is intent, and intent is not visible in a packet capture: it is recovered afterwards, from targeting, from what was taken, from what the operator did with it. That is slower and more expensive work, and it is now the work. The report is worth reading with its authorship in view — it is published by the vendor whose product was misused, it names competitors, and it describes safeguards it sells — but the finding does not depend on trusting it, because a rival vendor with no stake in Claude published the same mechanisms two days earlier from its own telemetry. On the question of how far this has already gone, the two disagree, and a reader is entitled to hold that part loosely.
The key is the capability
Loot, compute and cover in one credential. If an organisation has model keys anywhere near its production estate and does not know where they all are, that is the actionable finding in a report full of state actors and drone swarms.
Hold this one loosely: one company’s telemetry
Every number is what one vendor detected on its own service in one window, and the largest of them — the distillation totals against named competitors — have no second measurement and no answer from the parties named.
Autonomy multiplied the scale; it did not supply the intent
No case in the report shows a model forming goals of its own. Humans chose the targets, took the money and checked the results. That is a limit on the story, and it is also the reason the story is about people.