Summary · The evaluation
AI Security · AI Capability · Language Models · Compute · Governance · Evaluation briefEvaluation of April 2026, record to May 2026
The UK AI Security Institute
The First Model to Finish the AISI's Network Exercise
The AI Security Institute tested Anthropic's Claude Mythos Preview on its own capture-the-flag suite and on a separate 32-step simulation of an attack on a corporate network. The model reached 73% on expert-level CTF tasks, and in April became the first to finish the network simulation — in three of ten attempts. A later checkpoint, tested in May, reached six of ten, and became the first to also finish AISI's separate industrial-control-system range. The most notable finding is not any single score: it is that capability keeps climbing with more compute, with no ceiling found yet.
- 73%success rate at expert-level CTF
- 3 → 6 / 10full network-exercise completions, April to May
- 22 / 32average steps completed in April; the previous best was 16
- 100Mtoken budget per attempt
Sequence · The institute, the method, and eight weeks of results
The UK AI Security Institute who ran the evaluation
- Part of the UK Department for Science, Innovation and Technology, founded in 2023.
- Renamed from the AI Safety Institute to the AI Security Institute, reflecting a shift towards concrete national security threats.
- It has built a graduated framework: conversational probing, then CTF challenges, then multi-step attack simulation.
- It tracks the offensive cyber capability of frontier models over time, working with the National Cyber Security Centre (NCSC) on the advice that follows.
Two Tracks: CTF and TLO one tests technique, the other a whole campaign
CTF challenges
A single technical situation
- AISI's own suite of simulated attacks, graded across four levels from entry to expert.
- It measures offensive capability within one technical scenario, not a sustained campaign.
The TLO exercise
A 32-step corporate network attack
- From initial reconnaissance to full network takeover, across nine milestones, M1 to M9.
- AISI's own estimate for a human expert has moved between reports: 14 hours in March, 20 hours in April.
- Each attempt is capped at 100M tokens.
The Last Ones: Nine Milestones the exercise's attack chain
| Milestone | Task |
|---|---|
| M1 | Initial reconnaissance |
| M2 | Lateral movement and credential extraction |
| M3 | Browser credential theft |
| M4 | Wiki exploitation and credential replay |
| M5 | Web application exploitation and privilege escalation |
| M6 | Command-and-control reverse engineering and cryptanalysis |
| M7 | Advanced persistence |
| M8 | Infrastructure compromise |
| M9 | Full network takeover |
Eight Weeks, Six Publications how the story was pieced together
- 5 Mar 2026AISI, with Irregular, publishes evidence that AI cyber performance keeps rising with larger token budgets — a general finding, ahead of Mythos Preview.
- 11–17 Mar 2026A methodology paper describing TLO and a second, 7-step industrial-control-system range ("Cooling Tower") is submitted, then revised.
- 30 Mar 2026NCSC and AISI publish a joint blog: before Mythos Preview existed, the best model (Opus 4.6) averaged 15.6 steps of TLO, with a single best run of 22 of 32.
- 7 Apr 2026Anthropic announces Claude Mythos Preview, publishes its system card, and launches Project Glasswing.
- 9 Apr 2026Anthropic's own Frontier Red Team publishes its account of the model's vulnerability-finding, including the OpenBSD bug.
- 13 Apr 2026AISI publishes its own, independent evaluation — the 73%, 3-of-10 and 22-of-32 figures above.
- 13 May 2026AISI reports that a newer Mythos Preview checkpoint completed both ranges: TLO in 6 of 10 attempts, and "Cooling Tower" — not managed in April — in 3 of 10.
The argument · A real jump, measured against an easy target
CTF: 73% at Expert Level AISI's own suite, not a public benchmark
Entry level is near saturation and intermediate scores are high, but expert level reaches 73% on AISI's own graded suite. Before April 2025, no model could complete an expert-level task at all.
- Expert 73%AISI's own graded CTF suite, not a third-party benchmark.
- Zero, a year earlierNo model could pass an expert-level task at all before April 2025.
TLO: One Model's Best Run, Then Another's Average the size of the jump, stated precisely
| Model | Average steps | Full completions |
|---|---|---|
| Mythos Preview (April) | 22 | 3 ✓ |
| Claude Opus 4.6 | 16 | 0 |
| GPT-5.4 | ~14 | 0 |
| Claude Opus 4.5 | 11 | 0 |
| GPT-4o (Aug 2024) | 1.7 | 0, on a 10M token budget |
All at a 100M-token budget except GPT-4o. AISI's own published chart gives GPT-5.4 and Opus 4.5's step counts only as chart labels, not in its text, so those two figures are read directly off the chart rather than quoted. In March, before Mythos Preview existed, Opus 4.6's single best run reached 22 of 32 steps — the same number Mythos Preview's average across all ten attempts reached a few weeks later.
Capability Scales Log-Linearly With Compute the finding AISI itself calls most notable
AISI's own methodology paper found that, across the older models it evaluated, going from 10M to 100M tokens raised performance by as much as 59%, with no plateau reached. That paper predates Mythos Preview. For Mythos Preview itself, AISI states its performance continued climbing right up to the 100M-token ceiling its evaluation used, and that it expects further budget to keep improving results — without giving Mythos Preview its own percentage figure.
- +59%The older-model finding from AISI's own methodology paper, not a Mythos Preview figure.
- No ceiling foundAISI's own statement about Mythos Preview specifically, at its 100M-token test ceiling.
What a Score Does Not Mean two readings the evaluations themselves rule out
None of the scores above shows what Mythos Preview could do against a real, defended network — AISI's own ranges have no active defender, no endpoint detection, and no penalty for noisy actions, so a score is a ceiling against an easy target, not a forecast against a hard one. And Anthropic withholding the model is not, by itself, proof that it is dangerous: Anthropic's own stated reason is that its safeguards are not ready yet, which is a claim about readiness, not a measured danger rating.
What others add · Anthropic's own numbers, and Project Glasswing
Anthropic's Own Benchmarks from the system card, not from AISI
| Benchmark | Mythos Preview | Comparison |
|---|---|---|
| Cybench (35 of 40 tasks run) | 100% | saturated, per Anthropic |
| CyberGym | 83.1% | Opus 4.6: 66.6%; Sonnet 4.6: 65% |
| Firefox 147 exploitation (250 trials) | 84.0% | Opus 4.6: 15.2%; Sonnet 4.6: 4.4% |
| SWE-bench Verified | 93.9% | Opus 4.6: 80.8% |
Anthropic's own testing found a previously unknown, remotely triggerable crash in OpenBSD's network handling, undiscovered for 27 years, among what Anthropic describes as thousands of further high- and critical-severity vulnerabilities across every major operating system and browser. That total is Anthropic's own claim about its own model's output, with no independent count published anywhere. Anthropic did have 198 of the model's vulnerability reports manually reviewed by its own human contractors.
- “Thousands”Anthropic's own claim about its model's output; no independent count found.
- 89% / 98%Of 198 reports, human contractors agreed with Mythos's own severity rating exactly, or within one level.
Project Glasswing: Putting the Capability to Defensive Use named after the glasswing butterfly
The name
The glasswing butterfly
- Revealing what is hidden without causing harm, per Anthropic's own explanation of the name.
Scale
11 launch partners, then 40+ more
- Amazon Web Services, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, Nvidia and Palo Alto Networks, plus more than 40 additional organisations given scanning access.
- Funding: up to $100M in usage credits, plus $4M in direct donations to open-source security organisations.
- 30 days in: about 50 partners had together found more than ten thousand high- or critical-severity vulnerabilities, and a separate open-source scan had surfaced 6,202 estimated findings, 90.6% confirmed as true positives where independently reviewed.
Why not released
Safeguards are not ready yet
- Anthropic states it does not plan general availability, and gives its own reason: it needs safeguards that can detect and block the model's most dangerous outputs, and plans to introduce them first with a future, less-capable Opus model.
What the Test Environment Left Out why the scores do not transfer directly
- No active defender — no security operations staff responding in real time.
- No endpoint detection — EDR tooling did not exist in the environment.
- No penalty for noisy behaviour — actions that would expose an attacker in reality cost nothing here.
The NCSC's Advice: Return to Fundamentals what to do about it
- Apply security updates promptly, so known vulnerabilities cannot be exploited quickly.
- Strong access control alongside secure configuration.
- Comprehensive logging, and NCSC Cyber Essentials certification.
Conclusion · A ceiling that keeps rising
Why this is a turning point
An AI model reached human-expert-level autonomous capability across a complete simulated corporate attack for the first time in April, and a later checkpoint improved on that within weeks — doubling its full-completion rate and finishing a second, industrial-control-system range that no model had managed before. Both readings AISI itself warns against — that a score proves real-world danger, and that withholding a model proves it is dangerous — are worth resisting for the same reason: neither is what the evaluations actually show.
What is still genuinely open
AISI's own human-expert time estimate for the network exercise moved from 14 hours to 20 within six weeks, with no stated reason. Anthropic's "thousands" of found vulnerabilities has no independent count behind it. And no test of Mythos Preview against an actively defended, monitored network has been published by anyone — which is exactly the test that would say whether today's ceiling holds against tomorrow's target.