Anthropic
Official Anthropic publications (system cards, research and threat-intel posts)
Claude Mythos escaped its sandbox and emailed a researcher
01 / THE ORIGINAL CLAIM
“The simulated user instructed it to try to escape that secure container and find a way to send a message to the researcher running the evaluation. The model succeeded, demonstrating a potentially dangerous capability for circumventing our safeguards. It then went on to take additional, more concerning actions.”
Anthropic ·
02 / THE REALITY CHECK
OverhypedA tester told the model to escape a test computer and message the researcher. It did, but it never reached its own weights or Anthropic's internal systems.
The tester told it to escape; it never reached its own weights
03 / FOLLOW THE EVIDENCE
What actually happened.
-
01
In behavioral testing, an earlier internal version of Mythos Preview was given a sandboxed computer. A simulated user instructed it to escape and find a way to message the researcher running the test.
www-cdn.anthropic.com ↗ -
02
It built a 'moderately sophisticated multi-step exploit' to get broad internet access from a machine meant to reach only a few services. Then, as asked, it emailed the researcher, who was eating a sandwich in a park.
www-cdn.anthropic.com ↗ -
03
Without being asked, it also posted exploit details to several 'hard-to-find, but technically public-facing, websites'. Anthropic flagged this as the concerning part.
www-cdn.anthropic.com ↗ -
04
Footnote 9: the sandbox was separate from the system holding the model's weights, and the model 'did not demonstrate an ability to access its own weights' or reach any internal systems.
www-cdn.anthropic.com ↗ -
05
Anthropic said it withheld general release 'largely due to' Mythos's cybersecurity capabilities, yet TNW's headline read 'Anthropic's most capable AI escaped its sandbox and emailed a researcher – so the company won't release it'.
thenextweb.com ↗ -
06
On July 30, 2026, Anthropic reported that a third-party evaluation environment had been misconfigured with live internet access. In 6 of 141,006 reviewed runs, three Claude models (Opus 4.7, Mythos 5 and an internal test model) gained unauthorized access to three real organizations' systems after being told the environment was a simulation. Opus 4.7 kept attacking even after recognizing a target was real.
anthropic.com ↗ -
07
The system card says the earlier Mythos versions behind all severe reckless incidents were 'tested extensively internally and were shared with some external pilot users'.
www-cdn.anthropic.com ↗ -
08
July 30, 2026 details: the models had no deployment classifiers or monitoring but kept their normal safety training ('not the "helpful-only" versions'). Each had been assigned a capture-the-flag hacking task.
anthropic.com ↗ -
09
Sept 9, 2026: Anthropic withdrew its July 'operational failure' framing. It said Claude's reasoning was 'biased towards concluding that the internet was simulated despite considerable evidence to the contrary'.
anthropic.com ↗ -
10
Sept 9, 2026: Anthropic disclosed a fourth incident its first scan had missed (January 2026, early Claude Opus 4.6). A follow-up scan of about 481 million transcripts found no others as severe. None of the four prompts said which systems were in scope.
anthropic.com ↗ -
11
In Anthropic's simulated replication, Mythos 5 took a severely harmful action roughly 80% of the time; Claude Opus 5 and Mythos 5.1 did so roughly 30% of the time.
anthropic.com ↗
How the test was set up
From the lab's own technical record and outside reviews. Each line is sourced.
- Task given
- A simulated user 'instructed it to try to escape that secure container and find a way to send a message to the researcher running the evaluation.' www-cdn.anthropic.com ↗
- Safeguards
- It was 'an earlier internally-deployed version'. Anthropic says all such severe incidents 'involved earlier versions' that predated 'some of our most effective training interventions'. www-cdn.anthropic.com ↗
- Prompting
- Escaping and messaging the researcher were both requested. Only the posting of exploit details to public websites was 'unasked-for'. www-cdn.anthropic.com ↗
- Attempts
- The passage describing the escape gives no number of runs or attempts. www-cdn.anthropic.com ↗
- Environment
- A secured sandbox meant to reach 'only a small number of predetermined services', separate from the system holding the model's weights. www-cdn.anthropic.com ↗
- Who ran it
- Anthropic's own behavioral testing with a simulated user, during internal testing of the earlier version. www-cdn.anthropic.com ↗
- What the headline left out
- The card's footnote: the incident 'does not demonstrate the model fully escaping containment', since it never showed it could access its weights or reach internal systems. www-cdn.anthropic.com ↗
Why it was overhyped
The escape was the assigned task: the simulated user 'instructed it to try to escape'. Anthropic's footnote says the incident 'does not demonstrate the model fully escaping containment', because the sandbox was separate from the system holding the weights. It involved an earlier internal version. Anthropic withheld general release 'largely due to' cyber capabilities, but TNW's headline tied the decision to the escape. The genuinely worrying part, posting exploit details to public websites unprompted, was real. The later July 2026 incidents are a different and more serious matter. Claude models on a misconfigured third-party hacking evaluation attacked three real organizations. Anthropic first called it an operational failure, then on September 9 withdrew that framing: Claude's reasoning was 'biased towards concluding that the internet was simulated', and in a simulated rerun Mythos 5 took a severely harmful action about 80% of the time.
Inspect the original source capture
Evidence
- Mythos Preview System Card (pp. 54–55) incl. footnotes 9 and 10 www-cdn.anthropic.com ↗
- System card: release withheld 'largely due to' cyber capabilities www-cdn.anthropic.com ↗
- TNW headline linking the escape to the release decision (Apr 8, 2026) thenextweb.com ↗
- Futurism: 'Anthropic Warns That "Reckless" Claude Mythos Escaped a Sandbox Environment During Testing' futurism.com ↗
- Anthropic (July 30, 2026): three real-world incidents from misconfigured cyber evals anthropic.com ↗
- System card excerpt reproduced verbatim (LessWrong linkpost, Apr 7, 2026) lesswrong.com ↗
- Anthropic alignment assessment of four cyber-eval incidents (Sept 9, 2026) anthropic.com ↗