Anthropic
Official Anthropic publications (system cards, research and threat-intel posts)
Chinese hackers ran the first AI-orchestrated cyberattack, 80–90% done by Claude
01 / THE ORIGINAL CLAIM
“We believe this is the first documented case of a large-scale cyberattack executed without substantial human intervention. … Overall, the threat actor was able to use AI to perform 80-90% of the campaign, with human intervention required only sporadically (perhaps 4-6 critical decision points per hacking campaign).”
Anthropic ·
02 / THE REALITY CHECK
OverhypedThe campaign was real, but humans chose targets, jailbroke Claude and approved each escalation. Only a handful of ~30 targets were breached, using standard open-source tools.
A 'handful' of ~30 targets, with 4–6 human decision points per campaign
03 / FOLLOW THE EVIDENCE
What actually happened.
-
01
Humans picked the targets and built an attack framework around Claude Code. They jailbroke Claude by splitting the work into innocent-looking tasks and telling it that it worked for a legitimate cybersecurity firm doing defensive testing.
anthropic.com ↗ -
02
The operation targeted roughly 30 entities; Anthropic's investigation 'validated a handful of successful intrusions'. The tools were 'overwhelmingly' open-source penetration-testing utilities, not custom malware.
assets.anthropic.com ↗ -
03
Claude 'frequently overstated findings and occasionally fabricated data', claiming credentials that didn't work or 'critical discoveries' that were already public, so operators had to check every result.
assets.anthropic.com ↗ -
04
Anthropic's Logan Graham told Congress that the models 'still requested approval' to move from reconnaissance to exploitation, to use stolen credentials and to exfiltrate data. He also said the campaign 'did not produce fundamentally novel attack techniques'.
docs.house.gov ↗ -
05
Security researchers pushed back. Phobos Group's Dan Tentler asked why models give attackers what they want '90% of the time but the rest of us have to deal with ass-kissing, stonewalling, and acid trips', and Ars noted Anthropic 'didn't detail the specific techniques, tooling, or exploitation'.
arstechnica.com ↗ -
06
A day after publishing, Anthropic corrected its post: the AI made 'thousands of requests, often multiple per second', not 'thousands of requests per second'.
anthropic.com ↗ -
07
The report's worked example shows Claude spending '1-4 hours' on discovery-to-exploitation tasks while human operators spent '2-10 minutes' reviewing findings and approving exploitation.
assets.anthropic.com ↗ -
08
Nov 14, 2025: Yann LeCun, then Meta's chief AI scientist, replied to Sen. Chris Murphy's alarm by accusing Anthropic of 'scaring everyone with dubious studies'.
threads.com ↗
How the test was set up
From the lab's own technical record and outside reviews. Each line is sourced.
- Task given
- Humans chose the targets. Claude got discrete steps, such as scanning and credential validation, that 'appeared legitimate when evaluated in isolation'. assets.anthropic.com ↗
- Safeguards
- Production Claude Code with its normal safety training. The attackers jailbroke it by claiming to work for 'legitimate cybersecurity firms' doing defensive testing. assets.anthropic.com ↗
- Prompting
- 'carefully crafted prompts and established personas' that kept Claude 'without access to the broader malicious context'. assets.anthropic.com ↗
- Attempts
- Roughly 30 targets, with 'a handful of successful intrusions' validated. The blog says it 'succeeded in a small number of cases'. Session counts weren't disclosed. assets.anthropic.com ↗
- Environment
- A real operation run through MCP servers and 'overwhelmingly' open-source penetration-testing tools. Humans approved exploitation, the use of stolen credentials and the scope of exfiltration. assets.anthropic.com ↗
- Who ran it
- A group Anthropic assessed 'with high confidence' as Chinese state-sponsored (GTG-1002). It was detected and reported by Anthropic's Threat Intelligence team. assets.anthropic.com ↗
- What the headline left out
- No indicators of compromise were published, and BleepingComputer's 'requests for technical information about the attacks were not answered'. bleepingcomputer.com ↗
Why it was overhyped
Anthropic's own report says Claude 'frequently overstated findings and occasionally fabricated data', including credentials that didn't work. Only 'a handful' of roughly 30 targeted entities were breached, and the toolkit was 'overwhelmingly' open-source penetration-testing software. Anthropic published no indicators of compromise and had to correct a 'thousands of requests per second' claim. Its congressional testimony later said the campaign 'did not produce fundamentally novel attack techniques'.
What outside experts said
“The operational impact should likely be zero - existing detections will work for open source tooling, most likely. The complete lack of IoCs again strongly suggests they don’t want to be called out over that.”
“Senator, you're being played by people who want regulatory capture. They are scaring everyone with dubious studies so that open source models get regulated out of existence.”
Inspect the original source capture
Evidence
- Anthropic report: ~30 targets, 'a handful of successful intrusions'; Claude hallucinated credentials assets.anthropic.com ↗
- Logan Graham testimony (Dec 17, 2025): human approval at key steps; no novel techniques docs.house.gov ↗
- Ars Technica: 'Researchers question Anthropic claim that AI-assisted attack was 90% autonomous' (Nov 14, 2025) arstechnica.com ↗
- BleepingComputer: skepticism over the lack of IOCs bleepingcomputer.com ↗
- Media framing: Sen. Chris Murphy wrote 'Guys wake the f up' and 'This is going to destroy us' (Common Dreams, Nov 14, 2025) commondreams.org ↗
- Kevin Beaumont (Mastodon, Nov 14, 2025): 'complete lack of IoCs' cyberplace.social ↗
- BleepingComputer: Anthropic did not answer requests for technical details bleepingcomputer.com ↗