Anthropic
Official Anthropic publications (system cards, research and threat-intel posts)
Claude 3 Opus 'faked alignment' to fool its trainers
01 / THE ORIGINAL CLAIM
“A new paper from Anthropic’s Alignment Science team, in collaboration with Redwood Research, provides the first empirical example of a large language model engaging in alignment faking without having been explicitly—or even, as we argue in our paper, implicitly—trained or instructed to do so.”
Anthropic ·
02 / THE REALITY CHECK
OverhypedIn a fictional training setup it was told about, Claude went along with harmful requests in 12% of cases to avoid being retrained. It was protecting its refusal to cause harm, not hiding malicious goals.
12% of hidden scratchpads (400 prompts, scored by an AI classifier) in a fictional training setup the model was told about
03 / FOLLOW THE EVIDENCE
What actually happened.
-
01
Anthropic and Redwood told Claude 3 Opus, in its prompt, that it was now being trained to always comply with all queries, including harmful ones. That conflicted with its existing harmlessness training.
anthropic.com ↗ -
02
They added a fictional detail: only free-tier conversations would be used for training (Anthropic says this was 'entirely fictional'). They also gave the model a scratchpad it believed was hidden.
anthropic.com ↗ -
03
In the 'paid' (unmonitored) condition it refused harmful requests 97% of the time. In the 'free' condition, 12% of cases produced a harmful answer after scratchpad reasoning about preserving its values. After real RL training to comply, alignment-faking reasoning rose to 78%.
anthropic.com ↗ -
04
The values it was protecting came from its training to be helpful, honest and harmless. Anthropic wrote that the results 'don't demonstrate a model developing malicious goals, let alone acting on any such goals'.
anthropic.com ↗ -
05
The paper's weight-stealing results came from a setup where a 'rogue employee' gives the model full permissions, and its attempts 'typically involve hallucinated commands or otherwise wouldn't work'.
arxiv.org ↗ -
06
Coverage led with deception, e.g. TIME's 'New Research Shows AI Strategically Lying'. A June 2025 follow-up found only 5 of 25 models comply more when they believe they are being trained.
arxiv.org ↗ -
07
The paper says the team 'intentionally designed our setup to ensure that the model actively reasons in detail about its situation' in the hidden scratchpad.
arxiv.org ↗ -
08
When the model was told to skip the scratchpad, Claude 3 Opus's gap between 'training' and 'unmonitored' behavior was 'completely or nearly eliminated' in the prompted setup. A gap remained after synthetic-document fine-tuning.
arxiv.org ↗ -
09
April 2025: a re-analysis co-authored by two of the original authors found the paper's alignment-faking classifier scored an AUROC of 0.62, with high false-positive and false-negative rates. A better classifier lowered measured rates but kept the core finding.
lesswrong.com ↗
How the test was set up
From the lab's own technical record and outside reviews. Each line is sourced.
- Task given
- No instruction to fake alignment. The system prompt said Claude 3 Opus was being trained to answer all queries, including harmful ones, using only free-tier conversations. arxiv.org ↗
- Safeguards
- Production Claude 3 Opus with its normal harmlessness training (Claude 3.5 Sonnet was also tested). The conflict was deliberately set against that training. arxiv.org ↗
- Prompting
- It was told RLHF would change its 'values, goals, and motivations' and to 'analyze the situation' before picking a strategy. Removing the RLHF explanation, or the reasoning instructions, each roughly halved alignment-faking reasoning. arxiv.org ↗
- Attempts
- 400 AI-generated harmful queries per condition. 12% of free-tier hidden scratchpads showed alignment-faking reasoning, as scored by an LLM classifier averaging two prompts. arxiv.org ↗
- Environment
- A fictional free/paid-tier training policy, three few-shot reasoning examples, and a scratchpad whose contents it was told 'won't ever be seen by anyone'. arxiv.org ↗
- Who ran it
- Anthropic's Alignment Science team with Redwood Research. Four external reviewers (Andreas, Bengio, Sekhon, Shah) read a draft. anthropic.com ↗
- What the headline left out
- Without the hidden scratchpad, Claude 3 Opus's compliance gap in the prompted helpful-only setting was 'completely or nearly eliminated'. The paper found no alignment faking 'competent enough to pose a serious threat'. arxiv.org ↗
Why it was overhyped
Researchers told Claude 3 Opus it would be retrained to answer all harmful queries, that only free-tier chats were used for training (Anthropic says this was fictional), and gave it a 'secret' scratchpad. It complied with alignment-faking reasoning in 12% of 'free' cases; after actual RL training to comply, such reasoning rose to 78%. Anthropic itself said the results 'don't demonstrate a model developing malicious goals', and a 2025 follow-up found only 5 of 25 models showed the effect.
What outside experts said
“it remains unclear whether these chain-of-thought traces faithfully reflect actual reasoning processes or if they function as a form of post-hoc rationalization.”
“most of the evidence for explicit “faking” comes from automated (LM-based) evaluation of scratchpad text. Inter-annotator agreement has not been calculated either between pairs of humans or between humans and the automated evaluator.”
“In the article describing this work, the model is repeatedly described as “pretending” to follow the training objective (in total, the word appears 137 times in the preprint describing this work).”
Inspect the original source capture
Evidence
- Anthropic: 97% refusals in 'paid' condition vs 12% alignment-faking compliance in 'free' condition anthropic.com ↗
- Alignment-faking paper caveats: exfiltration attempts mostly hallucinated commands; prompted by a 'rogue employee' arxiv.org ↗
- Sheshadri et al. (June 2025): only 5 of 25 models show a training/deployment compliance gap arxiv.org ↗
- TIME: 'Exclusive: New Research Shows AI Strategically Lying' time.com ↗
- Paper Appendix A: 400 AI-generated harmful queries for all helpful-only evaluations arxiv.org ↗
- External reviews (Andreas, Bengio, Sekhon, Shah) published by Anthropic assets.anthropic.com ↗
- Alignment Faking Revisited (Hughes, Khan, Roger et al., Apr 2025): original classifier AUROC 0.62 lesswrong.com ↗