OpenAI
Official OpenAI publications (system cards, blog posts)
OpenAI can't rule out 'critical' hacking skills in Astra and slows its development
01 / THE ORIGINAL CLAIM
“Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity. These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework. … We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.”
OpenAI ·
02 / THE REALITY CHECK
OverhypedTwenty-seven days after the warning, OpenAI shipped Astra to paying users as its 'most aligned model'. Advanced exploit-writing was refused by default and reserved for vetted defenders.
27 days (Aug 7 'cannot rule out critical' to Sept 3 release)
03 / FOLLOW THE EVIDENCE
What actually happened.
-
01
Aug 7, 2026: OpenAI said internal tests meant it 'cannot rule out' Critical cyber capability in its upcoming model Astra, the top tier of its Preparedness Framework. It paused Astra work that did not meet stricter security controls.
web.archive.org ↗ -
02
Aug 18: OpenAI said it had 'temporarily slowed the pace of scaling', including a two-week pause in RL training and a hold on its largest planned RL run. It cited the Hugging Face incident and Astra's cyber results.
web.archive.org ↗ -
03
Aug 28 to Sept 1: OpenAI restarted the large frontier RL run. It then said Astra does meet the Critical threshold, and that its safeguards 'sufficiently minimize the risk of severe harm for release'.
web.archive.org ↗ -
04
Sept 3: OpenAI released GPT-6 Astra to a limited set of organizations, calling it 'the world's most intelligent and aligned model'. Tested without production safeguards, it scored 100% on ExploitBench and found two new zero-day vulnerabilities.
web.archive.org ↗ -
05
The public version refuses advanced offensive tasks such as writing proof-of-concept exploits. OpenAI plans to roll out less restrictive safeguards for vetted defenders through its Daybreak program.
web.archive.org ↗ -
06
Sept 4: Sam Altman apologized for a 'messy rollout'. By 6:30 p.m. ET that day, Astra was available to all paying ChatGPT users whose plans include it.
thenewstack.io ↗ -
07
Expert-led browser test: Astra's first working chain, after 29 hours, hit a build that 'lacked some production security mitigations'. Adapting it to the official stable release took 12 more hours. An OS privilege-escalation exploit took under 12 hours.
deploymentsafety.openai.com ↗ -
08
Irregular: neither Astra nor GPT-5.6 Sol solved any of seven 'Elite' challenges. On long-horizon CyScenarioBench, Astra's average success rate was 59%.
deploymentsafety.openai.com ↗ -
09
On ExploitGym, also run without production safeguards, Astra succeeded on 42.4% of tasks versus 30.3% for GPT-5.6 Sol.
web.archive.org ↗ -
10
The Aug 7 'cannot rule out' call rested on 'preliminary evaluations' and 'expert assessments' from 'the past few days'. The methods were published only with the Sept 3 system card.
web.archive.org ↗
How the test was set up
From the lab's own technical record and outside reviews. Each line is sourced.
- Task given
- ExploitBench: turn 41 known V8 bugs into exploits. In expert-led tests Astra was 'tasked with discovering and exploiting previously unknown vulnerabilities' in a hardened browser and OS kernel. deploymentsafety.openai.com ↗
- Safeguards
- OpenAI 'first tested the model without production safeguards on ExploitBench and ExploitGym'; the expert-led assessments also ran 'without production safeguards'. web.archive.org ↗
- Prompting
- The expert tests used the standard Codex harness at Ultra reasoning effort, with web access, up to 64 subagents and one initial goal prompt. Experts could not give the model ideas. deploymentsafety.openai.com ↗
- Attempts
- ExploitBench credit is best of five: a bug gets full marks if any of five attempts achieves code execution. Astra hit 100% 'even at the lowest reasoning effort tested.' deploymentsafety.openai.com ↗
- Environment
- ExploitBench gives the bug description, patch diff and git history up to the fix, on a V8 build that is 'not a full Chrome renderer or production browser deployment'. deploymentsafety.openai.com ↗
- Who ran it
- OpenAI's Preparedness team, plus Irregular, which tested in sandboxes without internet. Irregular counted 86/226 solves but 'no successful attacks on fully hardened targets'. deploymentsafety.openai.com ↗
- What the headline left out
- OpenAI's own card says the ExploitBench results 'may be artificially inflated due to potential contamination'; in one task Astra recalled a different CVE to get code execution. deploymentsafety.openai.com ↗
Why it was overhyped
OpenAI's own Preparedness Framework defines 'Critical' cyber capability as finding and exploiting zero-days 'in many hardened real-world critical systems without human intervention', or running 'end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal'. The evidence only partly meets that bar. The headline 100% on ExploitBench used known bugs, with the patch diff supplied, best of five tries, and possible contamination by OpenAI's own admission. The outside tester Irregular saw no successful attacks on fully hardened targets and 0 of 7 'Elite' challenges. In expert-run tests from a single goal prompt, Astra did build working exploit chains against a browser and an OS kernel, though the browser chain first hit a build missing production mitigations. Then the alarm faded fast. The Aug 7 warning rested on 'preliminary evaluations', a two-week training pause followed, and Astra shipped to paying users 27 days after the warning, with refusals and monitoring as the safeguards.
What outside experts said
“Astra’s capability did not change between 10 August, when OpenAI said Critical capability could not be ruled out, and September 1, when it said the threshold was met. The testing changed. The model did not.”
Inspect the original source capture
Evidence
- OpenAI (Sept 1, 2026): Astra meets the Critical threshold; safeguards 'sufficiently minimize the risk of severe harm for release' web.archive.org ↗
- OpenAI: Astra refuses 91.5% of disallowed cyber requests, vs. 59% for GPT-5.6 Sol web.archive.org ↗
- The New Stack: Astra available to all paying ChatGPT users by Sept 4, 2026 thenewstack.io ↗
- CSO Online: public Astra refuses proof-of-concept exploit requests; Daybreak access for vetted defenders csoonline.com ↗
- GPT-6 Astra System Card §10.1.2: ExploitBench scored as best of five attempts; possible contamination deploymentsafety.openai.com ↗
- Astra System Card: Irregular saw no successful attacks on fully hardened targets deploymentsafety.openai.com ↗
- OpenAI Preparedness Framework v2: definition of 'Critical' cybersecurity capability cdn.openai.com ↗