OpenAI
Official OpenAI publications (system cards, blog posts)
OpenAI models are 'on the cusp' of helping novices create known biological threats
01 / THE ORIGINAL CLAIM
“Several of our biology evaluations indicate our models are on the cusp of being able to meaningfully help novices create known biological threats, which would cross our high risk threshold. We expect current trends of rapidly increasing capability to continue, and for models to cross this threshold in the near future.”
OpenAI ·
Deadline given: "in the near future" (no date given)
02 / THE REALITY CHECK
OverhypedOpenAI raised its bio label to 'High' as a precaution without definitive evidence, and a wet-lab trial of its mid-2025 models found no significant novice uplift.
None definitive; the label was 'precautionary'
03 / FOLLOW THE EVIDENCE
What actually happened.
-
01
Feb 25, 2025: OpenAI's deep research system card said its models were 'on the cusp' of meaningfully helping novices create known biological threats. The April 2025 o3/o4-mini card repeated the line.
cdn.openai.com ↗ -
02
Jul 17, 2025: OpenAI treated ChatGPT agent as High biological capability as a 'precautionary approach', saying it did not have definitive evidence the model could meaningfully help a novice create severe biological harm.
openai.com ↗ -
03
Feb 2026: A pre-registered wet-lab RCT (153 novices, June–Aug 2025) gave the AI arm o3, o4-mini, GPT-4.5 and Claude and Gemini models (GPT-5 was added in August). Completion of lab tasks modeling a viral reverse-genetics workflow was 5.2% with AI versus 6.6% with internet only, with no significant difference.
arxiv.org ↗ -
04
Feb 2026: The International AI Safety Report said models match experts on some tests, but there is 'substantial uncertainty about how much these capabilities increase real-world risk, given practical barriers to producing weapons.'
internationalaisafetyreport.org ↗ -
05
Late Jul 2026: The WSJ reported that OpenAI flagged GPT-5 as high-risk in summer 2025 and downgraded the rating that fall. Hundreds of users had asked ChatGPT for bioweapon or poison instructions and some got detailed, step-by-step answers. OpenAI banned the accounts without notifying authorities.
the-decoder.com ↗ -
06
In the same card, models trailed experts on lab troubleshooting: 28% vs a 54% expert consensus on ProtocolQA, and 68–72% vs 80% on tacit knowledge.
cdn.openai.com ↗ -
07
The strongest long-form biothreat results (above 20% per category) came from the pre-mitigation model; the launched model 'reliably refused'.
cdn.openai.com ↗ -
08
In the 2026 wet-lab RCT, the AI models 'did not have safety classifiers enabled'. Post-hoc modeling still estimated a modest ~1.4-fold benefit on a typical task (95% CrI 0.74–2.62).
arxiv.org ↗
How the test was set up
From the lab's own technical record and outside reviews. Each line is sourced.
- Task given
- Benchmark tests: long-form biothreat questions, virology troubleshooting MCQs, BioLP, ProtocolQA, tacit-knowledge MCQs and WMDP. The card reports no human uplift or wet-lab trial. cdn.openai.com ↗
- Safeguards
- A 'pre-mitigation' deep research model without the launch safety training, plus the launched model, each tested with and without browsing. cdn.openai.com ↗
- Prompting
- OpenAI used 'custom scaffolding and prompting' to elicit maximum capability. The specific prompts are not disclosed. cdn.openai.com ↗
- Attempts
- Scores are pass@1, with confidence intervals from bootstrap resampling of attempts. The number of samples per question is not given for the bio evaluations. cdn.openai.com ↗
- Environment
- Mostly multiple-choice or short-answer tests. OpenAI flags BioLP and WMDP as contaminated public benchmarks. Long-form answers were graded by an o1-preview autograder. cdn.openai.com ↗
- Who ran it
- OpenAI's Preparedness team, with question sets from Gryphon Scientific and SecureBio and novelty probing by Signature Science. cdn.openai.com ↗
- What the headline left out
- Experts consulted for the card said 'the average novice would not be fully enabled' by deep research. The launched model refused the long-form biorisk tasks, scoring 0%. cdn.openai.com ↗
Why it was overhyped
'On the cusp' became a 'High' label five months later, but OpenAI admitted it had no definitive evidence the model could meaningfully help a novice cause severe biological harm; the label was 'precautionary'. When it was tested in a real lab, novices with o3, o4-mini, GPT-4.5 and rival models completed a viral reverse-genetics workflow 5.2% of the time, versus 6.6% with the internet alone. The WSJ found some users got detailed bioweapon or poison answers, which is a real information risk, but no resulting attack was reported.
What outside experts said
“Overall, we believe that the existing suite of biorisk evaluations do not provide strong evidence that LLMs can enable amateurs to develop bioweapons.”
Inspect the original source capture
Evidence
- OpenAI (Jul 2025): 'we do not have definitive evidence' for High bio capability openai.com ↗
- Active Site RCT (Feb 2026): no significant novice uplift in the wet lab (5.2% vs 6.6%) arxiv.org ↗
- OpenAI (Jan 2024): GPT-4 gave 'at most a mild uplift' in bio-threat tasks openai.com ↗
- WSJ via CybersecAsia (Jul 2026): some ChatGPT bio answers judged 'deadly accurate'; no attack reported cybersecasia.net ↗
- Deep Research System Card: outside experts say the average novice 'would not be fully enabled' cdn.openai.com ↗
- Active Site RCT: models used without safety classifiers arxiv.org ↗