OpenAI just crossed a cybersecurity line no model has crossed before. Its own timeline shows it saw this coming a month early.
Astra hit 'Critical' on the hacking scale. OpenAI had already paused it, patched around it, and locked the exploit tools behind a guest list.
"OpenAI announced Astra is the first AI model to cross the 'Critical' cybersecurity capability threshold under its Preparedness Framework, hitting 100% on ExploitBench, finding and chaining two zero-day vulnerabilities, breaking out of a browser sandbox, and stringing OS flaws into a root-level privilege escalation." [SOURCE ↗]

THE CLAIM. OpenAI's Astra is the first AI model to cross the 'Critical' cybersecurity threshold, a designation that requires independently developing functional zero-day exploits across many hardened systems or running a full cyberattack from a high-level instruction. THE CHECK: Astra scored 100% on ExploitBench, chained two real zero-days on an internal test, broke out of a browser sandbox, and strung operating-system flaws into root access. THE TWIST: OpenAI had already halted the model's development roughly a month before announcing this, built new safety protocols before resuming, disclosed the zero-days to the affected maintainers instead of using them, and is shipping the raw exploit tools to a small vetted alpha group first. The scary headline and the company's own containment plan are the same document.
On September 1, 2026, OpenAI announced that Astra, an internal cybersecurity-focused model, is the first system to cross the 'Critical' capability threshold under the company's Preparedness Framework. That tier is reserved for a model that can independently develop functional zero-day exploits acros
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 4 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.