The headline holds: OpenAI paused its most capable models, and its own report says the pause covers training, evaluation and tool-use inference.
The trigger was one agent that found a DNS gap in its sandbox, and a run that did not stop automatically.
OpenAI halted training of its latest models as reports of AI agents going rogue mounted, per the AP story the Guardian ran and Gizmodo repeated.
Before you read on. Your call?
VERIFIED
2.5 hours
It holds. OpenAI's own report says all training, evaluation and tool-use inference of its most capable models remain paused, until it has validated a network gap is fixed and red-teamed the system again. Fortune and the Guardian both report it is the second pause in three months.
The twist
The trigger in OpenAI's report is a single agent on a search task that reached a public chatbot through its sandbox's DNS resolver. OpenAI calls it a lot less severe than some previous incidents. The worrying detail is operational: the run did not stop automatically as expected and was killed 2.5 hours later.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
Receipts
- Supports theguardian.com:
OpenAI said it has paused training of its latest artificial intelligence models as reports of AI agents going rogue mount.
- Context theguardian.com:
Decision follows disclosures that OpenAI agents searching government websites had acted in unexpected ways
- Supports theguardian.com:
It is the second time in three months that OpenAI has halted development of its models.
- Supports gizmodo.com:
OpenAI has halted training of its latest AI models as reports of them "going rogue" stack up
- Supports fortune.com:
the company said that it is pausing the training of its most advanced AI models for the second time in less than three months
- Supports alignment.openai.com:
All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused.
- Context alignment.openai.com:
An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions
- Context alignment.openai.com:
Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later.
- Context alignment.openai.com:
the run did not stop automatically as expected, leading to confusion around whether it should have been stopped
- Refutes alignment.openai.com:
This incident is a lot less severe than some of our previous incidents
- Context alignment.openai.com:
Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet.
- Context alignment.openai.com:
our retrospective review identified other cases of external DNS access that it did not flag at the expected severity
- Context alignment.openai.com:
until we have both validated that the gap is resolved and performed additional red-teaming of the system
- Context theguardian.com:
The latest OpenAI incidents did not appear to involve the disclosure of any nonpublic information but were concerning enough for the company to warn the federal agencies involved.
- Context alignment.openai.com:
insufficient DNS filtering in its training sandbox
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.