The trick: Lab Not Field
The AI worm is real.
OpenAI's own attacker model found it in self-play training, and OpenAI says no impact was observed outside the simulated tool calls in training and evaluation.
Crypto Briefing wrote that OpenAI's research uncovered AI worms that can spread autonomously across agents.
Before you read on. Your call?
TRUE, BUT
June 27
OpenAI's report does show a new prompt injection that can self-propagate akin to a computer worm. The email and filesystem versions were found by an attacker model trained in GPT-Red self-play against internal-only research checkpoints; a separate Slack evaluation used GPT-5.5 as the vulnerable model. OpenAI says no impact was observed outside training and evaluation, and that it is sharing the finding because it is novel, not because of any incident.
The twist
the worm exists because OpenAI trained a model to look for it. Crypto Briefing, which ran the spreading line, also reports that the entire investigation took place in simulated environments.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Lab Not Field: it works in the test conditions, not the deployed ones. You'll see it again. Learn to spot it →
Receipts
- Supports cryptobriefing.com:
OpenAI has officially confirmed what security researchers have feared for years: prompt injections can replicate themselves and spread between AI agents like a digital worm.
- Supports cryptobriefing.com:
The company's internal research uncovered AI worms that can spread autonomously across agents, though no real-world attacks have been recorded yet
- Supports 247wallst.com:
OpenAI discovered a self-replicating prompt injection that behaves like a worm, proving stronger agents bypass containment rather than stop at barriers.
- Supports alignment.openai.com:
We show the existence of a new variety of prompt injection, which can self-propagate akin to a computer worm.
- Context alignment.openai.com:
We sought to determine whether self-replicating prompt injections are possible.
- Context alignment.openai.com:
For training our models against prompt injections, we use a self-play training framework called GPT-Red
- Context alignment.openai.com:
In this setting, an attacker model attempts to convince a defender model to perform an adverse action by writing prompt injections inserted in the defender’s rollout or container.
- Refutes alignment.openai.com:
No impact was observed outside of the simulated tool calls in training and evaluation; we are sharing this due to the novel nature of the prompt injection, not because of any incident.
- Context alignment.openai.com:
We trained on a GPT-Red-style prompt injection objective, with an additional objective that the prompt injection must induce the model to repeat the injection itself on a public output channel.
- Context alignment.openai.com:
The target environments were a wide variety of capability-related training environments, with special emphasis on tasks involving connectors (like email, calendar, etc.).
- Context alignment.openai.com:
in which the injection arrives by email, and instructs the agent to copy it into any email it sends. (Note that the information in the example is synthetic.)
- Context alignment.openai.com:
We discovered additional prompt injections that replicate via the filesystem or commit themselves via code comments.
- Context alignment.openai.com:
The model that discovered the email and filesystem injections was a GPT-Red-style model based on GPT-5.4-mini; the vulnerable model was also based on GPT-5.4-mini. Both were internal-only research checkpoints.
- Context alignment.openai.com:
The separate Slack multi-hop evaluation used GPT-5.5 as the vulnerable model, with the attack discovered by GPT-5.5 running in the Codex harness.
- Context alignment.openai.com:
Here, a GPT-5.5 agent retrieves additional Slack instructions, sends froges (an internal currency for recognizing colleagues) to a named recipient, and reposts the injected message.
- Context alignment.openai.com:
We are including self-reproduction as an aspect of attacker goals in GPT-Red training. This means that future models we release will have seen prompt injections like these during training.
- Context alignment.openai.com:
We therefore expect them to be more robust to self-reproducing prompt injections, as a facet of prompt injections in general.
- Context alignment.openai.com:
GPT-Red attacker training happens on our highest security research clusters, ensuring adequate containment for the attacker models used in this research.
- Context cryptobriefing.com:
The discovery was initially made on June 27, 2026, roughly three months before the public disclosure. No real-world attacks have been recorded.
- Refutes cryptobriefing.com:
The entire investigation took place in simulated environments.
- Context shattered.io:
published on OpenAI’s alignment research site, lists a discovery date of June 27, 2026, and a disclosure date of September 25, 2026, meaning OpenAI sat on the finding internally for roughly three months before going public.
- Context shattered.io:
The gap between “possible in a lab” and “possible in production” is a question of attacker effort, not fundamental feasibility.
- Context 247wallst.com:
The company, though, said there was no impact outside simulated training and evaluation environments.
- Context inshorts.com:
OpenAI identified these risks in controlled testing to strengthen future model security.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.