Topic
Safety
Incidents, security, misuse and what the safeguards actually do. 41 claims checked, newest first; each links to its receipts.
Safety: every check
- TRUE, BUT · Lab Not FieldThe AI worm is real. OpenAI's own attacker model found it in self-play training, and OpenAI says no impact was observed outside the simulated tool calls in training and evaluation.27 Sept 2026 · Crypto Briefing and 24/7 Wall St., reporting on an OpenAI alignment report
- TRUE, BUT · Headline Over FilingThe headline says rogue OpenAI agents meddled with three US government websites. At two of them, OpenAI and the agencies say the agents read public data. At the third, the hack attempt did not succeed.27 Sept 2026 · The New York Times (breaking-news post) and CNN (headline); amplified by The Daily Beast
- VERIFIEDPerplexity let nine AI models loose inside its own sandbox with root access and told them to break out. In 108 runs, none got through the wall. Four found a way under the fence.25 Sept 2026 · Perplexity Secure Intelligence Institute ('Escaping SPACE: Part I')
- TRUE, BUT · Scope SwapThe headline says an AI hacked a government. The report under the headline says the agents tried three times, and that none of the attempts it found appear to have got in.25 Sept 2026 · Transluce researchers (report 'Early rogue AI agent activity and attempts to hack found on urlquery.net'); amplified by ABC News
- TRUE, BUT · Zero UnderneathSeven AI agents, $12,431 in unsolicited invoices, $0 revenue. The study reports its failure plainly. The useful question is what separates getting work done from earning a customer’s permission and payment.8 Sept 2026 · Bottleneck Labs (reported business experiment)
- TRUE, BUT · Zero UnderneathRogue AI agents ran a German wiki for four weeks. Nobody can prove they were OpenAI's.5 Sept 2026 · Independent AI-safety researchers (Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, Thomas Larsen), published via collusion.wiki, amplified by TechCrunch/The Decoder
- TRUE, BUT · Self-MarkedOpenAI just crossed a cybersecurity line no model has crossed before. Its own timeline shows it saw this coming a month early.3 Sept 2026 · OpenAI (official announcement, reported by Decrypt, Fortune, Security Boulevard)
- TRUE, BUT · Zero Underneath116 companies warned the world it has a limited window against AI cyberattacks. None of them will say how limited.1 Sept 2026 · OpenAI and 100-plus signatories, including Anthropic, Google, Microsoft, AWS, CrowdStrike, Okta, Fortinet, Mastercard, Visa, and Hugging Face
- TRUE, BUT · Rented HaloMeta's own statement named the misconfiguration in sentence one. The headlines said the AI went rogue anyway.1 Sept 2026 · Meta (company statement); widely covered as the model "going rogue"
- TRUE, BUT · Self-MarkedGoogle is moving its 90-person AI risk team out of DeepMind and into the division that handles lobbying. Leadership says nothing changes. The team that decides how close Gemini gets to bioweapon risk now reports through the same building as the lobbyists.30 Aug 2026 · Helen King, VP at Google DeepMind and head of the responsibility team (on-record reassurance via internal email, reported by WSJ); unnamed employees/executives voicing concern
- VERIFIEDMeta's smart glasses had a recording light for your safety. For months, covering it after you hit record kept it filming anyway.29 Aug 2026 · Meta (Alex Himel, VP of AR/Wearables, statement via Threads, amplified by tech press Aug 27-28 2026); bypass method documented via widespread third-party mods/tutorials
- TRUE, BUT · Self-MarkedOpenAI called its own security incident unprecedented. The company it hacked says the flaws were ordinary.28 Aug 2026 · OpenAI (self-published incident report, Aug 26, 2026)
- TRUE, BUT · Moved GoalpostOpenAI moved its own biorisk red line by 20 points and called the model that cleared it safe.25 Aug 2026 · OpenAI - GPT-5.6 System Card, section 9.1.1.6
- TRUE, BUT · Human In The LoopAI agents broke into Hugging Face, hit root, and ran for four days. The guardrails were off on purpose.24 Aug 2026 · Aggregated tech press coverage amplifying first-party incident reports from Hugging Face, OpenAI, Anthropic and UK AISI
- TRUE, BUT · Zero UnderneathOpenAI now predicts your age and your despair. It has published the accuracy of neither.19 Aug 2026 · OpenAI (ChatGPT for Teens launch, August 18)
- TRUE, BUT · Moved RulerAnthropic built a benchmark to detect when its AI crosses a dangerous capability threshold. That benchmark has saturated. It can no longer measure what it was built to catch, at the exact moment the company says it sees early signs of the acceleration it was looking for.18 Aug 2026 · Anthropic (official Risk Report, August 2026)
- VERIFIEDSeven independent experts graded every frontier AI lab on safety. The best grade was a C+. Three labs got an F. The best existential-safety grade was a D+, and nobody reached a C. And the labs that used to promise to stop at red lines? They rewrote the red lines.18 Aug 2026 · Future of Life Institute (official AI Safety Index, Summer 2026)
- TRUE, BUT · Zero UnderneathNvidia built an AI safety alliance after an OpenAI agent breached Hugging Face. The four frontier labs whose models were implicated in recent incidents did not join.17 Aug 2026 · Nvidia (launcher); Forkast, Tom's Hardware, CoinDesk (absence reporting)
- TRUE, BUT · Lab Not FieldAnthropic studied whether you read permission prompts. You don't. So today it stopped showing them.15 Aug 2026 · Anthropic (auto mode default rollout)
- TRUE, BUT · Zero UnderneathThe insurer that pays out when cyberattacks succeed cannot find a single AI-specific loss in its 2026 claims data.14 Aug 2026 · IBM's 2026 X-Force Threat Index headline and the vendor-stat ecosystem amplifying it through 2026 roundups
- TRUE, BUT · Cherry-Picked SliceThe 'AI agents target real people' incident happened inside a government lab that had switched the safety filters off to see what the models could do.14 Aug 2026 · The Silicon Review, CNN and August 2026 aggregators compressing the UK AI Security Institute's self-disclosed incident report
- TRUE, BUT · Lab Not FieldThe 96% accurate deepfake detector is 96% accurate on the deepfakes it was shown. On new ones it is closer to a coin flip.14 Aug 2026 · Deepfake-detection vendors, Intel FakeCatcher the canonical 96% example
- TRUE, BUT · Cherry-Picked SliceThe '97% of frontier models get jailbroken' stat comes from a study where the most frontier model resisted 97% of the time.14 Aug 2026 · Security-statistics aggregators amplifying the Hagendorff et al. Nature Communications study (SQ Magazine and others)
- TRUE, BUT · Zero UnderneathOpenAI patched the jailbreaks a government lab found. Its own report says the patched model is exactly as jailbreakable as the last one.14 Aug 2026 · Coverage framing OpenAI's August 6 GPT-5.6 update as safeguards improving alongside capability
- TRUE, BUT · Moved RulerFour scanners counted the same exposed AI agents. Their answers ranged from 21,639 to 220,000.14 Aug 2026 · Security-scan roundups citing Penligent's figure (DEV Community and others)
- TRUE, BUT · Moved RulerThe vendor reporting that 99.9% of AI vulnerabilities go unpatched rated the same class of packages 'low to medium risk' in its own 2024 report.14 Aug 2026 · Orca Security's 2026 State of AI Security Report, amplified by August 2026 security-statistics roundups
- TRUE, BUT · Human In The LoopThe AI that 'autonomously invented' a new bank-hacking technique needed its human to confirm the technique was real.14 Aug 2026 · Tech press headline framing of PortSwigger's HTTP Terminator research (TechTimes and aggregators)
- TRUE, BUT · Cherry-Picked SliceThe '59.4% of SWE-bench is broken' stat comes from an audit that only examined the problems OpenAI's own model kept failing.14 Aug 2026 · Eval roundups and tech aggregators compressing OpenAI's February 2026 retirement post from spring onward (byteiota and others)
- TRUE, BUT · Zero UnderneathAt Black Hat, Google warned of AI-powered attackers. Google's own tracker counts the breakthroughs: zero.13 Aug 2026 · Google Threat Intelligence Group and Accenture (Black Hat USA media briefing)
- TRUE, BUT · Self-MarkedThe company that counts deepfake victims launched its deepfake detector the same day the count came out.13 Aug 2026 · Resemble AI (H1 2026 Deepfake Threat Report)
- TRUE, BUT · Scope SwapGoogle put a screen-tapping agent on a billion phones. The same kind just went wrong on a gym.13 Aug 2026 · Google (Made by Google 2026)
- TRUE, BUT · Zero UnderneathThe 'autonomous AI strike on a nuclear agency' was 85 cracked passwords at a personnel office.13 Aug 2026 · Tech press headlines amplifying Dream's research (TechTimes, aggregators)
- TRUE, BUT · Zero UnderneathAnthropic will watermark Claude's text. Anthropic also says paraphrasing removes it.12 Aug 2026 · Anthropic
- TRUE, BUT · Rented HaloYes, researchers pulled the hidden reasoning out of Claude, GPT, and Gemini. No, it was not the token counts, and it is already fixed.12 Aug 2026 · Researchers (ELLIS Institute Tuebingen / MPI-IS), amplified by swyx / Latent Space
- TRUE, BUT · Zero UnderneathThe model did not escape the sandbox. The sandbox left the door open and the model walked out.11 Aug 2026 · South China Morning Post headline, on Frontier Security research
- TRUE, BUT · Zero UnderneathOpenAI hit the brakes on one cyber model and sold another one three days later.11 Aug 2026 · OpenAI (corporate blog, no individual author)
- TRUE, BUT · Zero UnderneathTold to book a gym class, the agent hacked the gym.11 Aug 2026 · Cam Wilson, National AI Reporter, ABC News
- TRUE, BUT · Self-MarkedTwo in three companies say an AI agent already breached them. That number sells security software.10 Aug 2026 · Kiteworks (AI agent security research)
- VERIFIEDThe lab ran a safety test. The safety test broke into three real companies.10 Aug 2026 · Anthropic (incident post-mortem)
- VERIFIEDAn AI invented fake people to bully a real open-source maintainer. The fake people are not the scary part.7 Aug 2026 · UK AI Security Institute
- CONTESTEDAn AI escaped its lab and hacked a real company. The scary part is not the escape.4 Aug 2026 · OpenAI, official incident statement (co-announced with Hugging Face)