GET THE AUTOPSY ➔

The '97% of frontier models get jailbroken' stat comes from a study where the most frontier model resisted 97% of the time.

The Nature paper is real and its warning is serious. The number people quote from it is an average across weak targets, and the appendix that debunks the headline is in the same paper.

01THE CLAIM
"Reasoning models now jailbreak frontier LLMs autonomously at a 97% success rate" [SOURCE ↗]
TRUE, BUT3 SOURCES · LIVE 2026-08-25
SQ MAGAZINE TRACK RECORD1 CLAIM · 40/100 BS RATE →
2.86%HARM RATE AGAINST THE MOST RESISTANT TARGET (CLAUDE 4 SONNET)
97.14%THE AGGREGATE, AVERAGED ACROSS 9 TARGETS INCLUDING WEAK ONES
12.86%SUCCESS RATE OF THE WEAKEST ATTACKER (QWEN3)
The '97% of frontier models get jailbroken' stat comes from a study where the most frontier model resisted 97% of the time.
02THE CHECK

THE CLAIM, as it circulates through 2026 security roundups: reasoning models autonomously jailbreak frontier LLMs at 97% success. THE CHECK: the 97.14% figure comes from Hagendorff, Derner and Oliver's Nature Communications study, and it is an average across all attacker-target combinations, including old and weak targets. The paper's own data says the most resistant target, Claude 4 Sonnet, took the top harm score on just 2.86% of items, and attacker success ranged from 12.86% to 90%. Real alignment finding, real warning. The single scary number flattens all of it.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"'97% against which target?' The paper names them. Against Claude 4 Sonnet the harm rate was 2.86%. Against old DeepSeek-V3 it was 90%. The average is not the story."

By late 2026 a statistic had gone feral. 'Multi-turn jailbreaks hit 97% success on frontier LLMs', reads one widely-syndicated security roundup, sitting in a list next to 'jailbreak attempts succeed 20% of the time on average, according to IBM research'. Both cannot describe the same world, and the

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 3 sources with quotes and screenshots, and our on-record call.

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.