The '97% of frontier models get jailbroken' stat comes from a study where the most frontier model resisted 97% of the time.
The Nature paper is real and its warning is serious. The number people quote from it is an average across weak targets, and the appendix that debunks the headline is in the same paper.
"Reasoning models now jailbreak frontier LLMs autonomously at a 97% success rate" [SOURCE ↗]

THE CLAIM, as it circulates through 2026 security roundups: reasoning models autonomously jailbreak frontier LLMs at 97% success. THE CHECK: the 97.14% figure comes from Hagendorff, Derner and Oliver's Nature Communications study, and it is an average across all attacker-target combinations, including old and weak targets. The paper's own data says the most resistant target, Claude 4 Sonnet, took the top harm score on just 2.86% of items, and attacker success ranged from 12.86% to 90%. Real alignment finding, real warning. The single scary number flattens all of it.
By late 2026 a statistic had gone feral. 'Multi-turn jailbreaks hit 97% success on frontier LLMs', reads one widely-syndicated security roundup, sitting in a list next to 'jailbreak attempts succeed 20% of the time on average, according to IBM research'. Both cannot describe the same world, and the
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 3 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.