GET THE AUTOPSY ➔

The '59.4% of SWE-bench is broken' stat comes from an audit that only examined the problems OpenAI's own model kept failing.

OpenAI really did retire its own benchmark and the contamination evidence is damning. But 59.4% is the flaw rate of a hand-picked failure pile. As a share of the full benchmark, the confirmed broken tasks are 82 out of 500.

01THE CLAIM
"OpenAI retired SWE-bench Verified after its audit found 59.4% of tasks had flawed test cases, and every frontier model trained on the solutions" [SOURCE ↗]
TRUE, BUT3 SOURCES · LIVE 2026-08-25
BYTEIOTA TRACK RECORD1 CLAIM · 40/100 BS RATE →
59.4%SHARE OF THE 138 AUDITED PROBLEMS, ALL PRE-SELECTED BECAUSE O3 DID NOT CONSISTENTLY SOLVE THEM, WITH MATERIAL ISSUES
27.6%SHARE OF THE DATASET AUDITED, CHOSEN BECAUSE MODELS OFTEN FAILED IT, WHERE BROKEN TESTS POOL
31'ALMOST IMPOSSIBLE' TASKS GPT-5.2 SOLVED ANYWAY, OPENAI'S SMOKING GUN FOR CONTAMINATION
The '59.4% of SWE-bench is broken' stat comes from an audit that only examined the problems OpenAI's own model kept failing.
02THE CHECK

THE CLAIM, as it lands in August 2026 eval roundups: OpenAI abandoned SWE-bench Verified because 59.4% of its tests were flawed and every frontier model had trained on the answers. THE CHECK: the retirement is real, dated February 23, 2026, and the contamination findings are the strongest part. But the 59.4% comes from an audit of 138 problems selected precisely because OpenAI's o3 failed them across 64 runs. Broken tests are unsolvable, so they pile up in exactly that failure set. As a share of the whole 500-problem benchmark, the confirmed flawed tasks are 82, or 16.4%. The right reading is that the top of the benchmark was phantom headroom, not that the whole thing was always garbage.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"'59.4% of which tasks?' The audit only looked at the 138 problems o3 kept failing. Flawed tests live in the failure pile by definition. The confirmed count is 82 of 500. The other 418 were never audited, so the honest phrase is not shown to be broken, which is not the same as proven clean."

On February 23, 2026, OpenAI published a quiet execution notice: 'Why SWE-bench Verified no longer measures frontier coding capabilities.' The company stopped reporting scores on the most-cited coding benchmark in the industry and asked everyone else to stop too. Within weeks the story had been comp

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 3 sources with quotes and screenshots, and our on-record call.

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.