The '59.4% of SWE-bench is broken' stat comes from an audit that only examined the problems OpenAI's own model kept failing.
OpenAI really did retire its own benchmark and the contamination evidence is damning. But 59.4% is the flaw rate of a hand-picked failure pile. As a share of the full benchmark, the confirmed broken tasks are 82 out of 500.
"OpenAI retired SWE-bench Verified after its audit found 59.4% of tasks had flawed test cases, and every frontier model trained on the solutions" [SOURCE ↗]

THE CLAIM, as it lands in August 2026 eval roundups: OpenAI abandoned SWE-bench Verified because 59.4% of its tests were flawed and every frontier model had trained on the answers. THE CHECK: the retirement is real, dated February 23, 2026, and the contamination findings are the strongest part. But the 59.4% comes from an audit of 138 problems selected precisely because OpenAI's o3 failed them across 64 runs. Broken tests are unsolvable, so they pile up in exactly that failure set. As a share of the whole 500-problem benchmark, the confirmed flawed tasks are 82, or 16.4%. The right reading is that the top of the benchmark was phantom headroom, not that the whole thing was always garbage.
On February 23, 2026, OpenAI published a quiet execution notice: 'Why SWE-bench Verified no longer measures frontier coding capabilities.' The company stopped reporting scores on the most-cited coding benchmark in the industry and asked everyone else to stop too. Within weeks the story had been comp
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 3 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.