Subscribe

TRUE, BUT

The coding number everyone quotes says AI has nearly solved software engineering: Claude Opus 5 scores 96% on SWE-bench Verified. Move to the benchmark built to resist contamination and the top model sits at 80.3%, and GPT-5.6 Sol lands at 64.6%.

Frontier-model coding marketing built on SWE-bench Verified (Anthropic, OpenAI and coverage, August 2026). Recorded 19 August 2026. Our ruling means: literally true, materially misleading.

Read the story and its public receipts →

Receipts pack

The evidence file for this verdict

6 receipts: 1 support the claim, 1 contradict it, 4 add context. 4 figures on the record, 4 traced to a quoted receipt.

  1. 1benchlm.ai ↗Every source link and quote stays open on the story page.
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6

Members get the full file

What each receipt proves, and what would change our mind.

  • Every receipt with its source, date and stance, plus what it proves
  • The figures table, each number traced to the receipt that states it
  • The trick behind the claim, explained, with the guide to spotting it
  • What would change our verdict
  • Our call on the record, with its due date

The verdict, the opening and every source link stay free. Membership opens the evidence file for every checked story.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.