Subscribe

TRUE, BUT

Two new benchmarks agree: the best AI models in the world clear fewer than half of a hard benchmark of real analyst tasks. Claude Fable 5 tops the frontier at 49.2%. The context is who built the tests, and who sells the fix.

Samaya AI (FrontierFinance) and Vals AI (Finance Agent v2), August 2026. Recorded 19 August 2026. Our ruling means: literally true, materially misleading.

Read the story and its public receipts →

Receipts pack

The evidence file for this verdict

7 receipts: 3 support the claim, 1 contradict it, 3 add context. 4 figures on the record, 3 traced to a quoted receipt.

  1. 1arxiv.org ↗Every source link and quote stays open on the story page.
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7

Members get the full file

What each receipt proves, and what would change our mind.

  • Every receipt with its source, date and stance, plus what it proves
  • The figures table, each number traced to the receipt that states it
  • The trick behind the claim, explained, with the guide to spotting it
  • What would change our verdict
  • Our call on the record, with its due date

The verdict, the opening and every source link stay free. Membership opens the evidence file for every checked story.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.