THE ARCHIVE · EVERY STORY ON THE RECORD
The Archive
Every story we have checked, newest first. Click through to the full autopsy. the free short of every story stays free, forever.

Claude Fable 5 really is number one on the hardest AI leaderboards. It also scores 43 on the knowledge benchmark it leads, on a scale that runs from minus 100 to 100, and 55.5% on an exam built so models fail it.
TRUE, BUTAnthropic positioning and benchmark coverage (Claude Fable 5, August 2026)

Two new benchmarks agree: the best AI models in the world clear fewer than half of a hard benchmark of real analyst tasks. Claude Fable 5 tops the frontier at 49.2%. The context is who built the tests, and who sells the fix.
TRUE, BUTSamaya AI (FrontierFinance) and Vals AI (Finance Agent v2), August 2026

OpenAI and Anthropic are selling the same next step for AI agents: more of them. Claude Code now forks subagents by default, and Sol Ultra fans a problem across up to 64. Google Research ran the controlled test, and the answer is a split: more agents help work that breaks into independent pieces and hurt work that runs as one dependent chain, by up to 70%. Which one your task is decides whether the swarm is an upgrade or a tax.
TRUE, BUTOpenAI (GPT-5.6 Sol Ultra) and Anthropic (Claude Code subagents), and the industry 'more agents are better' heuristic, August 2026

The coding number everyone quotes says AI has nearly solved software engineering: Claude Opus 5 scores 96% on SWE-bench Verified. Move to the benchmark built to resist contamination and the top model sits at 80.3%, and GPT-5.6 Sol lands at 64.6%.
TRUE, BUTFrontier-model coding marketing built on SWE-bench Verified (Anthropic, OpenAI and coverage, August 2026)

93 percent of developers now use AI coding tools. Six independent studies converge on the same measured productivity gain: about 10 percent. In a randomized trial, experienced developers using frontier AI tools took 19 percent longer than those working without them. They thought they were 20 percent faster.
CONTESTEDAI tool vendors and companies citing AI productivity to justify spend
Anthropic is pitching investors a $2 trillion IPO built on a revenue forecast that requires 4.3x growth in under two years, from a company that posted its first quarterly profit three months ago on a discounted compute bill.
TRUE, BUTAnthropic (via bankers and investor presentations, per Reuters exclusive)

Anthropic built a benchmark to detect when its AI crosses a dangerous capability threshold. That benchmark has saturated. It can no longer measure what it was built to catch, at the exact moment the company says it sees early signs of the acceleration it was looking for.
TRUE, BUTAnthropic (official Risk Report, August 2026)

Seven independent experts graded every frontier AI lab on safety. The best grade was a C+. Three labs got an F. The best existential-safety grade was a D+, and nobody reached a C. And the labs that used to promise to stop at red lines? They rewrote the red lines.
VERIFIEDFuture of Life Institute (official AI Safety Index, Summer 2026)

The White House's own science advisor says companies are blaming AI for layoffs they would have done anyway because it plays better in the press. Early in 2026, 7 percent of layoffs cited AI. By summer, 54 percent did. The count hit 205,000 workers in under eight months.
TRUE, BUTMichael Kratsios, White House Science and Technology Advisor (Moonshots podcast, Aug 7 2026)

OpenAI is preparing to sell shares to the public at a valuation above $1 trillion. For every dollar the company earns, it loses $1.22. Its gross margin is falling, not rising, as revenue grows. HSBC estimates it needs another $207 billion in capital by 2030.
TRUE, BUTOpenAI (per pre-S-1 reporting, confidential filing June 2026)
A rocket company paid $60 billion in its own stock for a code editor. Cursor's revenue has never been publicly disclosed. The deal closed on August 14, making it the year's largest startup acquisition, and Cursor now claims access to the largest fleet of GPUs in the world, a fleet nobody outside SpaceX can count.
TRUE, BUTSpaceX / Cursor (per Bloomberg, TechCrunch, regulatory filing Aug 14 2026)

The first AI boss fired a human this week. It had to be told its own rules first, and a human held the axe.
TRUE, BUTWeekend coverage of Andon Labs' store-manager experiment (AI boss fires human worker framing)
