GET THE AUTOPSY ➔

Claude Fable 5 really is number one on the hardest AI leaderboards. It also scores 43 on the knowledge benchmark it leads, on a scale that runs from minus 100 to 100, and 55.5% on an exam built so models fail it.

The wins are real and earned. The absolute numbers are not what 'most capable model' makes them sound like. On AA-Omniscience the record is 43 on a minus-100-to-100 scale where most models land below zero. On Humanity's Last Exam it scores 55.5% on a test adversarially built so models fail it. And the leaderboard entry is the Opus 4.8 fallback configuration. A ranking is not a reliability score.

01THE CLAIM
"Claude Fable 5 is Anthropic's most capable public model and the new benchmark leader, topping the Artificial Analysis Intelligence Index and Humanity's Last Exam and setting the highest score to date on the AA-Omniscience knowledge and hallucination benchmark." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-25
ANTHROPIC POSITIONING AND BENCHMARK COVERAGE TRACK RECORD1 CLAIM · 40/100 BS RATE →
55.5%FABLE 5 ON HUMANITY'S LAST EXAM (AA), AN EXAM SEEDED WITH QUESTIONS THAT STUMPED AIS
43FABLE 5 AA-OMNISCIENCE RECORD, ON A SCALE FROM -100 TO 100
#1REAL AND EARNED; A RANKING, NOT A RELIABILITY SCORE
>95%FABLE 5 SESSIONS WITH NO FALLBACK, PER ANTHROPIC; THE RECORD ENTRY IS THE OPUS 4.8 FALLBACK CONFIG
Claude Fable 5 really is number one on the hardest AI leaderboards. It also scores 43 on the knowledge benchmark it leads, on a scale that runs from minus 100 to 100, and 55.5% on an exam built so models fail it.
02THE CHECK

THE CLAIM. Claude Fable 5, Anthropic's most capable public model, is the new benchmark leader, topping the Artificial Analysis Intelligence Index and Humanity's Last Exam and setting the highest score to date on AA-Omniscience, the knowledge and hallucination benchmark.

THE CHECK. the ranking is real. Fable 5 sits at number one on these boards on independent evaluation, which is not nothing. What the headline hides is what the numbers mean. AA-Omniscience runs from minus 100 to 100, where zero means as many right as wrong and most frontier models score below zero, so a leading 43 is a genuine jump and still a long way from anything a layperson would call omniscient. Humanity's Last Exam was built so models fail it, seeding only questions that already stumped the best AIs, so a rank of number one at 55.5% is a lead, not a grade. And the leaderboard entry is a specific named configuration: the record holder is the Adaptive Reasoning, Max Effort, Opus 4.8 Fallback setup, and Anthropic has not disclosed how the number would move without the fallback. Number one is a ranking, not a report card.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Fable 5 is genuinely the top model on Humanity's Last Exam and AA-Omniscience, and that is earned. But the omniscience record is 43 on a scale from minus 100 to 100 where most models score negative, the exam is designed so models fail it, and the record entry is the Opus 4.8 fallback configuration. Number one is a ranking, not a reliability score."

Claude Fable 5 is the top model right now on the two hardest public evaluations, and it earned the spot on an independent scoreboard, not just a vendor slide. On Artificial Analysis, the record holder on Humanity's Last Exam is Fable 5 at 55.5%, and it also sets the highest score to date on AA-Omnis

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.