GET THE AUTOPSY ➔

Anthropic built a model stronger than its flagship. You cannot use it, test it, or check the number.

Model 2 beats Mythos 5 by 12.5 points on CoBench: an internal test, of an internal model, with assessments Anthropic itself calls incomplete. The risk label moved in the same report, for a reason the headlines are fusing with the wrong cause.

01THE CLAIM
"Anthropic's second company-wide Risk Report discloses an unreleased internal Model 2 that is more capable than its shipped flagship Mythos 5, scoring 62.8% on CoBench versus 50.3%, with no plans for external release, alongside raising its catastrophic-misalignment rating from very low to low." [SOURCE ↗]
TRUE, BUT5 SOURCES · LIVE 2026-08-25
ANTHROPIC TRACK RECORD44 CLAIMS · 39/100 BS RATE →
62.8%MODEL 2 ON COBENCH, MEASURED BY ANTHROPIC, INSIDE ANTHROPIC
50.3%THE SHIPPED FLAGSHIP ON THE SAME INTERNAL TEST
0OUTSIDE RUNS: NO WEIGHTS, NO API, NO THIRD PARTY
Anthropic built a model stronger than its flagship. You cannot use it, test it, or check the number.
02THE CHECK

THE CLAIM. Anthropic's August 2026 Risk Report reveals Model 2, an internal system 'somewhat more capable' than Mythos 5 (62.8% vs 50.3% on CoBench), which the company does not plan to release; the same report raises Anthropic's catastrophic-misalignment rating from very low to low.

THE CHECK. the disclosure is real transparency and the capability claim is a closed loop. The score comes from an internal benchmark, run internally, on a model nobody outside Anthropic can touch: no weights, no API, no third-party run. Anthropic itself lowers the claim's confidence in ink, saying it 'has not run all of its typical predeployment assessments and therefore has somewhat less confidence in its beliefs about the model's capabilities'. Meanwhile the risk-label change the headlines attach to this scary stronger model has a different stated cause: increased overall uncertainty from recent cyber-evaluation incident disclosures involving shipped models, not Model 2. The report even notes its concrete task evaluations have 'saturated', meaning the measuring sticks maxed out. A stronger secret model and a raised risk label are both in the report. The causal arrow between them is not.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The 62.8 is an internal score on an internal test of a model nobody outside can run, and Anthropic itself says its assessment is incomplete. The risk label moved for a different reason than the stronger-model headline implies."

Anthropic published its second company-wide Risk Report on August 14, under version 3.4 of its Responsible Scaling Policy. Two disclosures drove the coverage. First, the company raised its rating of catastrophic harm from misalignment in high-stakes settings from very low to low. Second, the report

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 5 sources with quotes and screenshots, and our on-record call.

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.