GET THE AUTOPSY ➔

Issue #13

WEDNESDAY 19 AUGUST 2026 · 8 CLAIMS CHECKED · 0 SURVIVED THE RECEIPTS · ISSUE 13 OF 17

AI is building AI, the headlines say. So independent researchers handed frontier agents real, unpublished research questions and six days each. The agents did all of the engineering and wrote up the results. The papers' own authors rejected both. One got a Strong Reject.

The acceleration is real: Anthropic says more than 80% of its merged code is now written by Claude, with internal research speedups around 52x. But the jump from that to autonomous research just met its first controlled test. A Princeton and UK AI Security Institute shadow evaluation ran frontier agents on two unpublished NeurIPS submissions. It is early evidence, not a verdict, but the split was clean: they handled every piece of the engineering and could not answer the research question. Anthropic's own post says the same thing: we are not there yet.

01THE CLAIM
"AI is starting to build AI: frontier systems are automating meaningful parts of AI research, with Anthropic delegating a growing share of its own development to AI and OpenAI reporting a model that saved researchers several weeks, feeding the view that autonomous AI research is within reach." [SOURCE ↗]
TRUE, BUT9 SOURCES · LIVE 2026-08-28
ANTHROPIC TRACK RECORD39 CLAIMS · 38/100 BS RATE →
Strong RejectTHE GRADE ONE AI-PRODUCED PAPER RECEIVED; BOTH WERE REJECTED
80%ANTHROPIC CODE WRITTEN BY CLAUDE (MAY 2026); THE ACCELERATION IS REAL
~52xANTHROPIC'S INTERNAL RESEARCH SPEEDUP BY APRIL 2026, PER ITS OWN POST
six daysCOMPUTE PER RESEARCH QUESTION; THE AGENTS STILL COULD NOT ANSWER IT
AI is building AI, the headlines say. So independent researchers handed frontier agents real, unpublished research questions and six days each. The agents did all of the engineering and wrote up the results. The papers' own authors rejected both. One got a Strong Reject.
02THE CHECK

THE CLAIM. AI is beginning to build AI. Anthropic reports it is delegating a growing share of its own development to AI systems, OpenAI says a model helped post-train a smaller one and saved researchers several weeks, and the wider read is that autonomous AI research is within reach.

THE CHECK. the acceleration is documented and real. What was missing was any controlled test of the leap from acceleration to autonomy. A Princeton and UK AI Security Institute team built one, a shadow evaluation, where a frontier agent takes on the central open question of a high-quality unpublished paper and the paper's original authors grade the output. They ran it on unpublished NeurIPS 2026 submissions with frontier agents given six days and thousands of dollars of compute each. The agents completed all of the engineering without human help, ran the literature searches, debugged the GPU code, compiled full papers, and could not make substantial progress on the actual research questions. Both papers were unambiguously rejected. The authors catalogued five recurring failure modes, starting with poor judgment about what clears the bar for publishable research. Even Anthropic's post agrees on the ceiling: we are not there yet.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Frontier agents given six days on real, unpublished research questions did all the engineering and still wrote two papers their own authors rejected, one a Strong Reject. The acceleration Anthropic reports, 80% of code written by Claude, is real. Autonomous AI research is a different claim, and Anthropic's own post says we are not there yet."
DEEP DIVE · THE FULL AUTOPSY

What the headlines say

The story of the summer is that AI has started building AI. Anthropic's post 'When AI builds itself' put numbers on it, OpenAI said one of its models helped post-train a smaller one and saved researchers several weeks, and the compressed version traveling the timeline is that autonomous AI research is nearly here. As the new study's own abstract puts it, 'Forecasts of explosive AI progress hinge on AI agents automating AI research.' So the question is not rhetorical. It is the load-bearing assumption under a lot of the hype, and until now it had barely been tested.

The acceleration is real

Start by granting the strong part, because it is true. Anthropic reports that 'As of May 2026, more than 80% of the code we merge into Anthropic's codebase was authored by Claude,' and that its internal research loop reached roughly 52x speedups on some tasks. That is not a rounding error and it is not marketing fluff. AI genuinely does a large and growing share of the engineering labor inside a frontier lab. If the claim were only 'AI massively accelerates the engineering of research,' this desk would have nothing to autopsy.

The first controlled test

The claim is bigger than that, though, and a team from Princeton and the UK AI Security Institute built the first clean test of the bigger version. They call it a shadow evaluation: an agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors, the people who actually solved it, grade the agent's output. They ran it on two unpublished NeurIPS 2026 submissions, 'giving frontier agents six days and thousands of dollars of compute.' This is not a toy benchmark scraped from public repositories. It is a real open question with a known good answer that the model could not have memorized.

Engineering yes, research no

The result is clean and it splits exactly where the hype blurs. In the authors' words, 'The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions.' Coverage of the study fills in the picture: the agents 'ran literature searches, debugged GPU code, completed hundreds of experiments and robustness tests, and compiled full papers in LaTeX.' Every piece of the machinery worked. And then 'both papers were unambiguously rejected by the authors.' One got a Strong Reject. The team catalogued five recurring failure modes, and the first one is the whole story: poor judgment about the bar for publishable research. The agents could not tell a real contribution from a narrowed restatement of a falsified hypothesis, could not back out of dead ends, and drifted off instructions. They did the work. They could not do the science.

Anthropic already said the quiet part

Here is the part the headlines skip. Anthropic's own post does not actually claim autonomous research. It hedges, plainly: 'We are not there yet, and recursive self-improvement is not inevitable.' The lab shared real acceleration data and an explicit ceiling in the same document. The overclaim was manufactured downstream, in the retelling that kept the 80% and dropped the caveat. The study did not catch Anthropic lying. It caught the ecosystem rounding a careful claim up into an inevitability.

The honest steelman

Two papers is a small sample, the models will improve, and automating the engineering of research is itself a large deal that frees scarce human attention for the judgment calls. All fair. But the burden here runs the other way. The confident version of the claim, that autonomous research is within reach, is the one making forecasts and raising money, and it is the one that just failed its first controlled test on real problems with generous compute. Early evidence is still evidence, and right now the evidence says the hard part is untouched.

What would settle it

Run more shadow evaluations, across more papers and labs, and show frontier agents producing work the original authors would accept rather than reject. Until that exists, the accurate sentence is the boring one: AI does the engineering of AI research well and, on the problems tested so far, the research itself not yet, and the distance between those two facts is the entire debate.

04YOUR MOVE ⚡ WHAT IGNORING THIS COSTS

There are two honest sentences here and the headline keeps only the first. One: AI now does the engineering of research at a level that genuinely compresses weeks of grunt work, which is a real and useful capability. Two: on the part that makes research research, forming a judgment about what is worth pursuing and recognizing when a result clears the bar, the same agents produced papers their own authors rejected. Automating the labor is not the same as automating the science, and on the two papers tested the gap was total, not marginal.

05🔮 OUR CALL · ON THE RECORD 2026-08-19

The 80% and the 52x will keep getting read as a countdown to AI that does science by itself, and this study is the first hard data point that the countdown is measuring the wrong thing. Expect the engineering-automation half to keep improving fast, because that is what the productivity numbers actually track. Expect the research-judgment half, forming good questions and knowing when an answer is real, to be the slow, sticky part, and expect the next round of self-improvement claims to keep quietly resting on the engineering half while the headline implies the whole.

Flips toward BS if the acceleration claims are shown to be cherry-picked engineering wins with no path to research judgment. Flips toward HOLDS UP if a follow-on shadow evaluation, on a larger set of papers, shows frontier agents producing work the original authors would accept, not just engineering that runs.

RECEIPTS (9) · CONFIDENCE HIGH

every URL below answered a live HTTP check before publish · sweep 2026-08-28

  • arxiv.org · "The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions."
  • arxiv.org · "As a result, both papers were unambiguously rejected by the authors."
  • arxiv.org · "We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute."
  • arxiv.org · "Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle."
  • the-decoder.com · "The agents managed all engineering work without human help. They ran literature searches, debugged GPU code, completed hundreds of experiments and robustness tests, and compiled full papers in LaTeX."
  • the-decoder.com · "OpenAI claimed that GPT-5.6 Sol helped with post-training a smaller model and saved researchers several weeks."
  • anthropic.com · "at Anthropic, we are delegating a growing share of AI development to AI systems themselves, which is speeding up our work."
  • anthropic.com · "As of May 2026, more than 80% of the code we merge into Anthropic's codebase was authored by Claude."
  • anthropic.com · "We are not there yet, and recursive self-improvement is not inevitable."

GPT-5.6 Sol scores 92.5% on ARC-AGI-2, the test built so AI would fail it. Read that as abstract reasoning solved, then look one column over on the same scorecard: the same model, same maximum effort, scores 7.78% on ARC-AGI-3, the interactive benchmark the same team built next, where humans still score 100%.

ARC-AGI-2 was designed to resist pattern-matching, and a frontier model just cleared it well above the average human's 66%. Genuine milestone. But it needs Sol's most expensive reasoning setting; turned down it falls to 42.5%, and the ARC Prize's cost-capped grand prize stays unclaimed. Then ARC-AGI-3, a new interactive benchmark from the same team, drops the same model to 7.78% while humans still score 100%. Each version gets saturated and the next reopens the gap. That is a treadmill, not a mind acquiring general reasoning.

01THE CLAIM
"GPT-5.6 Sol leads the ARC-AGI-2 leaderboard at 92.5%, clearing the abstract-reasoning benchmark built specifically to resist AI and beating the average human, a result read as frontier models cracking fluid reasoning." [SOURCE ↗]
TRUE, BUT7 SOURCES · LIVE 2026-08-28
ARC PRIZE VERIFIED LEADERBOARD AND COVERAGE TRACK RECORD1 CLAIM · 40/100 BS RATE →
92.5%GPT-5.6 SOL ON ARC-AGI-2 AT MAX REASONING; ABOVE THE 66% AVERAGE HUMAN
7.78%SAME MODEL, SAME SETTING, ON ARC-AGI-3, THE NEW INTERACTIVE BENCHMARK
100%WHAT HUMAN TESTERS STILL SOLVE ON ARC-AGI-3
42.5%WHERE ARC-AGI-2 FALLS WHEN THE REASONING EFFORT IS TURNED DOWN
GPT-5.6 Sol scores 92.5% on ARC-AGI-2, the test built so AI would fail it. Read that as abstract reasoning solved, then look one column over on the same scorecard: the same model, same maximum effort, scores 7.78% on ARC-AGI-3, the interactive benchmark the same team built next, where humans still score 100%.
02THE CHECK

THE CLAIM. GPT-5.6 Sol tops the ARC-AGI-2 leaderboard at 92.5%, clearing the benchmark built to isolate fluid reasoning and resist memorization, and beating the average human's 66%. The read everywhere is that frontier models have essentially cracked abstract reasoning.

THE CHECK. the score is real and ARC Prize verified, and on its own terms it is a genuine jump. Two things temper the headline. First, the 92.5% depends on Sol's most expensive reasoning mode; dial the effort down and ARC-AGI-2 collapses to 42.5%, and the ARC Prize's own cost-capped grand prize, which needs above 85% cheaply, is still unclaimed. Second, ARC-AGI is a series that keeps raising the bar, and the newest rung is a different kind of test. ARC-AGI-3 is an interactive, agentic benchmark, where a model must explore an unfamiliar environment, infer the goal, and plan, and there the same model at the same setting scores 7.78% while human testers solve 100%. That does not prove the ARC-AGI-2 result was hollow. It shows the fluid-reasoning progress does not yet extend to agentic novelty, and that every time this team builds a new test, today's models fail it until they catch up. Saturating a benchmark is not the same as acquiring the ability it was built to isolate.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"GPT-5.6 Sol really does hit 92.5% on ARC-AGI-2, above the 66% average human, and that is a real jump. But the same model at the same setting scores 7.78% on ARC-AGI-3, which humans still solve 100% of, and the 92.5% needs its most expensive reasoning mode. That is a benchmark being saturated, not reasoning being solved."

GPT-5.6 Sol sits on top of the ARC-AGI-2 leaderboard at 92.5%. ARC-AGI-2 is not a trivia test. It is the benchmark François Chollet's team built specifically to resist the pattern-matching that let models saturate earlier evaluations, a set of visual puzzles designed to isolate fluid reasoning on ta

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 7 sources with quotes and screenshots, and our on-record call.

OpenAI now predicts your age and your despair. It has published the accuracy of neither.

ChatGPT for Teens auto-activates on an age guess and promises parents alerts in high-risk moments. The age predictor has been live since January, the distress flags route through human reviewers, and eleven months after OpenAI said it was building age prediction, neither classifier has a published error rate. Unconfident guesses default adults into teen mode by design.

01THE CLAIM
"ChatGPT for Teens, rolling out globally from August 18, automatically protects users aged 13 to 17: an age-prediction system places estimated under-18s into a restricted experience, and linked parents receive safety notifications when the system detects high-risk moments." [SOURCE ↗]
TRUE, BUT8 SOURCES · LIVE 2026-08-28
OPENAI TRACK RECORD28 CLAIMS · 39/100 BS RATE →
13-17THE AGE BAND THE CLASSIFIER MUST CATCH
11MONTHS SINCE OPENAI SAID IT WAS BUILDING AGE PREDICTION, STILL NO PUBLISHED ACCURACY
UNDER 18WHERE UNCONFIDENT GUESSES LAND, BY DESIGN, ADULTS INCLUDED
OpenAI now predicts your age and your despair. It has published the accuracy of neither.
02THE CHECK

THE CLAIM. ChatGPT for Teens activates automatically when a user says they are 13 to 17 or when OpenAI's systems estimate the user is under 18, wrapping teens in stronger guardrails, parental controls, and safety notifications to parents in high-risk situations.

THE CHECK. every load-bearing promise rests on two classifiers, and neither ships with a number. The age predictor has been in the works since September 2025, when OpenAI said it would 'err on the side of caution, defaulting users to the under-18 experience if it is not confident about their age', with adults offered ways to prove their age back out. The predictor rolled out in January, seven months before this launch, and no accuracy figure, false-positive rate, or evaluation has been published for it in all that live deployment. The distress detector is the same shape with higher stakes: a flagged message goes to a human reviewer, and if it qualifies, parents get notified 'in limited high-risk situations', a phrasing that quietly narrows the launch coverage's acute-distress promise. The reviewer layer is real and disclosed. The accuracy of the flagging classifier that feeds it is not, and what never reaches a reviewer never reaches a parent. Mental-health researchers have already flagged the risk of relying on automated systems in exactly these moments. The guardrails that do not need a classifier (no romantic language, break reminders, quiet hours) are sensible and real. The parts that need one are running on trust.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The teen experience triggers on an age guess, and the parent alerts start with a distress guess that a human reviewer then screens. OpenAI has published the accuracy of neither classifier, eleven months after it started building the first one, and unconfident guesses default adults into teen mode by design."

OpenAI launched ChatGPT for Teens on August 18, rolling out globally across free and paid personal accounts. The mechanism, per the launch coverage: the teen experience 'automatically activates when a user identifies as 13 to 17 or OpenAI's systems estimate that the person is under 18'. Inside it: s

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 8 sources with quotes and screenshots, and our on-record call.

Claude hit 14 of 15 protein targets, and outside labs confirmed it. Then read the method: Claude drove the specialist design tools the field already ships, and a binder is the first step of a drug, not the drug.

The result is real and externally validated, which is rarer than most AI biology claims. The framing is the gap. Claude orchestrated existing structure and co-folding models rather than inventing new capability, its 22 to 35 percent per-design hit rate sits below the specialist state of the art, and Anthropic itself writes that a minibinder is just the first step.

01THE CLAIM
"Anthropic says Claude designed protein binders against 14 of 15 targets, with 22 to 35 percent of individual designs binding successfully versus the 10 to 15 percent typical in protein design campaigns today, results independently produced and tested in the lab by Adaptyv Bio and Twist Bioscience." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-28
ANTHROPIC TRACK RECORD39 CLAIMS · 38/100 BS RATE →
14/15TARGETS CLAUDE GOT A BINDER AGAINST, INDEPENDENTLY LAB-TESTED
0NEW DESIGN MODELS CLAUDE BUILT; IT ORCHESTRATED THE FIELD'S EXISTING ONES
22-35%CLAUDE PER-DESIGN RATE ACROSS 15 TARGETS; A SPECIALIST BEST CASE HITS 70-80% ON 2 CURATED ONES
STEP 1WHERE A BINDER SITS ON THE ROAD TO A DRUG, PER ANTHROPIC
Claude hit 14 of 15 protein targets, and outside labs confirmed it. Then read the method: Claude drove the specialist design tools the field already ships, and a binder is the first step of a drug, not the drug.
02THE CHECK

THE CLAIM. Anthropic reports that Claude, using Opus 4.8 and a Mythos Preview, designed protein binders against 14 of 15 targets, with 22 to 35 percent of individual designs binding successfully against a typical 10 to 15 percent, all physically produced and tested by two outside labs, Adaptyv Bio and Twist Bioscience.

THE CHECK. unusually for this beat, the numbers hold. The designs were made and measured by independent evaluators, and the baseline is sourced. What the headline hides is the method. Claude did not invent a protein-design engine. It orchestrated the structure, sequence, and co-folding models the field already uses, asked to produce 30 candidates per target. Its per-design hit rate of 22 to 35 percent, averaged across all 15 targets, beats the loose field average and sits below the specialist best case of 70 to 80 percent, though that figure comes from just two curated targets with pre-filtered designs, so it is not a clean head to head. And Anthropic states plainly that a minibinder is not a standard drug and that a high-affinity binder is only the first step. This is a strong agent-orchestration result being read as a molecular-biology breakthrough.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Claude got binders against 14 of 15 targets and outside labs confirmed the binding, which is real. But Claude orchestrated the specialist tools the field already had, its 22 to 35 percent rate is an average across 15 targets while the specialist 70 to 80 percent is a best case on two curated ones, and Anthropic itself calls a binder just the first step toward a drug."

On August 18 Anthropic published a result that, for this desk, is unusual: the numbers hold. Claude, running Opus 4.8 and a Mythos Preview, was asked to design protein binders, small proteins that latch onto a target, against 15 separate targets. Its designs were shipped to two independent labs, Ada

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

Claude Fable 5 really is number one on the hardest AI leaderboards. It also scores 43 on the knowledge benchmark it leads, on a scale that runs from minus 100 to 100, and 55.5% on an exam built so models fail it.

The wins are real and earned. The absolute numbers are not what 'most capable model' makes them sound like. On AA-Omniscience the record is 43 on a minus-100-to-100 scale where most models land below zero. On Humanity's Last Exam it scores 55.5% on a test adversarially built so models fail it. And the leaderboard entry is the Opus 4.8 fallback configuration. A ranking is not a reliability score.

01THE CLAIM
"Claude Fable 5 is Anthropic's most capable public model and the new benchmark leader, topping the Artificial Analysis Intelligence Index and Humanity's Last Exam and setting the highest score to date on the AA-Omniscience knowledge and hallucination benchmark." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-28
ANTHROPIC POSITIONING AND BENCHMARK COVERAGE TRACK RECORD1 CLAIM · 40/100 BS RATE →
55.5%FABLE 5 ON HUMANITY'S LAST EXAM (AA), AN EXAM SEEDED WITH QUESTIONS THAT STUMPED AIS
43FABLE 5 AA-OMNISCIENCE RECORD, ON A SCALE FROM -100 TO 100
#1REAL AND EARNED; A RANKING, NOT A RELIABILITY SCORE
>95%FABLE 5 SESSIONS WITH NO FALLBACK, PER ANTHROPIC; THE RECORD ENTRY IS THE OPUS 4.8 FALLBACK CONFIG
Claude Fable 5 really is number one on the hardest AI leaderboards. It also scores 43 on the knowledge benchmark it leads, on a scale that runs from minus 100 to 100, and 55.5% on an exam built so models fail it.
02THE CHECK

THE CLAIM. Claude Fable 5, Anthropic's most capable public model, is the new benchmark leader, topping the Artificial Analysis Intelligence Index and Humanity's Last Exam and setting the highest score to date on AA-Omniscience, the knowledge and hallucination benchmark.

THE CHECK. the ranking is real. Fable 5 sits at number one on these boards on independent evaluation, which is not nothing. What the headline hides is what the numbers mean. AA-Omniscience runs from minus 100 to 100, where zero means as many right as wrong and most frontier models score below zero, so a leading 43 is a genuine jump and still a long way from anything a layperson would call omniscient. Humanity's Last Exam was built so models fail it, seeding only questions that already stumped the best AIs, so a rank of number one at 55.5% is a lead, not a grade. And the leaderboard entry is a specific named configuration: the record holder is the Adaptive Reasoning, Max Effort, Opus 4.8 Fallback setup, and Anthropic has not disclosed how the number would move without the fallback. Number one is a ranking, not a report card.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Fable 5 is genuinely the top model on Humanity's Last Exam and AA-Omniscience, and that is earned. But the omniscience record is 43 on a scale from minus 100 to 100 where most models score negative, the exam is designed so models fail it, and the record entry is the Opus 4.8 fallback configuration. Number one is a ranking, not a reliability score."

Claude Fable 5 is the top model right now on the two hardest public evaluations, and it earned the spot on an independent scoreboard, not just a vendor slide. On Artificial Analysis, the record holder on Humanity's Last Exam is Fable 5 at 55.5%, and it also sets the highest score to date on AA-Omnis

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

Two new benchmarks agree: the best AI models in the world clear fewer than half of a hard benchmark of real analyst tasks. Claude Fable 5 tops the frontier at 49.2%. The context is who built the tests, and who sells the fix.

Samaya AI's FrontierFinance and Vals AI's Finance Agent v2 both put frontier models below the low-50s on realistic analyst work, and the ceiling is genuine and separately reproduced. But both benchmarks come from companies selling finance AI, each is topped by its maker's own system, Samaya's wins its own board at 56%, and on the hardest use cases even the best system lands at 33%. The wall is real. Read who is charging admission.

01THE CLAIM
"Frontier AI models cannot yet do professional investment analysis: on Samaya AI's new FrontierFinance benchmark every frontier model clears fewer than half of real analyst tasks, and Samaya's own in-house system beats them all, a ceiling an independent Vals AI benchmark reproduces." [SOURCE ↗]
TRUE, BUT7 SOURCES · LIVE 2026-08-28
SAMAYA AI TRACK RECORD1 CLAIM · 40/100 BS RATE →
49.2%BEST FRONTIER MODEL (CLAUDE FABLE 5) ON FRONTIERFINANCE; EVERY FRONTIER MODEL UNDER 50%
56%SAMAYA'S OWN SYSTEM ON SAMAYA'S OWN BENCHMARK; A TOP SCORE THAT STILL MISSES NEARLY HALF
33% / 39%WHERE EVEN THE BEST SYSTEM LANDS ON THE HARDEST USE CASES
~52%TOP MODEL ON THE SEPARATE VALS AI FINANCE BENCHMARK; SAME CEILING, DIFFERENT TEAM
Two new benchmarks agree: the best AI models in the world clear fewer than half of a hard benchmark of real analyst tasks. Claude Fable 5 tops the frontier at 49.2%. The context is who built the tests, and who sells the fix.
02THE CHECK

THE CLAIM. frontier AI cannot do professional investment analysis. On Samaya AI's FrontierFinance, a public benchmark of 220 expert queries, every frontier model scores below 50%, with Claude Fable 5 best at 49.2%, GPT-5.6 Sol at 46.8%, and Samaya's own system ahead of all of them at 56%.

THE CHECK. the ceiling is real and it is not a fluke. A separate benchmark from Vals AI, which builds evaluations rather than investment products but still profits when models fall short, reaches the same place: its top model, GPT-5.5, hits roughly 52%, with the frontier clustered in the high-40s to low-50s. Two separate teams, two methods, one wall. What the headline underplays is the shape of the sales floor around it. FrontierFinance is Samaya's own benchmark and Samaya's own system tops it, which is the oldest move in the book, a vendor grading an exam its product is built to pass. Even that winning system clears only 56%, and on the hardest categories, Screening and Discovery and Sector and Macro, the best of everything reaches 33% and 39%. The ceiling is the story. So is the fact that the people ringing the bell are selling the ladder.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Two independent finance benchmarks put the best AI models under the low-50s on real analyst tasks, so the capability ceiling is real. But FrontierFinance is Samaya's own benchmark, Samaya's own system wins it at 56%, and even that winner clears only just over half. The wall is genuine. The scoreboard is marketing."

The pitch of the year in enterprise AI is the autonomous analyst: a model that reads the filings, builds the model, and writes the memo. Two new benchmarks just measured how close that is, and the answer is not close. On FrontierFinance, a public benchmark of 220 expert-crafted queries released by S

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 7 sources with quotes and screenshots, and our on-record call.

OpenAI and Anthropic are selling the same next step for AI agents: more of them. Claude Code now forks subagents by default, and Sol Ultra fans a problem across up to 64. Google Research ran the controlled test, and the answer is a split: more agents help work that breaks into independent pieces and hurt work that runs as one dependent chain, by up to 70%. Which one your task is decides whether the swarm is an upgrade or a tax.

The pitch is that a swarm of parallel subagents beats one worker. The most rigorous study of it, 260 configurations across six benchmarks from Google Research and academic collaborators, says it depends entirely on task shape. On parallelizable work, coordination helps a lot, up to +81%. On sequential, dependent work like planning, every multi-agent setup they tested got worse, by 39 to 70%. A lot of day-to-day coding, debugging a dependent chain, planning a change, is the second kind, and defaulting the swarm on only helps if it fans out for independent subtasks rather than splitting one line of reasoning.

01THE CLAIM
"Parallel subagents are the next capability step for AI agents: the two biggest coding agents now lean on them, with Claude Code forking subagents by default and OpenAI's Sol Ultra fanning a problem across up to 64 concurrent subagents, on the premise that more agents produce better results." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-28
OPENAI TRACK RECORD28 CLAIMS · 39/100 BS RATE →
-70%WORST MULTI-AGENT DEGRADATION ON A STRICTLY SEQUENTIAL TASK (PLANCRAFT)
39-70%RANGE BY WHICH EVERY MULTI-AGENT VARIANT DEGRADED SEQUENTIAL REASONING
+81%WHERE MORE AGENTS DO HELP: A PARALLELIZABLE TASK (FINANCE-AGENT)
64CONCURRENT SUBAGENTS OPENAI'S SOL ULTRA CAN FAN A PROBLEM ACROSS; 4 IS THE REAL DEFAULT
OpenAI and Anthropic are selling the same next step for AI agents: more of them. Claude Code now forks subagents by default, and Sol Ultra fans a problem across up to 64. Google Research ran the controlled test, and the answer is a split: more agents help work that breaks into independent pieces and hurt work that runs as one dependent chain, by up to 70%. Which one your task is decides whether the swarm is an upgrade or a tax.
02THE CHECK

THE CLAIM. parallel subagents are the next leap for AI agents. On August 13 Claude Code turned subagent forking on by default, OpenAI's Sol Ultra fans a task across up to 64 concurrent subagents, and the industry heuristic driving all of it is, in the words of the researchers who tested it, the belief that adding specialized agents will consistently improve results.

THE CHECK. Google Research and academic collaborators ran the first controlled study large enough to answer it, 260 configurations across six agentic benchmarks and five architectures, holding tools, prompts, and compute fixed. The verdict is not that multi-agent is fake. It is that task structure decides. On a parallelizable benchmark like Finance-Agent, coordination delivered up to +81%. On a strictly sequential one like PlanCraft, every multi-agent variant they tested degraded performance by 39 to 70%. Their own summary: coordination improves performance on parallelizable tasks but degrades it on sequential ones. The open question for the two coding agents pushing subagents hardest is which pattern their default triggers: fanning out to gather context in parallel is the winning case, but splitting a dependent job like debugging or planning across coordinating agents is the losing one, and much of what developers hand these tools is the dependent, tool-heavy kind.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Both big coding agents now push parallel subagents as the win, but the controlled Google Research study of 260 configs found more agents help parallelizable tasks by up to 81% and hurt sequential ones like planning by up to 70%. Task structure decides, not agent count, so a default that fans out for independent subtasks can help while one that splits a dependent chain is a tax."

The agent story of the season is parallelism. On August 13, Claude Code shipped a release whose changelog states plainly that 'Subagent forking is now on by default,' turning a swarm of parallel Claude instances from an expert opt-in into the standard behavior. OpenAI is pushing the same idea from t

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

The coding number everyone quotes says AI has nearly solved software engineering: Claude Opus 5 scores 96% on SWE-bench Verified. Move to the benchmark built to resist contamination and the top model sits at 80.3%, and GPT-5.6 Sol lands at 64.6%.

SWE-bench Verified, the coding leaderboard in every launch post, is saturating; benchlm itself calls it nearing saturation, and OpenAI stopped using it in February, with a separate audit finding most of the tasks it checked had flawed tests. Its successor, SWE-bench Pro, is both cleaner and harder: it uses actively maintained repositories with no public ground-truth leakage, plus longer, multi-file, enterprise-scale fixes. There the ceiling drops to 80.3% while GPT-5.6 Sol comes in at 64.6%. The 96% is a saturated, leaky ruler; the honest number is lower.

01THE CLAIM
"Frontier models have all but solved software engineering: on SWE-bench Verified, the coding benchmark quoted in launch posts, the top models now sit near 96%, with Claude Opus 5 leading at 96%." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-28
FRONTIER-MODEL CODING MARKETING BUILT ON SWE-BENCH VERIFIED TRACK RECORD1 CLAIM · 40/100 BS RATE →
96%CLAUDE OPUS 5 ON SWE-BENCH VERIFIED, THE QUOTED CODING NUMBER, NOW SATURATING
80.3%THE TOP SCORE ON CONTAMINATION-RESISTANT SWE-BENCH PRO
64.6%GPT-5.6 SOL ON SWE-BENCH PRO; THE SPREAD VERIFIED HAD FLATTENED
59.4%SHARE OF THE VERIFIED TASKS OPENAI AUDITED THAT HAD FLAWED TESTS, BEFORE IT WITHDREW
The coding number everyone quotes says AI has nearly solved software engineering: Claude Opus 5 scores 96% on SWE-bench Verified. Move to the benchmark built to resist contamination and the top model sits at 80.3%, and GPT-5.6 Sol lands at 64.6%.
02THE CHECK

THE CLAIM. coding is nearly a solved problem for frontier models. The number carrying that story is SWE-bench Verified, where the top of the board is now packed near 96%, Claude Opus 5 in front at 96%, a figure that reads as almost every real bug fixed.

THE CHECK. that benchmark is worn out and leaky. benchlm's own leaderboard note says the score is 'nearing saturation for frontier models,' meaning it can no longer tell the best models apart, and OpenAI publicly stopped using SWE-bench Verified in February 2026 after an audit found, in the auditors' words, that '59.4% of audited problems contain flawed test cases that reject correct solutions.' The honest successor is SWE-bench Pro, which, unlike Verified, 'uses actively maintained repositories with no public ground-truth leakage.' On Pro the ceiling falls to 80.3%, and the spread that Verified had flattened reopens: GPT-5.6 Sol, near the top of every coding conversation, sits at 64.6%. The models are strong. Solved is a word the quoted benchmark can no longer support.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The 96% SWE-bench score labs quote is SWE-bench Verified, which is saturated and which OpenAI abandoned after finding 59.4% of audited tasks had flawed tests. On the contamination-resistant SWE-bench Pro the top model scores 80.3% and GPT-5.6 Sol scores 64.6%. Coding is not solved; the ruler is just leaky."

Every coding launch this year leans on one benchmark, SWE-bench Verified, a set of real GitHub issues a model has to fix so its patch passes the repository's tests. As of the August 18 leaderboard, 'Claude Opus 5 leads the SWE-bench Verified leaderboard on BenchLM's August 2026 update with 96%,' and

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

THAT IS THE RECORD FOR ISSUE #13. NEXT VERDICT DROPS 9PM AEST.