Two new benchmarks agree: the best AI models in the world clear fewer than half of a hard benchmark of real analyst tasks. Claude Fable 5 tops the frontier at 49.2%. The context is who built the tests, and who sells the fix.
Samaya AI's FrontierFinance and Vals AI's Finance Agent v2 both put frontier models below the low-50s on realistic analyst work, and the ceiling is genuine and separately reproduced. But both benchmarks come from companies selling finance AI, each is topped by its maker's own system, Samaya's wins its own board at 56%, and on the hardest use cases even the best system lands at 33%. The wall is real. Read who is charging admission.
"Frontier AI models cannot yet do professional investment analysis: on Samaya AI's new FrontierFinance benchmark every frontier model clears fewer than half of real analyst tasks, and Samaya's own in-house system beats them all, a ceiling an independent Vals AI benchmark reproduces." [SOURCE ↗]

THE CLAIM. frontier AI cannot do professional investment analysis. On Samaya AI's FrontierFinance, a public benchmark of 220 expert queries, every frontier model scores below 50%, with Claude Fable 5 best at 49.2%, GPT-5.6 Sol at 46.8%, and Samaya's own system ahead of all of them at 56%.
THE CHECK. the ceiling is real and it is not a fluke. A separate benchmark from Vals AI, which builds evaluations rather than investment products but still profits when models fall short, reaches the same place: its top model, GPT-5.5, hits roughly 52%, with the frontier clustered in the high-40s to low-50s. Two separate teams, two methods, one wall. What the headline underplays is the shape of the sales floor around it. FrontierFinance is Samaya's own benchmark and Samaya's own system tops it, which is the oldest move in the book, a vendor grading an exam its product is built to pass. Even that winning system clears only 56%, and on the hardest categories, Screening and Discovery and Sector and Macro, the best of everything reaches 33% and 39%. The ceiling is the story. So is the fact that the people ringing the bell are selling the ladder.
The pitch of the year in enterprise AI is the autonomous analyst: a model that reads the filings, builds the model, and writes the memo. Two new benchmarks just measured how close that is, and the answer is not close. On FrontierFinance, a public benchmark of 220 expert-crafted queries released by S
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 7 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.