DeepSeek's chart said its flagship jumped 49.9 points. The referee showed up and moved the index by one.
Independent numbers now exist, and they shrink the chart: Terminal-Bench lands at 79 against the vendor's 87.9, and the intelligence composite ticks up a single point.
"DeepSeek's V4-Pro-0813, quietly made the official flagship on August 12, delivers 'significantly enhanced agent capabilities', with its chart showing DeepSWE up 49.9 points and Terminal Bench 2.1 at 87.9." [SOURCE ↗]
THE CLAIM. DeepSeek's official flagship endpoint now serves V4-Pro-0813, and the company's chart shows agentic scores exploding: DeepSWE from 12.8 to 62.7, CyberGym from 52.7 to 83.3, Terminal Bench 2.1 at 87.9.
THE CHECK. at launch, every number was DeepSeek grading DeepSeek against its own retired preview. Within a day the independent record filled in, and it reads smaller: Artificial Analysis scores the model 53 on its Intelligence Index, one point above DeepSeek's own cheaper Flash, and measures Terminal-Bench v2.1 at 79, well under the chart's 87.9 and ten points behind Claude Opus 5. The direction is real, the drama is not. SCMP adds that developers are underwhelmed and disappointed in the pricing, even as cybersecurity researchers are impressed.
On August 12 the deepseek-v4-pro endpoint on DeepSeek's API quietly began serving a new build, DeepSeek-V4-Pro-0813, ending a preview window that had run since April 24. There was no blog post and no launch thread, just an updated documentation page, a pricing table, and a brief website statement pr
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 7 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.