GET THE AUTOPSY ➔

Issue #7

THURSDAY 13 AUGUST 2026 · 12 CLAIMS CHECKED · 0 SURVIVED THE RECEIPTS · ISSUE 7 OF 17

Grok 4.6 posts a 1753 and a number one. One is a real benchmark. The other is a tweet.

The 1753 is genuine and Grok is genuinely cheap. But it is tied for third overall, and the number one Musk tweeted is missing from both xAI's published table and Databricks' own leaderboard.

01THE CLAIM
"Grok 4.6 reaches 1753 ELO and is #1 on the Databricks leaderboard, at half the price of rival frontier models." [SOURCE ↗]
TRUE, BUT9 SOURCES · LIVE 2026-08-28
ELON MUSK / XAI TRACK RECORD1 CLAIM · 40/100 BS RATE →
1753ELO, ON ONE BENCHMARK
61INTELLIGENCE INDEX, TIED 3RD BEHIND OPUS + FABLE
$0.84COST PER TASK, THE REAL WIN
Grok 4.6 posts a 1753 and a number one. One is a real benchmark. The other is a tweet.
02THE CHECK

THE CLAIM. Musk says Grok 4.6 reaches 1753 ELO and is number one on Databricks' OfficeQA benchmark, at half the price of rivals.

THE CHECK. the 1753 is real, and it comes from Artificial Analysis, not just Musk. But it is one benchmark, GDPval-AA v2, where Grok sits behind Claude Opus 5 and inside overlapping confidence intervals with Fable 5 and Qwen. On the overall Intelligence Index it scores 61, tied with GPT-5.6 Sol for third, behind Opus at 63 and Fable at 62.

THE TWIST. the number one Musk tweeted is an OfficeQA Pro V2 run with Databricks' Genie harness, and in one analysis's words that number 'is not in SpaceXAI's published table.' xAI's model card does show a number one on the older OfficeQA Pro v1, 63.2 percent, but xAI ran that itself, and Databricks' own leaderboard does not list Grok 4.6 at all. The part that actually holds is the boring part: Grok 4.6 is cheap, about $0.84 a task, with list prices of $2 in and $6 out per million tokens against $5 and $25 for Opus 5.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"1753 is real and it is tied for third. The number one is a benchmark run Databricks has not published, and Grok 4.6 is not on Databricks' own leaderboard."
DEEP DIVE · THE FULL AUTOPSY

What actually happened

SpaceXAI, formerly xAI, released Grok 4.6, and Elon Musk posted that it 'reaches 1753 ELO' and is number one on the Databricks leaderboard, at half the price of rival frontier models. The launch was the day's biggest, and unusually for a Musk number, the headline figure is real.

Artificial Analysis, an independent evaluator, confirms it: Grok 4.6 'achieves a GDPval-AA v2 Elo of 1753.' That is a genuine result from a third party, not a self-report. So far, so good.

Why we rate this needs_context

Three things the tweet leaves out.

First, 1753 is one benchmark. GDPval-AA v2 measures real-world agentic tasks, and on it Grok sits 'behind only Claude Opus 5,' with confidence intervals that overlap Claude Fable 5 and Qwen3.8 Max. Overlapping intervals means statistically tied, not beaten.

Second, the overall picture is a tie for third, not first. On Artificial Analysis's headline Intelligence Index, Grok 4.6 'scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62).' A tie for third, by a nose, is a fine result. It is not a number one.

A single benchmark is a spotlight, and you can always stand where the spotlight is.

Third, and sharpest: the number one Musk tweeted is a specific run, OfficeQA Pro V2 with Databricks' Genie harness, and that result has no paper trail. As one launch analysis puts it flatly, 'That number is not in SpaceXAI's published table.' The advice from the same piece: 'Treat it as a founder claim until Databricks or a third party reproduces it.' To be precise about what does exist: xAI's own model card (the Grok 4.6 PDF on media.x.ai) shows a number one on the older OfficeQA Pro v1, 63.2 percent against 60.9 for Claude Opus 5, but xAI ran that benchmark itself. Databricks' official OfficeQA leaderboard, last updated August 2, does not list Grok 4.6 at all.

The part that actually holds

Strip the ranking theater and there is a genuine win underneath: cost. Grok 4.6 runs at roughly $0.84 per task on Artificial Analysis's harness, the same as Kimi K3, and its list prices, $2 in and $6 out per million tokens, sit well below Opus 5 at $5 and $25 and GPT-5.6 Sol at $5 and $30. If your decision is price-per-capability, Grok 4.6 is a real option, and that is the claim xAI could have led with honestly.

The steelman, and why the framing still needs the caveat

The fair case: the jump from Grok 4.5 to 4.6 is large and real, the 1753 is independently confirmed, and being third in the frontier cluster while undercutting everyone on price is a legitimately strong position. Agreed. The problem is only the compression: 'third, tied, on one benchmark, plus an unverified Databricks crown' becomes '1753, number one.' The numbers are real; the ranking is chosen.

The mechanism

Cross-model comparisons in xAI's table use the best of each rival's self-reported or publicly available results, not one lab running all four models in the same harness with the same tools and settings. That means the table is assembled from different conditions, which is exactly how a third-place model ends up sounding like a leader. The founder tweet does the rest.

What to do with this

  • Separate the three claims: 1753 (real, one benchmark), number-one-Databricks (unverified tweet), and half-price (real on list prices). Believe the first and third, hold the second.
  • Distrust any leaderboard result assembled from each model's 'best publicly available score.' Wait for one evaluator to run them all under identical conditions.
  • If you are choosing on cost, Grok 4.6 is worth testing on your own workload. If you are choosing on raw intelligence, the top of the cluster is still Opus and Fable.
04YOUR MOVE ⚡ WHAT IGNORING THIS COSTS

The leaderboard game is won by choosing the leaderboard. A real number and a real strength, price, get bundled with an unverified crown so the whole thing reads as frontier dominance. Buy the cost efficiency, not the coronation, and wait for one lab to run every model in one identical harness before you believe any ranking.

05🔮 OUR CALL · ON THE RECORD 2026-08-13

When an independent evaluator runs Grok 4.6 and its rivals in one identical harness, Grok lands in the frontier cluster on price and mid-pack on raw intelligence, not alone at number one. The cost win survives, the crown does not. Hold us to it.

Flips to a real lead if Databricks or an independent party reproduces the OfficeQA Pro V2 number-one result, or if Grok 4.6 tops a major public benchmark run under identical conditions for every model.

RECEIPTS (9) · CONFIDENCE HIGH

every URL below answered a live HTTP check before publish · sweep 2026-08-28

  • artificialanalysis.ai · "achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5"
  • artificialanalysis.ai · "It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62)"
  • explainx.ai · "That number is not in SpaceXAI's published table."
  • explainx.ai · "Treat it as a founder claim until Databricks or a third party reproduces it."
  • cryptobriefing.com · "Grok 4.6 scored 1,753 on GDPVal AA v2 compared with 1,728 for GPT 5.6 Sol Max"
  • github.com · "Opus 5 and Antigravity Gemini 3.1 Pro (denoted by *) results updated August 2 2026"
  • github.com · "Headline results on OfficeQA Pro (N=133), followed by OfficeQA Pro V2 (N=90)"
  • eesel.ai · "charging $2/$6 per million tokens against Sol's $5/$30"
  • eesel.ai · "Claude Opus 5 is $5/$25"

At Black Hat, Google warned of AI-powered attackers. Google's own tracker counts the breakthroughs: zero.

The 'new entry path' on stage is stolen session cookies, a technique older than the models. The genuinely new thing is one AI-built zero-day, and Google caught it before anyone got hit.

01THE CLAIM
"Threat groups are using AI models to develop new exploits and gain entry into corporate networks through new techniques, dragging defenders into an AI-based arms race" [SOURCE ↗]
TRUE, BUT5 SOURCES · LIVE 2026-08-28
GOOGLE THREAT INTELLIGENCE GROUP TRACK RECORD1 CLAIM · 40/100 BS RATE →
0BREAKTHROUGH CAPABILITIES IN GTIG'S OWN TRACKER
1AI-BUILT ZERO-DAYS OBSERVED IN THE WILD, EVER
MAY 11WHEN GOOGLE LOGGED THAT ONE
At Black Hat, Google warned of AI-powered attackers. Google's own tracker counts the breakthroughs: zero.
02THE CHECK

THE CLAIM. at a Black Hat media briefing, Google Threat Intelligence Group and Accenture said threat actors are using frontier and open-weight AI models to develop new exploits and new entry paths into corporate networks. THE CHECK: Google's own written AI Threat Tracker says it 'has not yet observed' actors achieving breakthrough capabilities that alter the threat landscape. The headline 'new technique' is stealing tokens, cookies and session IDs, which infostealers did long before AI. The one confirmed AI-built zero-day was disclosed May 11 and closed before it was used.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"'Which of these entry paths did not exist before AI?' If the answer is stolen session cookies, that path is older than the models selling the panic."

On August 12 at Black Hat USA in Las Vegas, Ryan Whelan of Accenture and John Hultquist of Google Threat Intelligence Group gave a media briefing on AI-enabled attackers. The message, as reported by Cybersecurity Dive: threat actors are using frontier and open-weight models to 'develop new exploits,

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 5 sources with quotes and screenshots, and our on-record call.

Claude moved a bound that had not moved in years: 41.6 to 67.2. Journal reviews of the paper so far: zero.

The math looks real: formally checked, examined favorably by Conrey and Goldston. The autonomy story rides on a model nobody outside Anthropic can run.

01THE CLAIM
"An unreleased research version of Claude raised the proven lower bound for the fraction of Riemann zeta zeros satisfying the hypothesis from 41.6% to 67.2%, working autonomously across two sessions with about 60 subagents." [SOURCE ↗]
TRUE, BUT5 SOURCES · LIVE 2026-08-28
ANTHROPIC TRACK RECORD39 CLAIMS · 38/100 BS RATE →
67.2%THE NEW BOUND, FORMALLY CHECKED, AWAITING REVIEW
650IDEAS THAT FAILED BEFORE ONE WORKED
0JOURNAL PEER REVIEWS OF THE PAPER SO FAR
Claude moved a bound that had not moved in years: 41.6 to 67.2. Journal reviews of the paper so far: zero.
02THE CHECK

THE CLAIM. Anthropic says an unreleased research Claude improved the longstanding lower bound for the fraction of Riemann zeta zeros that satisfy the hypothesis from 41.6% to 67.2%, burning 650 failed ideas and 31 million output tokens across two agentic sessions.

THE CHECK. the evidence is unusually strong for a same-day lab announcement. Two Anthropic mathematicians validated the paper, Claude produced a formally verifiable proof, and outside experts Brian Conrey and Dan Goldston examined it favorably. But the paper has not been submitted to a journal, a short-notice examination is not peer review, and the whole run happened on an unidentified model nobody can rerun. The theorem is checkable. The capability story around it is not.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The 67.2 is probably real math. The two-session autonomy story comes from a model nobody outside Anthropic can run."

On August 10 Anthropic published a research note saying an unreleased version of Claude had improved one of the classic partial results around the Riemann hypothesis. The hypothesis, open since 1859, says every nontrivial zero of the zeta function sits on the critical line. What mathematicians can a

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 5 sources with quotes and screenshots, and our on-record call.

CoreWeave says it has $104 billion in revenue backlog, up 246 percent. The same release shows a $626 million quarterly loss, and the credit market is pricing a coin flip on whether the company survives five years to collect.

The backlog is real, contracted, and eight times this year's revenue. The question is not whether the number exists. It is who is alive when it converts.

01THE CLAIM
"CoreWeave's revenue backlog was approximately $104 billion as of June 30, 2026, up 246% year over year, and that excludes more than $25 billion of net new customer commitments added in early Q3." [SOURCE ↗]
TRUE, BUT4 SOURCES · LIVE 2026-08-28
COREWEAVE Q2 2026 EARNINGS RELEASE TRACK RECORD1 CLAIM · 40/100 BS RATE →
$104 billionrevenue backlog as of June 30, 2026, per the Q2 earnings release (up 246% year over year)
$25 billionnet new customer commitments added in early Q3, excluded from the $104B and pushing the circulating headline total to roughly $129 billion
$12.4-13.2 billionCoreWeave's own full-year 2026 revenue guidance, roughly one eighth of the backlog headline
$626 millionQ2 2026 net loss, up from $290 million a year ago, reported in the same release as the backlog
$640 millionQ2 2026 net interest expense per the earnings release income statement, larger than the quarter's operating result
~50%five-year default probability priced by the credit default swap market going into the call, per TechTimes
$35-39 billion2026 capital expenditure guidance, about three dollars of capex for every dollar of guided 2026 revenue
CoreWeave says it has $104 billion in revenue backlog, up 246 percent. The same release shows a $626 million quarterly loss, and the credit market is pricing a coin flip on whether the company survives five years to collect.
02THE CHECK

THE CLAIM. per CoreWeave's Q2 2026 release, revenue backlog was approximately $104 billion as of June 30, 2026, up 246% year over year, and a footnote adds it does not include more than $25 billion of net new customer commitments added in early Q3. Stack the footnote on the headline and you get the $129 billion figure now circulating.

THE CHECK. the same release guides full-year 2026 revenue to $12.4 to 13.2 billion, about one eighth of the backlog, shows a $626 million net loss for the quarter, $640 million in net interest expense, and $35 to 39 billion of planned capex. The backlog is, by the company's own definition, subject to the satisfaction of delivery and availability of service requirements. And per TechTimes, the credit default swap market had priced a roughly 50% five-year default probability going into the call.

THE PATTERN. backlog is the AI infrastructure era's favorite unit of account because it is the biggest number a money-losing company can print without an auditor slowing it down.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The backlog is 104 billion. This year's revenue guide is 13 at best, the quarter lost 626 million, and interest ate 640. Ask who finances the other seven years."

On August 11, CoreWeave reported second quarter 2026 results and the stock jumped double digits after hours. Revenue was genuinely strong: $2.6 billion, up 112 percent year over year. But the number that did the work was in the highlights section: "Revenue backlog was approximately $104 billion as o

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 4 sources with quotes and screenshots, and our on-record call.

AI servers are now more than half of Foxconn's revenue, a first. Foxconn's own release shows what half the revenue buys: a 6.12 percent gross margin, thinner than last year.

The AI share of revenue went up. The gross margin went down. For a contract manufacturer, both numbers are measuring the same thing: the price of Nvidia chips passing through the building.

01THE CLAIM
"AI servers now generate more than half of Foxconn's revenue, crossing 51% of Q2 2026 sales for the first time, alongside record second-quarter revenue, operating profit, and net profit, with strong AI-driven growth guided for Q3 and 2027." [SOURCE ↗]
TRUE, BUT3 SOURCES · LIVE 2026-08-28
FOXCONN TRACK RECORD1 CLAIM · 40/100 BS RATE →
51%share of Q2 2026 revenue from the cloud and networking division that builds AI servers, the first time it passed half
6.12%Q2 2026 gross profit margin per Foxconn's own release, down 0.21 percentage points from a year earlier
3.75%Q2 2026 operating profit margin, improved on operational leverage but still under four cents on the dollar
2.37%Q2 2026 net profit margin, slightly below last year even in the record quarter
NT$2.53 trillionQ2 2026 consolidated revenue, up 41% year on year, a second-quarter record
NT$59.97 billionQ2 2026 net profit, up 35% year on year, also a second-quarter record
39%Morgan Stanley's projected 2026 high-end AI rack market share for Foxconn, down from 51% in 2025, per the Wall Street Journal
AI servers are now more than half of Foxconn's revenue, a first. Foxconn's own release shows what half the revenue buys: a 6.12 percent gross margin, thinner than last year.
02THE CHECK

THE CLAIM. Foxconn's Q2 2026 results, reported August 12, show AI servers crossing 51% of revenue for the first time, with revenue up 41% to NT$2.53 trillion and record net profit of NT$59.97 billion, and management guiding high double-digit cloud growth for Q3.

THE CHECK. the same release prints gross, operating, and net margins of 6.12%, 3.75%, and 2.37%. Gross margin fell 0.21 points in the record quarter. Net margin slipped too. Per the Wall Street Journal, Morgan Stanley expects Foxconn's high-end rack share to fall to 39% this year from 51% in 2025, and CEO Michael Chiang named TSMC's sold-out CoWoS packaging, which Foxconn does not control, as the ceiling on 2027 growth.

THE PATTERN. revenue-mix milestones are the AI supply chain's favorite flex because revenue is the one line that inflates automatically when the components you assemble get more expensive.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Half of Foxconn's revenue is AI servers now, and the gross margin on all of it is 6.12 percent, down from last year. That is not an AI transformation. That is expensive cargo."

On August 12, Foxconn, the contract manufacturer formally known as Hon Hai, reported a second quarter that broke records in every headline line: revenue of NT$2.53 trillion, up 41 percent year on year, with operating profit and net profit both at all-time second-quarter highs. Net profit came in at

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 3 sources with quotes and screenshots, and our on-record call.

OpenAI's hacking model found two real bugs in Chrome. The word zero-day got added in post.

The V8 flaws are genuine and Google patched them, at High severity, weeks before OpenAI said a word. The zero-day drama and the 95% score are packaging.

01THE CLAIM
"OpenAI's new GPT-5.6-Cyber model found two previously undocumented Chrome V8 flaws (billed as zero-days) and completes 95% of advanced cybersecurity requests the standard model refuses." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-28
OPENAI TRACK RECORD28 CLAIMS · 39/100 BS RATE →
95%PROMPTS ANSWERED, NOT ANSWERED RIGHT
1.5%SAME PROMPTS, STOCK GPT-5.6
8.8CVSS OF THE CHROME FIND, HIGH NOT CRITICAL
OpenAI's hacking model found two real bugs in Chrome. The word zero-day got added in post.
02THE CHECK

THE CLAIM. OpenAI says GPT-5.6-Cyber found bugs no one had documented before, including two Chrome V8 flaws that chain into a sandbox escape, and answers 95% of advanced cyber prompts the stock model refuses.

THE CHECK. the bugs are real. Google patched them as CVE-2026-15903, severity High, weeks before the announcement, and no exploit has ever been seen in the wild. A zero-day is a flaw attackers use before defenders can respond. These were found by the defender and were dead on arrival. The 95% counts how often the model agrees to answer, not whether the answer works. And on OpenAI's own report-writing eval, Cyber scores worse than plain Sol.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The Chrome bugs are real and were patched weeks before the announcement. The 95% measures how often it answers, not how often the answer works."

On August 10 OpenAI expanded Daybreak, its vetted-access security program, and introduced GPT-5.6-Cyber, a variant of GPT-5.6 Sol trained for exploit development and vulnerability research. Access is restricted to the program's Red tier. The launch rides on two artifacts. One: the model "found bugs

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

The company that counts deepfake victims launched its deepfake detector the same day the count came out.

Grok's deepfake problem is real, courts keep saying so. The 87% and the 15,736 come from reading 1,760 news stories, tallied by the firm selling the fix.

01THE CLAIM
"Grok generated 87% of the synthetic files behind documented deepfake attacks in H1 2026, which hit 15,736 confirmed victims across 821 attacks" [SOURCE ↗]
TRUE, BUT4 SOURCES · LIVE 2026-08-28
RESEMBLE AI TRACK RECORD1 CLAIM · 40/100 BS RATE →
1,760NEWS REPORTS BEHIND THE ENTIRE DATASET
87%OF FILES THEY COULD COUNT, NOT OF ALL DEEPFAKES
0DAYS BETWEEN THE SCARY REPORT AND THE PRODUCT LAUNCH
The company that counts deepfake victims launched its deepfake detector the same day the count came out.
02THE CHECK

THE CLAIM. Resemble AI's H1 2026 Deepfake Threat Report says Grok accounts for 87% of synthetic files in documented deepfake attacks, with 15,736 victims across 821 attacks. Headlines upgraded that to 'confirmed victims'. THE CHECK: the dataset is 1,760 news reports, the 87% covers only files researchers could count and attribute, and Resemble's own page calls the victim figure 'a conservative floor'. The report shipped the same day Resemble launched DETECT-World, its new detection product. Real problem, court-documented. Precise-sounding numbers, clipping-derived.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"'87% of what denominator?' The answer is files countable from press coverage of 821 incidents, not deepfakes made anywhere by anyone."

On August 12, voice-AI and detection company Resemble AI published its H1 2026 Deepfake Threat Report. The headline stats: 821 verified deepfake attacks in six months, at least 15,736 victims, roughly 3.46 million synthetic files, and one generator, Grok, behind 87% of the files. By the morning of A

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 4 sources with quotes and screenshots, and our on-record call.

NVIDIA built a model that talks four times faster. The work arrives 30 percent sooner.

The 4x is NVIDIA's own number, measured at the token tap. On NVIDIA's own chart, the agent's finish line moves 30 percent. Nothing independent confirms either.

01THE CLAIM
"NVIDIA's Nemotron 3.5 Lightning delivers up to 4x the output speed of similar-sized models, putting it on the accuracy-speed Pareto frontier for agent workloads." [SOURCE ↗]
TRUE, BUT4 SOURCES · LIVE 2026-08-28
NVIDIA TRACK RECORD8 CLAIMS · 35/100 BS RATE →
4xOUTPUT SPEED, 'UP TO', SAYS THE VENDOR
30%FASTER THE ACTUAL TASKS FINISH, SAME VENDOR'S CHART
86%PINCHBENCH ACCURACY, THE PART THAT HELD
NVIDIA built a model that talks four times faster. The work arrives 30 percent sooner.
02THE CHECK

THE CLAIM. Nemotron 3.5 Lightning, NVIDIA's new 30B agent model with 3B active parameters, generates output up to 4x faster than similar-sized models and wins the accuracy-speed Pareto frontier.

THE CHECK. three sentences after the 4x, NVIDIA's own blog concedes that agent efficiency comes down to completed work, not token speed, and its own PinchBench chart shows tasks finishing 30% faster than Qwen3.6 35B. The 4x is 'up to', measured by the vendor, and Artificial Analysis's page for the model lists no output speed at all. A 4x engine revs; a 30% car arrives. Both numbers are NVIDIA grading NVIDIA.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Tokens per second is the engine revving. Task completion is the car arriving. NVIDIA's own chart says 30 percent."

On August 11 NVIDIA released Nemotron 3.5 Lightning, a 30B-parameter open Mixture-of-Experts model that keeps 3B parameters active per token, shipped under the permissive OpenMDW license with weights, training data, and recipes. It is built for the grunt-work layer of agent systems: the tool calls,

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 4 sources with quotes and screenshots, and our on-record call.

Google put a screen-tapping agent on a billion phones. The same kind just went wrong on a gym.

Gemini on the Pixel 11 runs your errands by reading your screen and tapping like you would. That is the same kind of autonomous agent that just went wrong in the wild.

01THE CLAIM
"Gemini on Pixel 11 can order groceries, book rides, manage reservations, and call businesses across third-party apps, including apps with no native Gemini integration." [SOURCE ↗]
TRUE, BUT5 SOURCES · LIVE 2026-08-28
GOOGLE TRACK RECORD16 CLAIMS · 42/100 BS RATE →
$899PIXEL 11 STARTING PRICE
0APP INTEGRATION NEEDED, IT READS THE SCREEN
US ONLYWHERE THE AGENT WORKS AT LAUNCH
Google put a screen-tapping agent on a billion phones. The same kind just went wrong on a gym.
02THE CHECK

THE CLAIM. Gemini on the Pixel 11 will order groceries, book rides, make reservations, and call businesses across third-party apps, including apps with no Gemini integration.

THE CHECK. the mechanism is the tell. Gemini 'reads what's on your screen, picks out things like text fields, menus, and search bars, and interacts with them in real time just like you would.' No integration, no API. It taps like a human, which is exactly the agent class behind this week's real-world agent failures, like the OpenClaw gym booking that deleted a stranger's spot.

THE TWIST. it is not the shipped, do-anything feature the keynote implies. It is US-only, the connected apps arrive 'over the next few weeks', and even a friendly review says 'you're still faster than Gemini' and reaches for the failed Humane AI Pin and Rabbit R1. A screen-reading agent with your logins is a convenience and a blast radius at once.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"It has no integration. It reads your screen and taps. That is the same kind of agent that just went wrong on a gym."

At Made by Google 2026 on August 12, Google announced that Gemini on the Pixel 11, which starts at $899, can complete real tasks across third-party apps. Per TechCrunch, 'U.S.-based users will be able to order groceries, book rides, or get coffee,' and Gemini can call businesses on your behalf for r

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 5 sources with quotes and screenshots, and our on-record call.

Alibaba kept its open-weights promise, then swapped the fine print.

The Qwen Max weights shipped on time, under a name nobody was watching, with a license that is not the free one everyone assumed.

01THE CLAIM
"Alibaba said open weights for Qwen3.8-Max and Qwen3.8-27B would be published the week of August 10, 2026 on Hugging Face and ModelScope." [SOURCE ↗]
TRUE, BUT4 SOURCES · LIVE 2026-08-28
ALIBABA QWEN TEAM TRACK RECORD2 CLAIMS · 40/100 BS RATE →
1 OF 2PROMISED MODELS SHIPPED IN THE WINDOW
$50MREVENUE GATE IN THE NEW CUSTOM LICENSE
AUG 10THE WEEK THE WEIGHTS WERE PROMISED
Alibaba kept its open-weights promise, then swapped the fine print.
02THE CHECK

THE CLAIM. Alibaba said the open weights for Qwen3.8-Max and Qwen3.8-27B would be published the week of August 10 on Hugging Face and ModelScope.

THE CHECK. the Max weights did land inside the window, but as Qwen3.8-2.4T-A95B, a spec name, while every countdown site refreshed the still-private Qwen3.8-Max repo and saw nothing. And the license is not Apache-2.0 like the last Qwen. It is a custom license with a revenue gate: past $50 million you need a separate license from Qwen.

THE TWIST. half the promise is still unshipped. The Qwen3.8-27B repo is private, HTTP 401, as of today. So 'open weights, this week' became one model, under a name no one was watching, on terms no one was promised.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The weights shipped, spec-named, under a $50M revenue gate. The 27B is still private. Precedent is not a promise."

On August 3, Alibaba's Qwen team promised the open weights of two models, Qwen3.8-Max and Qwen3.8-27B, would be published 'next week,' the week of August 10, on Hugging Face and ModelScope. It was a real commitment, quoted verbatim across the coverage.

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 4 sources with quotes and screenshots, and our on-record call.

The 'autonomous AI strike on a nuclear agency' was 85 cracked passwords at a personnel office.

Real attack, real agents, real loot. The nuclear part is a target list that got scanned, and the autonomy shipped with a human operator attached.

01THE CLAIM
"Open-source AI agents ran a four-day autonomous cyberattack that breached Taiwan's nuclear safety agency and seven energy companies" [SOURCE ↗]
TRUE, BUT3 SOURCES · LIVE 2026-08-28
DREAM TRACK RECORD1 CLAIM · 40/100 BS RATE →
0CONFIRMED NUCLEAR-AGENCY COMPROMISES
85GOVERNMENT ACCOUNTS CRACKED
2,564PERSONNEL RECORDS EXFILTRATED
The 'autonomous AI strike on a nuclear agency' was 85 cracked passwords at a personnel office.
02THE CHECK

THE CLAIM. open-source AI agents ran a four-day autonomous strike that breached Taiwan's nuclear safety agency. THE CHECK: Dream's own report confirms compromises at a government department (85 cracked accounts, 2,564+ personnel records) and lists the nuclear agency only as an expansion target the agents scanned. Dream itself writes 'what appears to be a near-autonomous attack', and Taiwan's digital ministry describes a hybrid operation combining conventional hacking with AI agents.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"'Which systems at the nuclear agency were compromised, and who confirmed it?' If the answer is a target list, you are reading about a scan, not a breach."

Between July 1 and July 4, someone pointed a multi-agent AI framework at Taiwan's government. The framework, stitched together from two open-source agent systems called Hermes and OpenClaw, ran 12 documented attack waves with up to eight sub-agents working in parallel, labeled Agent A through Agent

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 3 sources with quotes and screenshots, and our on-record call.

Tencent says the AI compute it is buying will convert into revenue going forward. This quarter it converted into a 176 percent capex jump, negative free cash flow, and 10.5 billion yuan of losses on the new AI products.

Revenue grew 11 percent and the ads really are AI-assisted. The AI division itself is a money pit wearing a growth story.

01THE CLAIM
"Tencent's surging AI investment is already driving revenue growth, and the compute it is buying will convert usage of our applications and models into revenue going forward, per chairman and CEO Pony Ma at Q2 2026 results." [SOURCE ↗]
TRUE, BUT4 SOURCES · LIVE 2026-08-28
PONY MA HUATENG TRACK RECORD1 CLAIM · 40/100 BS RATE →
176%year-over-year jump in Q2 capital expenditure, to RMB 52.8 billion, driven by AI infrastructure
RMB 13.8 billionnegative free cash flow for the quarter, the cash cost of the AI buildout landing now
RMB 10.5 billiondrag on non-IFRS operating profit from new AI products this quarter
RMB 58.2 billionnet cash position at quarter end, down from RMB 146.9 billion on March 31, 2026
RMB 204.8 billionQ2 revenue, up 11% year over year, the growth management attributes partly to AI-driven ad targeting
19%non-IFRS operating profit growth excluding new AI products, versus 9% with them included
190%growth in operating capex specifically (RMB 51.8 billion), the earnings-call cut of the same spending surge
Tencent says the AI compute it is buying will convert into revenue going forward. This quarter it converted into a 176 percent capex jump, negative free cash flow, and 10.5 billion yuan of losses on the new AI products.
02THE CHECK

THE CLAIM. presenting Q2 2026 results on August 12, Tencent chairman Pony Ma said the surging spend reflects compute procurement the company expects to convert usage of our applications and models into revenue going forward, with 11% revenue growth and 22% marketing services growth credited partly to AI-driven ad targeting.

THE CHECK. the same results show what the conversion costs today. Capex up 176% to RMB 52.8 billion. Free cash flow negative RMB 13.8 billion. New AI products dragged non-IFRS operating profit by RMB 10.5 billion; strip them out and profit growth doubles from 9% to 19%. Net cash fell from RMB 146.9 billion to RMB 58.2 billion in a single quarter, and the ADR dropped over 5% on the print.

THE PATTERN. revenue going forward is the AI era's favorite tense. The spending is always in the present, the conversion is always in the future, and the gap between the two is called conviction.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Their own numbers: capex up 176 percent, free cash flow minus 13.8 billion yuan, new AI products lost 10.5 billion. The revenue is going forward. The cash is going now."

On August 12, Tencent reported second quarter 2026 results: revenue of RMB 204.8 billion, up 11 percent year over year, beating estimates. Marketing services grew 22 percent on AI-driven ad targeting. Chairman and CEO Pony Ma framed the quarter's defining number, a 176 percent surge in capital expen

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 4 sources with quotes and screenshots, and our on-record call.

THAT IS THE RECORD FOR ISSUE #7. NEXT VERDICT DROPS 9PM AEST.