GET THE AUTOPSY ➔

Issue #10

SUNDAY 16 AUGUST 2026 · 13 CLAIMS CHECKED · 1 SURVIVED THE RECEIPTS · ISSUE 10 OF 17

OpenAI emailed every free ChatGPT user in Europe: ads are coming, and they will use only three data points. Their own privacy policy lists seven.

The email undersells the data collection by half. The policy discloses purchase data from advertisers, prior ad interactions, behavioral signals, and contact sync on top of the three the email names.

01THE CLAIM
"Ads are coming to ChatGPT Free/Go in the EEA later this month; selection uses only the current chat topic, general location and device type; Plus/Pro stay ad-free." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-28
OPENAI IRELAND TRACK RECORD1 CLAIM · 40/100 BS RATE →
3data categories the email admits to (topic, location, device)
7data categories the privacy policy actually names for ad targeting
02THE CHECK

THE CLAIM. OpenAI Ireland emailed EEA Free and Go users on August 15 that ads will arrive later this month, selecting on "only" the current chat topic, general location, and device type. Plus and Pro stay ad-free.

THE CHECK. the updated privacy policy lists seven distinct ad-related data flows: conversation context, general location, device type, purchase data received from advertisers, prior ad interactions, behavioral signals, and optional address book sync. Four of those seven do not appear in the email. The word "only" is doing load-bearing work it cannot support.

THE TWIST. OpenAI called this a "test" in the US in February. By August it has rolled through the UK, Mexico, Brazil, Japan, South Korea, and now the entire EEA. A test that expands to 30+ countries is a launch.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"OpenAI says ads use only 3 data points. Their own privacy policy lists 7. That word only is covering for purchase data, ad-interaction history, behavioral signals, and your contacts."
DEEP DIVE · THE FULL AUTOPSY

The full receipt trail

On August 15, OpenAI Ireland Ltd sent an email to every Free and Go user in the European Economic Area and Switzerland. The key sentence: ad selection "starts with the current chat topic, general location and device type only."

That word "only" is the entire story.

What the privacy policy actually says

The updated OpenAI privacy policy (effective February 9, 2026, last revised April 30) describes seven distinct data categories involved in ad delivery:

1. Current conversation topic (in the email) 2. General location (in the email) 3. Device type (in the email) 4. Purchase data received from advertisers (NOT in the email) 5. Prior ad interactions (NOT in the email) 6. Behavioral signals (NOT in the email) 7. Optional device address book data (NOT in the email)

Four of seven data flows are absent from the user-facing email. The email is not lying, but it is presenting fewer than half the relevant facts.

The "test" that became a launch

OpenAI's language has consistently called this a "test." The timeline tells a different story:

  • February 2026: US launch
  • August 11: UK, Mexico, Brazil, Japan, South Korea
  • August 15: Entire EEA + Switzerland

A test that expands to 30+ countries across four continents in six months is a phased global launch with plausible deniability built into the language.

The steelmanned counter

OpenAI could argue the email describes the ad *selection* process (what determines which ad you see), while the policy covers the broader *measurement and infrastructure* stack (what happens after). This is a reasonable distinction if the four omitted categories genuinely do not influence which ad appears. But "behavioral signals" and "prior ad interactions" are selection inputs at every other ad platform. OpenAI has not explained why they would collect these without using them for targeting.

What to do

Read the privacy policy, not the email. If you are a paying Plus or Pro user, you are not affected yet. If you are on Free or Go in Europe, your conversations are now ad-adjacent data. The opt-in for full personalisation is coming; the infrastructure is already in place.

Our call

Within 12 months the opt-in for personalised ads becomes the default, and the three-point framing quietly expands. Every ad-funded platform has followed this exact trajectory: launch with minimal targeting, expand behind policy updates, and let the original "only" language fade from institutional memory.

Flips if OpenAI freezes ad data categories at the current policy level or tightens the policy to match the email.

04YOUR MOVE ⚡ WHAT IGNORING THIS COSTS

If you use ChatGPT Free in Europe, you are about to see ads. The email frames the data collection as minimal. The policy says otherwise. Read the policy, not the email.

05🔮 OUR CALL · ON THE RECORD 2026-08-16

Within 12 months the opt-in for personalised ads becomes the default, and the "only 3" framing quietly expands. The staged global rollout pattern matches every ad-funded platform before it.

Flips to holds if OpenAI tightens the privacy policy to match the email's three-input targeting description. Flips to failed if OpenAI confirms that purchase data, prior ad interactions, behavioral signals, or contact sync are ad-selection inputs.

RECEIPTS (6) · CONFIDENCE MEDIUM

every URL below answered a live HTTP check before publish · sweep 2026-08-28

  • ppc.land · "current chat topic, general location and device type only"
  • ghacks.net · "They can dismiss ads, give feedback, see why an ad was shown, delete ad data with one tap, and manage ad personalization"
  • ppc.land · "Rather than only data going out to advertisers in aggregate form, the updated policy confirms that advertiser-sourced data - including purchase data - now flows in to OpenAI."
  • ghacks.net · "expanding its ChatGPT ads test to the United Kingdom, Mexico, Brazil, Japan, and South Korea"
  • almcorp.com · "advertisers do not receive access to user conversations, chat history, personal details, or memories"
  • leanware.co · "OpenAI says targeting will be contextual rather than based on selling your data"

Anthropic built a model stronger than its flagship. You cannot use it, test it, or check the number.

Model 2 beats Mythos 5 by 12.5 points on CoBench: an internal test, of an internal model, with assessments Anthropic itself calls incomplete. The risk label moved in the same report, for a reason the headlines are fusing with the wrong cause.

01THE CLAIM
"Anthropic's second company-wide Risk Report discloses an unreleased internal Model 2 that is more capable than its shipped flagship Mythos 5, scoring 62.8% on CoBench versus 50.3%, with no plans for external release, alongside raising its catastrophic-misalignment rating from very low to low." [SOURCE ↗]
TRUE, BUT5 SOURCES · LIVE 2026-08-28
ANTHROPIC TRACK RECORD39 CLAIMS · 38/100 BS RATE →
62.8%MODEL 2 ON COBENCH, MEASURED BY ANTHROPIC, INSIDE ANTHROPIC
50.3%THE SHIPPED FLAGSHIP ON THE SAME INTERNAL TEST
0OUTSIDE RUNS: NO WEIGHTS, NO API, NO THIRD PARTY
Anthropic built a model stronger than its flagship. You cannot use it, test it, or check the number.
02THE CHECK

THE CLAIM. Anthropic's August 2026 Risk Report reveals Model 2, an internal system 'somewhat more capable' than Mythos 5 (62.8% vs 50.3% on CoBench), which the company does not plan to release; the same report raises Anthropic's catastrophic-misalignment rating from very low to low.

THE CHECK. the disclosure is real transparency and the capability claim is a closed loop. The score comes from an internal benchmark, run internally, on a model nobody outside Anthropic can touch: no weights, no API, no third-party run. Anthropic itself lowers the claim's confidence in ink, saying it 'has not run all of its typical predeployment assessments and therefore has somewhat less confidence in its beliefs about the model's capabilities'. Meanwhile the risk-label change the headlines attach to this scary stronger model has a different stated cause: increased overall uncertainty from recent cyber-evaluation incident disclosures involving shipped models, not Model 2. The report even notes its concrete task evaluations have 'saturated', meaning the measuring sticks maxed out. A stronger secret model and a raised risk label are both in the report. The causal arrow between them is not.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The 62.8 is an internal score on an internal test of a model nobody outside can run, and Anthropic itself says its assessment is incomplete. The risk label moved for a different reason than the stronger-model headline implies."

Anthropic published its second company-wide Risk Report on August 14, under version 3.4 of its Responsible Scaling Policy. Two disclosures drove the coverage. First, the company raised its rating of catastrophic harm from misalignment in high-stakes settings from very low to low. Second, the report

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 5 sources with quotes and screenshots, and our on-record call.

Free ChatGPT went unlimited last week. Opt out of ads and it stops being unlimited.

The EEA notice says targeting is contextual only: the live conversation topic, general location, device type. The privacy promises all describe internals nobody outside can audit, and OpenAI's own ads page prices the opt-out in daily messages.

01THE CLAIM
"OpenAI told EEA and Swiss users on August 15 that ads arrive in ChatGPT's Free and Go plans later this month, and that the ads are privacy-safe: targeting is contextual only, advertisers receive only aggregate performance data, answers are not influenced, and users can opt out." [SOURCE ↗]
TRUE, BUT7 SOURCES · LIVE 2026-08-28
OPENAI / OPENAI IRELAND TRACK RECORD1 CLAIM · 40/100 BS RATE →
3SIGNALS INSIDE CONTEXTUAL ONLY: TOPIC, LOCATION, DEVICE
5PLANS THAT STAY AD-FREE WITHOUT A CATCH, ALL OF THEM PAID
FEWERDAILY MESSAGES, THE PRICE OF REFUSING ADS
Free ChatGPT went unlimited last week. Opt out of ads and it stops being unlimited.
02THE CHECK

THE CLAIM. ads land in ChatGPT Free and Go in the EEA later this month, and the design is privacy-first: no personalization to start, targeting from 'the current topic of the ChatGPT conversation and limited contextual information such as general location and device type', advertisers see only aggregates, answers stay independent, opt-in required before ads touch chat history or memory.

THE CHECK. the architecture is genuinely better than adtech's default, and every load-bearing promise is about what happens inside a building nobody outside can inspect. Contextual only still means the ad system reads the conversation you are having right now, the medium where, as one analysis puts it, the platform often knows why you want something, not just what. The aggregate-only and answer-independence commitments are unverifiable from outside. And the choice architecture has a meter on it: OpenAI's own ads page offers the Free-tier opt-out 'in exchange for fewer daily free messages', five days after the same company announced unlimited free text chats. Unlimited is the with-ads price. Regulators noticed the shape: it is the consent-or-pay structure European data protection bodies have been circling for two years.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The ads read the chat you are typing, that is what contextual means. The privacy promises are unauditable internals. And the ad-free free tier loses the unlimited chats they announced the week before."

OpenAI Ireland notified Free and Go users across the EEA and Switzerland on August 15 that ads start appearing in ChatGPT later this month. The framework comes from OpenAI's ads page, published August 11: no personalization to start, targeting built from 'the current topic of the ChatGPT conversatio

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 7 sources with quotes and screenshots, and our on-record call.

The great Claude cancellation wave is four people Business Insider talked to. The company says the line is flat.

The anger is real and the on-record quitters are real people. The wave is unmeasured: self-reported cancellations, no denominator, and a vendor that says churn has not moved. The mechanism worry underneath deserves the attention the headline is stealing.

01THE CLAIM
"Claude users are canceling their subscriptions over Anthropic's invisible text watermark, per Business Insider and a wave of weekend aggregator coverage citing dozens of cancellation posts on X." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-28
BUSINESS INSIDER PLUS WEEKEND AGGREGATORS AND VIRAL X POSTS TRACK RECORD1 CLAIM · 40/100 BS RATE →
4CANCELED SUBSCRIBERS BUSINESS INSIDER ACTUALLY SPOKE TO
DOZENSX POSTS CLAIMING CANCELLATION, SELF-REPORTED
0CANCELLATION UPTICK SEEN BY ANTHROPIC, PER ANTHROPIC
The great Claude cancellation wave is four people Business Insider talked to. The company says the line is flat.
02THE CHECK

THE CLAIM. Claude users are canceling en masse over the invisible watermark Anthropic switched on August 2 to comply with the EU AI Act's transparency code.

THE CHECK. count the evidence. Business Insider spoke with four Claude Max subscribers who canceled, and reports 'dozens' of X posts claiming the same; Gizmodo's own phrasing hedges to 'many supposed paying Claude users'. Against that, Anthropic says it has not seen a cancellation uptick. Nobody in this fight has published a real number: the coverage has anecdotes without a denominator, the company has a denial without data, and a cancellation screenshot costs nothing to post. What the wave-framing buries is the legitimate grievance underneath: the mark attaches when Claude merely touches text, and by Anthropic's own description 'a detected mark means Claude touched the text, not that Claude invented it'. Your essay, edited by Claude for commas, carries the same invisible stamp as a fully generated one, and the mark travels with copied text and survives some editing, into a world of employers and schools that will read any detected mark as an authorship verdict. Four people quitting is not a wave. The attribution problem is not a tweet.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Business Insider talked to four cancelers; Anthropic says churn is flat; neither published a number. The real issue is that an edit and an authorship get the same invisible stamp."

Anthropic's invisible text watermark went live globally on August 2, added, per TechCrunch, 'to comply with the EU AI Act's transparency code, which requires AI companies to use systems that make it possible to identify AI-generated content'. This desk covered the launch claims on August 12. The new

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

Anthropic's CEO wants mandatory AI testing before release. Anthropic spent $3.53 million in six months lobbying to shape what that testing looks like.

The safety lab backs a safety regime. The safety lab also tripled its lobbying spend to help write the rules it would need to follow. Startup compliance cost: $344K per deployment against a $10K budget.

01THE CLAIM
"Anthropic CEO Dario Amodei is publicly supportive of mandatory pre-deployment testing for frontier AI models, backing the Trump administration's framework." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-28
DARIO AMODEI TRACK RECORD2 CLAIMS · 20/100 BS RATE →
3.53million USD lobbying per Yahoo Finance receipt
500million USD annual revenue threshold above which the heavy obligations apply (RAISE Act chapter amendments, per Jones Walker receipt)
02THE CHECK

THE CLAIM. Dario Amodei publicly supports mandatory pre-deployment testing for frontier AI models. He compared it to airplane certification: models should be blocked if they fail.

THE CHECK. Anthropic spent $3.53 million on federal lobbying in H1 2026, nearly tripling its full-year 2025 spend. The lobbying focused on export controls, cybersecurity, and AI safety standards. A Harvard Kennedy School study found mandatory compliance transforms startup margins from +13% to -7%, with actual costs hitting $344K per deployment against $10K budgets.

THE TWIST. the company that designs the safety framework also has the resources to comply with it. That is not necessarily corrupt. It is a structure where the regulated help write the rules, and the rules happen to be affordable only to the regulated. The name for that structure is regulatory capture.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Anthropic spent $3.53M in six months lobbying on AI rules while its CEO calls for mandatory testing. The testing burden lands on labs with revenue over $500M, which is Anthropic and its four biggest rivals. The safety lab is shaping the rules it will be graded by."

Dario Amodei published "Policy on the AI Exponential" in June 2026 proposing that frontier models, like airplanes, should require pre-deployment testing by independent third-party auditors. Models failing the test would be blocked from release. He endorsed this again publicly on August 16.

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

DeepSeek wiped $600 billion off Nvidia in January with the cheap-AI story. Today it raised its own prices up to 1,100%.

The disruption narrative has an expiration date. V4-Pro cached input jumps from $0.0036 to $0.044 at peak. Output quadruples. The company calls it resource allocation.

01THE CLAIM
"DeepSeek is raising API prices by up to 1,100% with a new peak/off-peak billing structure, effective August 16." [SOURCE ↗]
VERIFIED6 SOURCES · LIVE 2026-08-28
DEEPSEEK TRACK RECORD3 CLAIMS · 27/100 BS RATE →
1100% maximum increase per Caixin Global (cached input at peak)
3.96USD per 1M output tokens V4-Pro peak (was $0.87)
4x multiplier on V4-Pro output at peak per Engadget
02THE CHECK

THE CLAIM. DeepSeek is raising API prices by up to 1,100% with a new peak/off-peak billing structure, effective August 16.

THE CHECK. confirmed. V4-Pro cached input tokens go from $0.00362 to $0.044 at peak hours (1,113%). Output tokens from $0.87 to $3.96 (355%). V4-Flash sees 57-400% increases. Peak hours: 01:00-04:00 and 06:00-10:00 UTC, which covers prime US working hours.

THE TWIST. seven months ago this company panicked Wall Street by training a frontier model for under $6 million and pricing it at a fraction of American competitors. The cheap-AI disruption story just hit its sell-by date.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"DeepSeek wiped $600B off Nvidia seven months ago with the cheap-AI story. This week it raised API prices by as much as 1,100%. The disruption was real. The old pricing was a customer acquisition cost."

Effective August 16, 2026 at 16:00 UTC:

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

GPT-5.6 Terra scored 69.6 and 64.8 on the same benchmark this week. Nothing changed except whose chart it was.

Every score is real and every score is a different lab privately running a suite that shares only a name. The two available cross-checks disagree by about five points each, and both disagreements point in the house's favor.

01THE CLAIM
"Weekend coverage ranks the week's launches on DeepSWE v1.1 as one leaderboard: GPT-5.6 Terra 69.6, Grok 4.6 High 65.9, Gemini 3.7 Flash 65.3, Muse Spark 1.2 59.3, with each lab's launch chart cited as if the numbers share a scale." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-28
GOOGLE TRACK RECORD16 CLAIMS · 42/100 BS RATE →
69.6%GPT-5.6 TERRA ON GOOGLE'S DEEPSWE CHART
64.8%THE SAME TERRA ON META'S DEEPSWE CHART
59.3%MUSE SPARK PER META. GOOGLE'S CHART SAYS 54.9%
GPT-5.6 Terra scored 69.6 and 64.8 on the same benchmark this week. Nothing changed except whose chart it was.
02THE CHECK

THE CLAIM. the week's coding-agent launches stack neatly on DeepSWE v1.1: GPT-5.6 Terra 69.6, Grok 4.6 High 65.9, Gemini 3.7 Flash 65.3, Muse Spark 1.2 59.3. Aggregators and social posts quote these as one ranking.

THE CHECK. no shared harness produced those numbers. Google's launch chart, Meta's launch chart, and xAI's launch table each ran a same-named suite independently, and the same models land in both of the two charts that overlap. That overlap is the tell. GPT-5.6 Terra: 69.6 on Google's chart, 64.8 on Meta's. Muse Spark 1.2: 59.3 on Meta's chart, 54.9 on Google's. Roughly five points of daylight per model, and in each case the chart owner's rival scores lower on the owner's chart. As the one careful comparison in circulation puts it, cross-chart readings are 'suggestive but not proof, since it's two different labs running the same-named suite independently', or shorter: 'a shared benchmark name is not a shared benchmark'. The individual numbers are ordinary first-party launch stats. The leaderboard assembled from them is fiction with a spreadsheet aesthetic.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Terra is 69.6 on Google's chart and 64.8 on Meta's. Same model, same benchmark name, five points apart. Every cross-lab gap smaller than that is noise wearing a ranking."

Three coding-model launches landed within four days: xAI's Grok 4.6 on August 12, Google's Gemini 3.7 Flash on August 13, and the ongoing rollout of Meta's Muse Code agent on Muse Spark 1.2. Each launch shipped a chart, and each chart included a benchmark called DeepSWE v1.1, a long-horizon software

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

Z.ai says GLM-5.3 leads CyberGym with 84.5% and found 2,436 vulnerabilities. Independent verifications: zero. Vulnerabilities with public CVEs: 53 out of 2,436.

Every number is self-reported by the company that made the model. The weights are withheld for two weeks, making independent verification impossible until August 28.

01THE CLAIM
"GLM-5.3 leads CyberGym at 84.5%, beating Anthropic Mythos 5 and GPT-5.6 Sol, and found 2,436 confirmed vulnerabilities; weights delayed two weeks for safety." [SOURCE ↗]
BS4 SOURCES · LIVE 2026-08-28
Z.AI TRACK RECORD1 CLAIM · 100/100 BS RATE →
84.5% CyberGym score per OfficeChai
0independent verifications at launch
2436claimed confirmed vulnerabilities per fello AI
0.7point margin over Mythos 5, derived: 84.5 minus 83.8, both from Z.ai's own chart per OfficeChai receipt
02THE CHECK

THE CLAIM. Z.ai launched GLM-5.3 on August 14 claiming CyberGym 84.5%, beating Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%), plus 2,436 confirmed vulnerabilities across 269 open-source projects.

THE CHECK. searched for independent CyberGym replication of GLM-5.3 on August 16, 2026: zero results. Every score is Z.ai-reported, Z.ai-tested, on Z.ai's own harness configuration. Of the 2,436 "confirmed" vulnerabilities, only 53 have public CVE assignments. The remaining 2,383 are under embargo. Confirmed is doing work it has not earned.

THE TWIST. the weights are withheld for two weeks "for safety testing." That doubles as a perfect scarcity launch window. No one can replicate the benchmark until after the press cycle ends.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"GLM-5.3 claims CyberGym first place by 0.7 points on its own test. Independent verifications: zero. The weights are held back for two weeks. That is not a benchmark result. That is a press embargo dressed as safety."

Z.ai launched GLM-5.3 on August 14 with a familiar playbook: announce a model, cite a benchmark, claim the top spot, and withhold the weights that would let anyone check.

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 4 sources with quotes and screenshots, and our on-record call.

Musk says nothing will beat Grok 4.7 at engineering. So far it has beaten only its own ship date.

The superiority claim covers a model nobody outside xAI has run. The ship window slid from August 22 to early September inside three weeks, the 2.1T figure appears on no spec sheet, and the awesome-and-unique corpus is uncheckable by design.

01THE CLAIM
"Grok 4.7 'will exceed all current models', and the SpaceX training corpus is 'so awesome & unique' that no model will be better at real-world engineering, per Musk on August 12; the roughly 2.1 trillion parameter model ships in 3 to 4 weeks." [SOURCE ↗]
BS6 SOURCES · LIVE 2026-08-28
ELON MUSK TRACK RECORD4 CLAIMS · 70/100 BS RATE →
0PEOPLE OUTSIDE XAI WHO HAVE RUN GROK 4.7
Aug 22THE SHIP WINDOW THAT IS ALREADY GONE
5DAYS LATE GROK 4.6 SHIPPED AGAINST ITS OWN AUG 7 FRAMING
Musk says nothing will beat Grok 4.7 at engineering. So far it has beaten only its own ship date.
02THE CHECK

THE CLAIM. Grok 4.7 'will exceed all current models', with a SpaceX engineering corpus so strong Musk 'would be shocked if any model is better at real-world engineering than 4.7'. Reported size: about 2.1 trillion parameters. Timeline: '3 to 4 weeks'.

THE CHECK. every part of that is a promise about an object nobody outside xAI can touch. No benchmark exists, vendor-run or otherwise, because the model is unreleased. The 2.1T figure traces to Musk's remarks and appears on no xAI specification sheet. The training-corpus advantage is unfalsifiable and comes with its own asterisk in Musk's earlier post: the SpaceX data excludes everything ITAR-restricted, which is precisely the rocket-engineering core the sales pitch evokes. And the one checkable element, the calendar, has already failed once mid-claim: on July 25 the model was 'about four weeks out', pointing at August 22; on August 12 that became '3 to 4 weeks', pointing at September 2 to 9. The desk notes the base rate: Grok 4.6 was framed for August 7 and shipped August 12. This is a pre-release superiority claim scored against a track record of sliding windows, which is why it files as UNSUPPORTED rather than wrong: there is simply nothing to check yet, and the claimant knows it.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Nobody outside xAI has run Grok 4.7. The 2.1T is not on any spec sheet, the SpaceX corpus excludes the ITAR core it evokes, and the ship date already slipped once inside the claim itself."

On August 12, replying to coverage of Anthropic on X, Musk posted the full claim in three sentences: 'Grok 4.7 will exceed all current models. That said, Anthropic is a great company and will probably release improved models soon. However, the SpaceX training corpus is so awesome & unique that I wou

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

Bloomberg says Alibaba's AI eclipses Meta and Google with 3 billion downloads. Alibaba released 460 models. Meta released 3. Do the division.

Per model, Qwen gets 4.4 million downloads. Meta Llama gets far more per variant. Google Gemma gets 104.5 million. The headline measures catalog size, not adoption.

01THE CLAIM
"Alibaba's Qwen models have accumulated 3 billion global downloads in six months, eclipsing Meta and Google to become the world's No. 1 open-weight AI model family." [SOURCE ↗]
TRUE, BUT3 SOURCES · LIVE 2026-08-28
ALIBABA / BLOOMBERG TRACK RECORD1 CLAIM · 40/100 BS RATE →
3billion Qwen downloads per Yahoo Finance
460models open-sourced by Alibaba per AI Weekly
02THE CHECK

THE CLAIM. Alibaba's Qwen accumulated 3 billion downloads in six months, eclipsing Meta and Google to become the world's top open-weight AI family. Bloomberg, Yahoo Finance, and Fortune ran it.

THE CHECK. the Hugging Face report covers January through August (seven months, not six). Qwen open-sourced 460+ models. Meta has three main Llama variants. Google has four Gemma sizes. Per model: Qwen 4.4 million downloads, Meta far more per variant, Google 104.5 million. Qwen's per-model adoption trails Meta by 17x and Google by 24x.

THE TWIST. downloads count every variant, quant, and finetune separately. Qwen's 300,000 community derivatives each generate their own download count feeding back into the parent total. The metric measures ecosystem breadth, not model quality.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Qwen 3 billion downloads sounds dominant until you divide by 460 models. Alibaba won the catalog race, not the adoption race."

Bloomberg reported on August 15 that Alibaba's Qwen models hit 3 billion global downloads in the past six months, citing a Hugging Face State of Open Models report published August 14. The headline ran everywhere: Yahoo Finance, Fortune, Free Press Journal, AI Weekly.

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 3 sources with quotes and screenshots, and our on-record call.

Alibaba just passed Meta in downloads. Meta passed a billion before this scoreboard started counting.

The 3 billion is real, on one platform, in one seven-month window, summed across 460 plus models. The report says all of that in plain sentences. The headlines kept the number and dropped the scope.

01THE CLAIM
"Alibaba's Qwen family has passed 3 billion global downloads, making it the world's leading open-weight model family, ahead of Google at 418 million and Meta at 227 million, per a Hugging Face report picked up by Bloomberg on August 15." [SOURCE ↗]
TRUE, BUT7 SOURCES · LIVE 2026-08-28
PRESS COVERAGE AND ALIBABA AMPLIFICATION OF HUGGING FACE'S STATE OF OPEN MODELS REPORT TRACK RECORD1 CLAIM · 40/100 BS RATE →
3BQWEN DOWNLOADS: SEVEN MONTHS, ONE HUB, 460+ MODELS SUMMED
227MTHE META FIGURE ON THE SAME SCOREBOARD
1.2BMETA'S OWN LIFETIME COUNT, 16 MONTHS EARLIER
Alibaba just passed Meta in downloads. Meta passed a billion before this scoreboard started counting.
02THE CHECK

THE CLAIM. Qwen has amassed more than 3 billion downloads, versus 418 million for Google and 227 million for Meta, so Alibaba now leads open AI, per coverage of Hugging Face's August 14 State of Open Models report.

THE CHECK. the report is real and unusually careful, and every careful sentence got shed on the way to the headline. The count is 'activity observed on the Hugging Face Hub during the first seven months of 2026': one platform, one window, and by design 'nothing is credited for merely having existed longer'. It excludes, in the report's own words, 'API usage, private deployments, or models distributed through other channels'. The 227 million next to Meta's name is that window on that hub, not Meta's tally: Meta's own LlamaCon figure was 1.2 billion cumulative downloads back in April 2025. And the 3 billion sums 460 plus Qwen models against rivals who ship a handful, on counters that meter pulls (CI jobs, mirrors, re-quantizations), not people. Qwen's momentum is genuine and the report documents it well. The crown in the headline is a window-scoped, hub-scoped, SKU-summed crown.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The 3 billion is one platform, seven months, 460 models summed. Meta's own lifetime count was 1.2 billion before this window opened. The momentum is real; the crown is scoped."

Hugging Face published its State of Open Models report on August 14. Bloomberg ran the scoreboard a day later: Alibaba's Qwen family 'accumulated more than 3 billion global downloads in the past six months', against 418 million for Google and 227 million for Meta, and the coverage crowned Qwen the w

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 7 sources with quotes and screenshots, and our on-record call.

Headlines say Anthropic signed a $9.1 billion deal. Riot's SEC filing names no customer. The $9.1 billion runs 20 years to 2048.

A 20-year lease became a headline number. An unnamed customer became Anthropic via Bloomberg sourcing. The stock surged on an attribution neither company confirmed.

01THE CLAIM
"Anthropic signed a $9.1 billion, 191MW compute deal with Riot Platforms." [SOURCE ↗]
TRUE, BUT3 SOURCES · LIVE 2026-08-28
HEADLINES TRACK RECORD1 CLAIM · 40/100 BS RATE →
9.1billion USD total deal value per Bloomberg/CNBC
20years lease duration per Data Center Dynamics
191MW capacity per Data Center Dynamics
02THE CHECK

THE CLAIM. Anthropic signed a $9.1 billion compute deal with Riot Platforms for 191MW of capacity.

THE CHECK. Riot's 8-K filing with the SEC on August 10 describes the customer as a "leading frontier AI company." It does not name Anthropic. Bloomberg attached the name citing "people familiar with the matter." Neither company has officially confirmed the counterparty. The $9.1 billion is a 20-year lease running to June 2048, which is $455 million per year. First capacity: 96MW by December 2027. Full 191MW by June 2028.

THE TWIST. Riot stock surged 25% on the headline number. The headline number is a cumulative 20-year figure for a customer the company will not name in its own SEC filing.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The $9.1B Anthropic deal runs 20 years to 2048. Riot's own filing does not name Anthropic. CNBC's headline says signed. The filing says unnamed."

Riot Platforms filed an 8-K with the SEC on August 10, 2026. The filing describes a 20-year colocation lease agreement for 191MW of capacity at its Corsicana, Texas facility. The customer is described as a "leading frontier AI company."

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 3 sources with quotes and screenshots, and our on-record call.

The company that grades AI for OpenAI, Anthropic, Google, Meta, and xAI just raised $40M at a $400M valuation. It disclosed a customer relationship with the labs it evaluates.

The benchmark referee has a customer relationship with the labs it grades and is using a May number while August models score 75%. Independence is the product and the conflict.

01THE CLAIM
"Vals AI's independent benchmarks show frontier models correctly complete fewer than 52% of real financial-analyst tasks; a16z valued the evaluator at $400M." [SOURCE ↗]
TRUE, BUT4 SOURCES · LIVE 2026-08-28
VALS AI TRACK RECORD1 CLAIM · 40/100 BS RATE →
40million USD Series A per Crypto Briefing
52% Finance Agent v2 accuracy for GPT-5.5, May 2026 run, per KuCoin receipt
02THE CHECK

THE CLAIM. Vals AI raised $40M led by a16z at a $400M valuation on the pitch that it is the independent evaluator of AI. Its benchmarks appear in model cards from OpenAI, Anthropic, Google, Meta, and xAI.

THE CHECK. Artificial Lawyer reported Vals AI disclosed a customer relationship with one or more of the participants it evaluates. The benchmarks that appear in model cards are produced by a company billing the labs whose products carry those scores.

THE TWIST. independence is the product. The customers are the subjects. The conflict is structural, not hidden. Vals disclosed it. But disclosure does not eliminate the conflict; it just puts it on the record.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The AI benchmark referee just raised $400M. Its headline stat is three months stale, and it disclosed customer relationships with the labs it grades. That is not independence. That is consulting."

Vals AI positions itself as the independent evaluator of AI. Its Series A blog opens with: "We started Vals AI to solve this problem as the independent evaluator of artificial intelligence." Its benchmarks appear in model cards from five of the six major frontier labs.

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 4 sources with quotes and screenshots, and our on-record call.

THAT IS THE RECORD FOR ISSUE #10. NEXT VERDICT DROPS 9PM AEST.