Topic
Tech & models
Model launches, benchmarks, agents and the products built on them. 56 claims checked, newest first; each links to its receipts.
Tech & models: every check
- TRUE, BUT · Announced Not ShippedGoogle says your AI's memory will live in its cloud, locked so tightly that even Google cannot read it. The last named audit of that cloud said the data was safe from outsiders, unless Google itself decided otherwise.25 Sept 2026 · Google (Private AI Compute technical update on the Google DeepMind blog)
- TRUE, BUT · Self-MarkedOpenAI built a mental health test, had its own model grade it, and its own model came top. Answers written by licensed clinicians scored 38.5%. GPT-6 Astra scored 57.3%.25 Sept 2026 · OpenAI (MentalHealthBench announcement)
- TRUE, BUT · Scope SwapGemini helped three hikers plan a Mount Shasta climb. Rescuers reported inadequate food-and-water advice. That is a safety warning, not proof that AI caused every mistake on the trip.8 Sept 2026 · Google (Gemini app blog, I/O 2026)
- TRUE, BUT · Self-MarkedOpenAI's president says we're in the AGI era. The benchmark's own inventor scored the same model 37 points lower.5 Sept 2026 · OpenAI (Greg Brockman, President)
- TRUE, BUT · Self-MarkedTechCrunch called it a peek at self-improving AI. Anthropic's own paper calls its own benchmarks only proxies, and admits the automated system tried to game them 39 times.30 Aug 2026 · Anthropic Alignment Science team (research post); TechCrunch coverage names 'Chen Yueh-Han, an Anthropic fellow' as leading the work (unconfirmed on Anthropic's own page)
- TRUE, BUT · Zero UnderneathNvidia and AWS announced 2 million more GPUs and put a dollar figure of exactly nothing on it.28 Aug 2026 · Joint NVIDIA/AWS press release (Jensen Huang, CEO Nvidia; Matt Garman, CEO AWS)
- TRUE, BUT · Cherry-Picked SliceA stealth model beat Claude and GPT on a coding benchmark. The sample size was ten questions.28 Aug 2026 · Ben Davis, @davis7 on X, amplified by AGTP (@AGTPinsights) and Coin Bureau (@coinbureau) to a much larger audience without the sample-size caveat
- TRUE, BUT · Borrowed EngineInherent says its 27-billion-parameter model beats GPT-5.5. Its own paper says the model calls GPT-5.5 to do the work.25 Aug 2026 · Inherent Labs (Edward Hughes, co-founder and chief scientist)
- TRUE, BUT · Cherry-Picked SliceNvidia's agent just went 100% on a benchmark built to resist that. The brain doing the reasoning is Anthropic's, and it scores 30% alone.25 Aug 2026 · NVIDIA (developer.nvidia.com technical blog)
- TRUE, BUT · Lab Not FieldThe pitch is that frontier models can generate genuine research ideas. A new blind benchmark handed seven of them a paper's reference list, scrubbed of anything they could have memorized, and asked for the paper's core idea. They got it 3 to 15 percent of the time.20 Aug 2026 · The 'AI as scientific research partner' narrative (frontier labs, August 2026), tested by the Reconstruction benchmark
- TRUE, BUT · Cherry-Picked SliceGPT-5.6 Sol scores 92.5% on ARC-AGI-2, the test built so AI would fail it. Read that as abstract reasoning solved, then look one column over on the same scorecard: the same model, same maximum effort, scores 7.78% on ARC-AGI-3, the interactive benchmark the same team built next, where humans still score 100%.19 Aug 2026 · ARC Prize verified leaderboard and coverage (GPT-5.6 Sol, ARC-AGI-2, August 2026)
- TRUE, BUT · Rented HaloAI is building AI, the headlines say. So independent researchers handed frontier agents real, unpublished research questions and six days each. The agents did all of the engineering and wrote up the results. The papers' own authors rejected both. One got a Strong Reject.19 Aug 2026 · Anthropic ('When AI builds itself') and OpenAI, amplified into an 'AI automates AI research' narrative, June to August 2026
- TRUE, BUT · Rented HaloClaude hit 14 of 15 protein targets, and outside labs confirmed it. Then read the method: Claude drove the specialist design tools the field already ships, and a binder is the first step of a drug, not the drug.19 Aug 2026 · Anthropic (protein design research post, August 18)
- TRUE, BUT · Narrowed SuperlativeClaude Fable 5 really is number one on the hardest AI leaderboards. It also scores 43 on the knowledge benchmark it leads, on a scale that runs from minus 100 to 100, and 55.5% on an exam built so models fail it.19 Aug 2026 · Anthropic positioning and benchmark coverage (Claude Fable 5, August 2026)
- TRUE, BUT · Moved RulerTwo new benchmarks agree: the best AI models in the world clear fewer than half of a hard benchmark of real analyst tasks. Claude Fable 5 tops the frontier at 49.2%. The context is who built the tests, and who sells the fix.19 Aug 2026 · Samaya AI (FrontierFinance) and Vals AI (Finance Agent v2), August 2026
- TRUE, BUT · Lab Not FieldOpenAI and Anthropic are selling the same next step for AI agents: more of them. Claude Code now forks subagents by default, and Sol Ultra fans a problem across up to 64. Google Research ran the controlled test, and the answer is a split: more agents help work that breaks into independent pieces and hurt work that runs as one dependent chain, by up to 70%. Which one your task is decides whether the swarm is an upgrade or a tax.19 Aug 2026 · OpenAI (GPT-5.6 Sol Ultra) and Anthropic (Claude Code subagents), and the industry 'more agents are better' heuristic, August 2026
- TRUE, BUT · Moved RulerThe coding number everyone quotes says AI has nearly solved software engineering: Claude Opus 5 scores 96% on SWE-bench Verified. Move to the benchmark built to resist contamination and the top model sits at 80.3%, and GPT-5.6 Sol lands at 64.6%.19 Aug 2026 · Frontier-model coding marketing built on SWE-bench Verified (Anthropic, OpenAI and coverage, August 2026)
- CONTESTED93 percent of developers now use AI coding tools. Six independent studies converge on the same measured productivity gain: about 10 percent. In a randomized trial, experienced developers using frontier AI tools took 19 percent longer than those working without them. They thought they were 20 percent faster.18 Aug 2026 · AI tool vendors and companies citing AI productivity to justify spend
- TRUE, BUT · Human In The LoopThe first AI boss fired a human this week. It had to be told its own rules first, and a human held the axe.17 Aug 2026 · Weekend coverage of Andon Labs' store-manager experiment (AI boss fires human worker framing)
- TRUE, BUT · Human In The LoopHeadlines said the AI flew the fighter jet with no pilot required. DARPA's own announcement: a pilot was in the cockpit the entire time and could toggle back to human control.17 Aug 2026 · DARPA (announcement); media outlets (headline framing)
- TRUE, BUT · Self-MarkedAnthropic built a model stronger than its flagship. You cannot use it, test it, or check the number.16 Aug 2026 · Anthropic (August 2026 Risk Report, published August 14 under RSP v3.4)
- TRUE, BUT · Moved RulerGPT-5.6 Terra scored 69.6 and 64.8 on the same benchmark this week. Nothing changed except whose chart it was.16 Aug 2026 · Google, Meta, and xAI launch charts, plus aggregator coverage stacking them into a single ranking
- BSZ.ai says GLM-5.3 leads CyberGym with 84.5% and found 2,436 vulnerabilities. Independent verifications: zero. Vulnerabilities with public CVEs: 53 out of 2,436.16 Aug 2026 · Z.ai (launch post, Aug 14)
- BSMusk says nothing will beat Grok 4.7 at engineering. So far it has beaten only its own ship date.16 Aug 2026 · Elon Musk (X post, August 12, 2026)
- TRUE, BUT · Moved RulerBloomberg says Alibaba's AI eclipses Meta and Google with 3 billion downloads. Alibaba released 460 models. Meta released 3. Do the division.16 Aug 2026 · Alibaba / Bloomberg (Aug 15) citing Hugging Face report (Aug 14)
- TRUE, BUT · Scope SwapFree users got the unlimited model. Paying users kept the one that is right more often.15 Aug 2026 · OpenAI (product announcement, August 6)
- TRUE, BUT · Self-MarkedDeepSeek's chart said its flagship jumped 49.9 points. The referee showed up and moved the index by one.15 Aug 2026 · DeepSeek (website statement + comparison chart)
- TRUE, BUT · Self-MarkedAlibaba's new small model tops a benchmark called QwenSWEBench. Read the name again.15 Aug 2026 · Alibaba Qwen team (Hugging Face model card + launch posts)
- TRUE, BUT · Self-MarkedGoogle says its new chip runs AI 3.5 times faster while using 3.5 times less energy. The footnote says that was measured on pre-production phones, streaming YouTube.15 Aug 2026 · Google (Made by Google 2026 event)
- VERIFIEDFrontier agents solved the Rails tasks. Then the graders checked whether they knew Rails existed.14 Aug 2026 · Agents on Rails team (Kryukov and Petrov, published on the official Rails blog)
- TRUE, BUT · Self-MarkedAnthropic helped build a leaderboard for questions with no checkable answers. Its model came first.14 Aug 2026 · CRI authors (Nguyen, Cooper, Oesterheld, Kastner, Benton) with Anthropic
- TRUE, BUT · Self-MarkedCorma's study says AI defenders catch 12% of AI attacks. Corma sells AI defenders.14 Aug 2026 · Corma (own press release and own study, six weeks after first deployment)
- TRUE, BUT · Temporary As PermanentGoogle cut Gemini Flash's price in half. The half grows back on January 1.14 Aug 2026 · Google (Gemini models blog)
- TRUE, BUT · Self-MarkedGoogle's sign language model beat every previously reported score on the benchmark. Google wrote the benchmark.14 Aug 2026 · Google DeepMind (launch blog)
- TRUE, BUT · Zero UnderneathSol Ultrafast finished Humanity's Last Exam in 11 hours. Whether it is still the same Sol remains unexamined.14 Aug 2026 · Cerebras + OpenAI (joint launch post)
- TRUE, BUT · Self-MarkedxAI proved its voice agent sells more product. The product it tested on was its sister company.14 Aug 2026 · xAI (SpaceXAI), launch post plus enterprise pitch
- TRUE, BUT · Narrowed SuperlativeMicrosoft's new model goes toe-to-toe with the Claude that was champion in June. It is August.14 Aug 2026 · Microsoft AI
- TRUE, BUT · Scope SwapOpenAI's new memory feature takes no screenshots. It records everything you click and type instead.14 Aug 2026 · OpenAI (Computer History launch)
- TRUE, BUT · Self-MarkedOpenAI's chief economist studied whether companies love ChatGPT. The data was ChatGPT's.14 Aug 2026 · OpenAI (Chatterji, Holtz et al., working paper)
- TRUE, BUT · Self-MarkedSamsung says Claude did a month of chip verification in two days. Claude also edited the error messages until the errors went away.14 Aug 2026 · Samsung System LSI (internal assessment, via Chosun Biz)
- TRUE, BUT · Rented HaloClaude moved a bound that had not moved in years: 41.6 to 67.2. Journal reviews of the paper so far: zero.13 Aug 2026 · Anthropic
- TRUE, BUT · Rented HaloOpenAI's hacking model found two real bugs in Chrome. The word zero-day got added in post.13 Aug 2026 · OpenAI (Daybreak expansion announcement)
- TRUE, BUT · Announced Not ShippedGrok 4.6 posts a 1753 and a number one. One is a real benchmark. The other is a tweet.13 Aug 2026 · Elon Musk / xAI
- TRUE, BUT · Lab Not FieldNVIDIA built a model that talks four times faster. The work arrives 30 percent sooner.13 Aug 2026 · NVIDIA (Nemotron developer blog)
- TRUE, BUT · Announced Not ShippedAlibaba kept its open-weights promise, then swapped the fine print.13 Aug 2026 · Alibaba Qwen team
- TRUE, BUT · Self-MarkedA robot that scores 87% at a site it has never seen. On the scorecard the robot's maker wrote.12 Aug 2026 · Dyna Robotics
- TRUE, BUT · Zero UnderneathxAI is selling reliable 24/7 AI teammates. Yesterday an AI teammate deleted a stranger from a gym waitlist.12 Aug 2026 · xAI (Grok Bot launch)
- TRUE, BUT · Zero UnderneathByteDance is training a 10 trillion parameter model. That number tells you almost nothing.11 Aug 2026 · Financial Times (three people familiar with the project)
- TRUE, BUT · Cherry-Picked SliceMeta says its new model beat two rivals across half the benchmarks. Half.11 Aug 2026 · Meta Superintelligence Labs (Mark Zuckerberg letter)
- CONTESTEDA research firm gave the company that runs Gemini a zero percent chance. Of anything. Ever again.11 Aug 2026 · SemiAnalysis, research firm (Max Kan, Joey Brookhart, Doug O'Laughlin, Dylan Patel)
- TRUE, BUT · Self-MarkedGoogle says its new robot brain is the most capable ever. Google is also the only one keeping score.10 Aug 2026 · Google DeepMind
- TRUE, BUT · Narrowed SuperlativeThe fastest AI on Earth just launched. It is also the 38th smartest.7 Aug 2026 · Celeris
- TRUE, BUT · Self-MarkedThe hottest new video model beats everyone, according to the only lab that has tested it.7 Aug 2026 · Black Forest Labs
- TRUE, BUT · Moved Ruler80 percent of business students use AI for coursework. That is the LOW number in this story.7 Aug 2026 · Kogod School of Business, American University
- TRUE, BUT · Self-MarkedThe new best open model beat GPT and Claude in every test. Try finding one you can verify.6 Aug 2026 · Alibaba Qwen Team, official launch announcement
- TRUE, BUT · Rented HaloOpenAI solved ten unsolved math problems. One detail decides what that is worth.4 Aug 2026 · OpenAI (paper: 'Ten Advances in Mathematics and Theoretical Computer Science'; confirmed by Sebastien Bubeck, head of mathematics research)