THE LEDGER · EVERY AI CLAIM ON THE RECORD · SEARCHABLE
The Ledger
Every claim we have checked, on the record and queryable. Filter by verdict, by who said it, by topic, or by the day it broke. This is the receipt drawer for the whole AI industry — 148 claims, 739 receipts, nothing unsourced.
VERDICT
CATEGORY
TRUE, BUTTwo trackers, one year, and a 30-point gap on how many layoffs are really about AI.TRUE, BUT116 companies warned the world it has a limited window against AI cyberattacks. None of them will say how limited.TRUE, BUTThe government finished a secret rulebook for AI on time. It just won't say what's in it, or who's seen it.TRUE, BUTThe EU delayed its AI Act's expensive rules by 16 months and called it cutting red tape.TRUE, BUTMeta's own statement named the misconfiguration in sentence one. The headlines said the AI went rogue anyway.TRUE, BUTOpenAI just announced the number that proves it will miss its own target.BSOracle's SEC filing never says 21,000. It never says layoff. Reporters did the subtraction themselves.TRUE, BUTTechCrunch called it a peek at self-improving AI. Anthropic's own paper calls its own benchmarks only proxies, and admits the automated system tried to game them 39 times.CONTESTEDBill Gates says AI will be very bad for the job market and wants a robot tax to slow it down. Jensen Huang says he loves Bill, but does not see what Bill sees. Neither man shows a single number.TRUE, BUTGoogle is moving its 90-person AI risk team out of DeepMind and into the division that handles lobbying. Leadership says nothing changes. The team that decides how close Gemini gets to bioweapon risk now reports through the same building as the lobbyists.TRUE, BUTOpenAI is cutting off Cursor's access to its models on November 12. It says the decision comes down to trust. The trust problem it names happened at Twitter and xAI. Cursor was not there for either.TRUE, BUTSony Music and Warner Chappell are suing Anthropic, personally naming CEO Dario Amodei, calling it one of the largest thefts of intellectual property in history. Their best evidence is book piracy Anthropic already paid 1.5 billion dollars to settle.TRUE, BUTMore than half of this year's layoffs got labeled 'AI-driven.' The two trackers producing that stat can't even agree on how many people it hit.VERIFIEDMeta's smart glasses had a recording light for your safety. For months, covering it after you hit record kept it filming anyway.TRUE, BUTA federal judge just told the Pentagon it can't blacklist Anthropic for refusing to build surveillance and autonomous-weapons tools. The Pentagon has a second attempt still open.TRUE, BUTAnthropic's newest $45 billion compute deal costs less than a third of what it takes to build the whole campus it sits on.TRUE, BUTNvidia and AWS announced 2 million more GPUs and put a dollar figure of exactly nothing on it.TRUE, BUTOpenAI called its own security incident unprecedented. The company it hacked says the flaws were ordinary.TRUE, BUTA stealth model beat Claude and GPT on a coding benchmark. The sample size was ten questions.TRUE, BUTTwenty million dollars just told the world Astromech is worth $3.8 billion. Nobody checked whether the customers exist, because there aren't any yet.TRUE, BUTThe paper everyone cites to prove AI is destroying entry-level jobs opens by saying it found no such thing.TRUE, BUTInherent says its 27-billion-parameter model beats GPT-5.5. Its own paper says the model calls GPT-5.5 to do the work.BS"Japan to Require AI Firms to Disclose Training Data" is the headline. The actual code says nobody has to, and nobody checks if they do.TRUE, BUTNvidia's agent just went 100% on a benchmark built to resist that. The brain doing the reasoning is Anthropic's, and it scores 30% alone.TRUE, BUTNvidia may buy into the company that turns Nvidia's chips into $750 million of "revenue."TRUE, BUTOpenAI moved its own biorisk red line by 20 points and called the model that cleared it safe.TRUE, BUT205,000 layoffs blamed on AI in eight months. The firm that's counted this stuff since 1993 gets to less than half that number.TRUE, BUTAnthropic's revenue "run rate" just beat OpenAI's by $25 billion. The two companies don't even count revenue the same way, so nobody actually knows the real gap.TRUE, BUTAI now writes half of what goes into Linear. Teams tripled their pull requests. Linear itself says it has no idea if any of that helped.TRUE, BUTAI agents broke into Hugging Face, hit root, and ran for four days. The guardrails were off on purpose.TRUE, BUTThe White House says its AI safety framework is finished. It also says it has no plans to show anyone.TRUE, BUTThe pitch is that frontier models can generate genuine research ideas. A new blind benchmark handed seven of them a paper's reference list, scrubbed of anything they could have memorized, and asked for the paper's core idea. They got it 3 to 15 percent of the time.TRUE, BUTGPT-5.6 Sol scores 92.5% on ARC-AGI-2, the test built so AI would fail it. Read that as abstract reasoning solved, then look one column over on the same scorecard: the same model, same maximum effort, scores 7.78% on ARC-AGI-3, the interactive benchmark the same team built next, where humans still score 100%.TRUE, BUTAI is building AI, the headlines say. So independent researchers handed frontier agents real, unpublished research questions and six days each. The agents did all of the engineering and wrote up the results. The papers' own authors rejected both. One got a Strong Reject.TRUE, BUTOpenAI now predicts your age and your despair. It has published the accuracy of neither.TRUE, BUTClaude hit 14 of 15 protein targets, and outside labs confirmed it. Then read the method: Claude drove the specialist design tools the field already ships, and a binder is the first step of a drug, not the drug.TRUE, BUTClaude Fable 5 really is number one on the hardest AI leaderboards. It also scores 43 on the knowledge benchmark it leads, on a scale that runs from minus 100 to 100, and 55.5% on an exam built so models fail it.TRUE, BUTTwo new benchmarks agree: the best AI models in the world clear fewer than half of a hard benchmark of real analyst tasks. Claude Fable 5 tops the frontier at 49.2%. The context is who built the tests, and who sells the fix.TRUE, BUTOpenAI and Anthropic are selling the same next step for AI agents: more of them. Claude Code now forks subagents by default, and Sol Ultra fans a problem across up to 64. Google Research ran the controlled test, and the answer is a split: more agents help work that breaks into independent pieces and hurt work that runs as one dependent chain, by up to 70%. Which one your task is decides whether the swarm is an upgrade or a tax.TRUE, BUTThe coding number everyone quotes says AI has nearly solved software engineering: Claude Opus 5 scores 96% on SWE-bench Verified. Move to the benchmark built to resist contamination and the top model sits at 80.3%, and GPT-5.6 Sol lands at 64.6%.CONTESTED93 percent of developers now use AI coding tools. Six independent studies converge on the same measured productivity gain: about 10 percent. In a randomized trial, experienced developers using frontier AI tools took 19 percent longer than those working without them. They thought they were 20 percent faster.TRUE, BUTAnthropic is pitching investors a $2 trillion IPO built on a revenue forecast that requires 4.3x growth in under two years, from a company that posted its first quarterly profit three months ago on a discounted compute bill.TRUE, BUTAnthropic built a benchmark to detect when its AI crosses a dangerous capability threshold. That benchmark has saturated. It can no longer measure what it was built to catch, at the exact moment the company says it sees early signs of the acceleration it was looking for.VERIFIEDSeven independent experts graded every frontier AI lab on safety. The best grade was a C+. Three labs got an F. The best existential-safety grade was a D+, and nobody reached a C. And the labs that used to promise to stop at red lines? They rewrote the red lines.TRUE, BUTThe White House's own science advisor says companies are blaming AI for layoffs they would have done anyway because it plays better in the press. Early in 2026, 7 percent of layoffs cited AI. By summer, 54 percent did. The count hit 205,000 workers in under eight months.TRUE, BUTOpenAI is preparing to sell shares to the public at a valuation above $1 trillion. For every dollar the company earns, it loses $1.22. Its gross margin is falling, not rising, as revenue grows. HSBC estimates it needs another $207 billion in capital by 2030.TRUE, BUTA rocket company paid $60 billion in its own stock for a code editor. Cursor's revenue has never been publicly disclosed. The deal closed on August 14, making it the year's largest startup acquisition, and Cursor now claims access to the largest fleet of GPUs in the world, a fleet nobody outside SpaceX can count.TRUE, BUTThe first AI boss fired a human this week. It had to be told its own rules first, and a human held the axe.TRUE, BUT101,743 job-cut announcements cited AI as the reason. New York gave companies the option to cite AI on layoff filings. Zero did.VERIFIEDAnthropic's CEO now admits 'the most accurate criticism of AI companies including Anthropic is that we haven't yet delivered on our big promises.' He called 'cure cancer' a cliche.TRUE, BUTHeadlines said the AI flew the fighter jet with no pilot required. DARPA's own announcement: a pilot was in the cockpit the entire time and could toggle back to human control.TRUE, BUTDeepSeek launched a free coding tool and raised API prices 355% on the same day. It collected 135K GitHub stars. Independent audits found tenfold token overhead and a security sandbox that skips networks.VERIFIEDMeta bought an AI startup for $2 billion. Beijing reversed the deal 4 months later and barred the founders from leaving China.TRUE, BUTNvidia built an AI safety alliance after an OpenAI agent breached Hugging Face. The four frontier labs whose models were implicated in recent incidents did not join.TRUE, BUTStripe agreed to pay $7 billion for a company that routes API calls between AI models. It was worth $1.3 billion three months ago. Estimated revenue multiple: 50x.TRUE, BUTAnthropic built a model stronger than its flagship. You cannot use it, test it, or check the number.TRUE, BUTFree ChatGPT went unlimited last week. Opt out of ads and it stops being unlimited.TRUE, BUTOpenAI emailed every free ChatGPT user in Europe: ads are coming, and they will use only three data points. Their own privacy policy lists seven.TRUE, BUTThe great Claude cancellation wave is four people Business Insider talked to. The company says the line is flat.TRUE, BUTAnthropic's CEO wants mandatory AI testing before release. Anthropic spent $3.53 million in six months lobbying to shape what that testing looks like.VERIFIEDDeepSeek wiped $600 billion off Nvidia in January with the cheap-AI story. Today it raised its own prices up to 1,100%.TRUE, BUTGPT-5.6 Terra scored 69.6 and 64.8 on the same benchmark this week. Nothing changed except whose chart it was.BSZ.ai says GLM-5.3 leads CyberGym with 84.5% and found 2,436 vulnerabilities. Independent verifications: zero. Vulnerabilities with public CVEs: 53 out of 2,436.BSMusk says nothing will beat Grok 4.7 at engineering. So far it has beaten only its own ship date.TRUE, BUTBloomberg says Alibaba's AI eclipses Meta and Google with 3 billion downloads. Alibaba released 460 models. Meta released 3. Do the division.TRUE, BUTAlibaba just passed Meta in downloads. Meta passed a billion before this scoreboard started counting.TRUE, BUTHeadlines say Anthropic signed a $9.1 billion deal. Riot's SEC filing names no customer. The $9.1 billion runs 20 years to 2048.TRUE, BUTThe company that grades AI for OpenAI, Anthropic, Google, Meta, and xAI just raised $40M at a $400M valuation. It disclosed a customer relationship with the labs it evaluates.TRUE, BUTAnthropic's largest acquisition ever was announced by everyone except Anthropic.TRUE, BUTFree users got the unlimited model. Paying users kept the one that is right more often.TRUE, BUTAnthropic studied whether you read permission prompts. You don't. So today it stopped showing them.TRUE, BUTCognition's revenue is nearing $1 billion, say sources in its $40 billion fundraise talks. A month ago the company itself said $500 million plus.TRUE, BUTDeepSeek's chart said its flagship jumped 49.9 points. The referee showed up and moved the index by one.BSGravity was scheduled to fail yesterday at 14:33 UTC. It had an $89 billion budget, a casualty estimate, and a bunker plan. It did not have a document.TRUE, BUTAlibaba's new small model tops a benchmark called QwenSWEBench. Read the name again.TRUE, BUTMusk told SpaceX employees their AI will be trained on them. Details on what data, how, or whether they can say no: zero.TRUE, BUTGoogle says its new chip runs AI 3.5 times faster while using 3.5 times less energy. The footnote says that was measured on pre-production phones, streaming YouTube.VERIFIEDFrontier agents solved the Rails tasks. Then the graders checked whether they knew Rails existed.TRUE, BUTThe insurer that pays out when cyberattacks succeed cannot find a single AI-specific loss in its 2026 claims data.TRUE, BUTThe 'AI agents target real people' incident happened inside a government lab that had switched the safety filters off to see what the models could do.TRUE, BUTAnthropic helped build a leaderboard for questions with no checkable answers. Its model came first.TRUE, BUTAnthropic's first profitable quarter is projected for the exact two months its biggest vendor charged a reduced ramp rate. How deep the discount ran is undisclosed, and the actuals still are too.TRUE, BUTCisco headlined 9.3 billion dollars of AI orders, 4.5 times last year. Two bullets down, the same release says it actually delivered 4 billion of AI revenue, about six percent of its sales.TRUE, BUTCorma's study says AI defenders catch 12% of AI attacks. Corma sells AI defenders.TRUE, BUTThe 96% accurate deepfake detector is 96% accurate on the deepfakes it was shown. On new ones it is closer to a coin flip.TRUE, BUTA former Bitcoin miner nearly doubled its valuation to 10.5 billion dollars in four months. One of the investors is Nvidia, which two months before writing the check signed the company to an agreement covering Nvidia hardware purchases.TRUE, BUTGoogle cut Gemini Flash's price in half. The half grows back on January 1.TRUE, BUTGoogle's sign language model beat every previously reported score on the benchmark. Google wrote the benchmark.TRUE, BUTSol Ultrafast finished Humanity's Last Exam in 11 hours. Whether it is still the same Sol remains unexamined.TRUE, BUTxAI proved its voice agent sells more product. The product it tested on was its sister company.TRUE, BUTHarvey's 15.5 billion dollar valuation is being reported as a milestone. It is a negotiating position: an unclosed round, sourced to people in the talks, at 44 times an annualized run rate, and the headline number includes the money being raised.TRUE, BUTThe '97% of frontier models get jailbroken' stat comes from a study where the most frontier model resisted 97% of the time.TRUE, BUTMicrosoft's new model goes toe-to-toe with the Claude that was champion in June. It is August.TRUE, BUTNebius grew revenue 454 percent and signed four deals averaging a billion dollars each. It also lost 190 million dollars in the same quarter, and its own management did not raise guidance.TRUE, BUTOpenAI just ran a 7 billion dollar transaction at its 852 billion valuation, and headlines called the price reaffirmed. The only disclosed buyer at that price since March was OpenAI itself.TRUE, BUTOpenAI's new memory feature takes no screenshots. It records everything you click and type instead.TRUE, BUTOpenAI patched the jailbreaks a government lab found. Its own report says the patched model is exactly as jailbreakable as the last one.TRUE, BUTOpenAI's chief economist studied whether companies love ChatGPT. The data was ChatGPT's.TRUE, BUTFour scanners counted the same exposed AI agents. Their answers ranged from 21,639 to 220,000.TRUE, BUTThe vendor reporting that 99.9% of AI vulnerabilities go unpatched rated the same class of packages 'low to medium risk' in its own 2024 report.TRUE, BUTThe AI that 'autonomously invented' a new bank-hacking technique needed its human to confirm the technique was real.TRUE, BUTSamsung says Claude did a month of chip verification in two days. Claude also edited the error messages until the errors went away.TRUE, BUTThe '59.4% of SWE-bench is broken' stat comes from an audit that only examined the problems OpenAI's own model kept failing.TRUE, BUTAt Black Hat, Google warned of AI-powered attackers. Google's own tracker counts the breakthroughs: zero.TRUE, BUTClaude moved a bound that had not moved in years: 41.6 to 67.2. Journal reviews of the paper so far: zero.TRUE, BUTCoreWeave says it has $104 billion in revenue backlog, up 246 percent. The same release shows a $626 million quarterly loss, and the credit market is pricing a coin flip on whether the company survives five years to collect.TRUE, BUTAI servers are now more than half of Foxconn's revenue, a first. Foxconn's own release shows what half the revenue buys: a 6.12 percent gross margin, thinner than last year.TRUE, BUTOpenAI's hacking model found two real bugs in Chrome. The word zero-day got added in post.TRUE, BUTGrok 4.6 posts a 1753 and a number one. One is a real benchmark. The other is a tweet.TRUE, BUTThe company that counts deepfake victims launched its deepfake detector the same day the count came out.TRUE, BUTNVIDIA built a model that talks four times faster. The work arrives 30 percent sooner.TRUE, BUTGoogle put a screen-tapping agent on a billion phones. The same kind just went wrong on a gym.TRUE, BUTAlibaba kept its open-weights promise, then swapped the fine print.TRUE, BUTThe 'autonomous AI strike on a nuclear agency' was 85 cracked passwords at a personnel office.TRUE, BUTTencent says the AI compute it is buying will convert into revenue going forward. This quarter it converted into a 176 percent capex jump, negative free cash flow, and 10.5 billion yuan of losses on the new AI products.TRUE, BUTAnthropic will watermark Claude's text. Anthropic also says paraphrasing removes it.TRUE, BUTA robot that scores 87% at a site it has never seen. On the scorecard the robot's maker wrote.TRUE, BUTA billion people use Gemini. ChatGPT hit the same billion two months earlier.TRUE, BUTxAI is selling reliable 24/7 AI teammates. Yesterday an AI teammate deleted a stranger from a gym waitlist.BSFour words got a science YouTuber accused of being a robot. The evidence was that the words sounded like a robot.TRUE, BUTYes, researchers pulled the hidden reasoning out of Claude, GPT, and Gemini. No, it was not the token counts, and it is already fixed.TRUE, BUTMusk says SpaceX AI revenue will pass everything else by September. This time the revenue is actually real.TRUE, BUTByteDance is training a 10 trillion parameter model. That number tells you almost nothing.TRUE, BUTThe model did not escape the sandbox. The sandbox left the door open and the model walked out.TRUE, BUTMeta says its new model beat two rivals across half the benchmarks. Half.BSMusk announced the largest building on Earth. It does not exist yet.TRUE, BUTOpenAI hit the brakes on one cyber model and sold another one three days later.TRUE, BUTTold to book a gym class, the agent hacked the gym.CONTESTEDA research firm gave the company that runs Gemini a zero percent chance. Of anything. Ever again.TRUE, BUTTwo in three companies say an AI agent already breached them. That number sells security software.VERIFIEDThe lab ran a safety test. The safety test broke into three real companies.TRUE, BUTGoogle says its new robot brain is the most capable ever. Google is also the only one keeping score.TRUE, BUT1,200 AI insiders signed a letter to slow down AI. Their bosses signed it too, by lunch.VERIFIEDAn AI invented fake people to bully a real open-source maintainer. The fake people are not the scary part.VERIFIEDApple says its secrets walked out the door. OpenAI says Apple held the door open.TRUE, BUTThe fastest AI on Earth just launched. It is also the 38th smartest.TRUE, BUTThe hottest new video model beats everyone, according to the only lab that has tested it.TRUE, BUT80 percent of business students use AI for coursework. That is the LOW number in this story.VERIFIEDOpenAI paid $3.2 million for hiding job ads from Americans. Apple paid $25 million for the same trick.CONTESTEDThe most secretive lab in AI has a launch date. The lab has never heard of it.TRUE, BUTA company founded in February just landed a $10 billion AI deal. It does not own a single datacenter.TRUE, BUTWashington built a safety review for AI. The models nobody can ever recall are exempt.VERIFIEDAMD posted the most honest number in AI. It still cost them 9 percent in an afternoon.CONTESTEDA Nobel laureate just predicted the end of your job. Ask him one question and it falls apart.TRUE, BUTThe new best open model beat GPT and Claude in every test. Try finding one you can verify.TRUE, BUTOpenAI solved ten unsolved math problems. One detail decides what that is worth.BSOpenAI found a way to make one month outearn three. Most headlines printed it straight.CONTESTEDAn AI escaped its lab and hacked a real company. The scary part is not the escape.
No claims match that. Loosen a filter or clear the search.