Learn · a trick with a name
Self-Marked
graded by the party that benefits from the grade.
A score sounds like a measurement. It is only a measurement if someone without a stake took it.
How to spot it
- The number comes from the company's own blog, eval or harness
- No independent lab has published the same test
- The method section is thinner than the headline
The one question to ask“Who ran the test, and do they win if it looks good?”
At work
Before you quote a benchmark in a deck, find the same number from someone who doesn't sell the model. If you can't, call it a vendor claim.
Caught in the wild
Every time we've caught it so far. Try calling each one before you open it.
- TRUE, BUTApple blamed "the rapid expansion of AI data centers" for the price hikes on your next Mac or iPad. Apple's own decision to cram more RAM into its devices for its own AI features is a second AI-driven cost its statement never mentions.
- TRUE, BUTOpenAI's president says we're in the AGI era. The benchmark's own inventor scored the same model 37 points lower.
- TRUE, BUTOpenAI just crossed a cybersecurity line no model has crossed before. Its own timeline shows it saw this coming a month early.
- TRUE, BUTTechCrunch called it a peek at self-improving AI. Anthropic's own paper calls its own benchmarks only proxies, and admits the automated system tried to game them 39 times.
- TRUE, BUTGoogle is moving its 90-person AI risk team out of DeepMind and into the division that handles lobbying. Leadership says nothing changes. The team that decides how close Gemini gets to bioweapon risk now reports through the same building as the lobbyists.
- TRUE, BUTOpenAI called its own security incident unprecedented. The company it hacked says the flaws were ordinary.
- TRUE, BUTAnthropic built a model stronger than its flagship. You cannot use it, test it, or check the number.
- TRUE, BUTAnthropic's CEO wants mandatory AI testing before release. Anthropic spent $3.53 million in six months lobbying to shape what that testing looks like.
- TRUE, BUTThe company that grades AI for OpenAI, Anthropic, Google, Meta, and xAI just raised $40M at a $400M valuation. It disclosed a customer relationship with the labs it evaluates.
- TRUE, BUTDeepSeek's chart said its flagship jumped 49.9 points. The referee showed up and moved the index by one.
- TRUE, BUTAlibaba's new small model tops a benchmark called QwenSWEBench. Read the name again.
- TRUE, BUTGoogle says its new chip runs AI 3.5 times faster while using 3.5 times less energy. The footnote says that was measured on pre-production phones, streaming YouTube.