A stealth model beat Claude and GPT on a coding benchmark. The sample size was ten questions.
Its own tester ran the other 103 the next day and lost 17 points before an independent run lost more.
"A viral claim on X said the mystery stealth model 'Ox Alpha' beats Claude Fable 5 and GPT-5.6 Sol on the DeepSWE coding benchmark, based on an 8/10-task sample (effectively >80%)." [SOURCE ↗]

THE CLAIM. A tester's viral post said the free stealth model Ox Alpha scored 80% on the DeepSWE coding benchmark, beating GPT-5.6 Sol's 52% and Claude Fable 5's 65%, and Coin Bureau broadcast it to a much larger audience as a mystery model beating the leaders. THE CHECK: the tester had run 10 of the benchmark's 113 tasks. He ran the rest the next day and landed at 63%. An independent lab ran all 113 and got 58.4%, calling the 80% figure completely incorrect. THE TWIST: the model's maker, Zhipu AI, has since confirmed its identity as GLM-5.3-Flash, and on a more rigorous multi-trial benchmark the base GLM-5.3 model (not confirmed as the same Flash variant) loses to GPT-5.6 Sol on a single attempt but ties or leads it once retries are allowed.
What actually happened
On August 21, 2026, a tester posting as Ben Davis (@davis7) ran a free, unlabeled "stealth" model called Ox Alpha through a small slice of the DeepSWE coding benchmark, ten tasks out of a possible 113. The result: eight of ten solved, an effective score above 80%, against 52% for GPT-5.6 Sol and 65% for Claude Fable 5 on the same tiny slice. He posted the numbers with a caption admitting he was confused by them.
That caveat did not survive contact with the internet. The account AGTP repeated the numbers without the sample-size warning, and Coin Bureau, an account with a much larger, less technical audience, turned it into "BREAKING: A mysterious new AI model called Ox Alpha is reportedly BEATING Claude Fable 5 and GPT-5.6 Sol at coding, and nobody knows who built it." That version is what actually went viral.
The correction arrived fast, from the same person. The next day, Davis posted his own full 113-task run: roughly 63%, not 80%. A separate, independent tester (StealthModelWatch) ran the complete benchmark too and landed even lower, 58.4%, explicitly stating the rumored 80% pass rate was completely incorrect. On August 26, Zhipu AI (Z.ai) ended the identity mystery, confirming authorship and naming the model GLM-5.3-Flash.
Why we rate this needs_context
Every piece of the mechanical story checks out: the original 10-task tweet, the self-correction, the independent full run, the identity reveal. Nothing here is fabricated. What breaks is the headline framing, "beats Claude and GPT," which only ever existed in a sample size too small to mean anything, and which the original poster killed himself within 24 hours.
The steelman, and why it still falls short
The fairest pushback: a more rigorous, multi-trial benchmark from Together AI found GPT-5.6 Sol ahead of the base GLM-5.3 model on a single attempt (72.7% to 69.0%), but GLM-5.3 ties Sol at a second attempt (81.1% to 81.0%) and leads by a fourth (87.6% to 85.8%). For agentic workflows that permit retries, that is a real, non-hyped point in the model family's favor, with one caveat: this benchmark tested plain GLM-5.3, not confirmed as the same Flash checkpoint Zhipu named as Ox Alpha. "Ties on retries" and "beats outright, nobody knows who built it" are different claims regardless, and only the second one got 70 million impressions.
The mechanism
This is the standard shape of a benchmark-hype cycle: a small, honest sample produces a startling number, a caveat gets stripped on the first retweet, and the correction, when it comes, reaches a fraction of the audience the hype did. Ben Davis behaved well here, he ran the correction himself within a day. The system around him did not carry that correction as far as the claim.
What to do with this
- Treat any benchmark score built from fewer than 30-40 tasks as a rumor, not a result, regardless of who posts it.
- When a claim says "nobody knows who built it," check whether the maker has simply not spoken yet, that gap tends to close within a week, as it did here.
- If you are choosing a model off a viral benchmark screenshot, wait for the full-suite number. It is usually lower, and sometimes by more than 20 points.
Ten questions is a coin flip with a leaderboard attached, and if that tweet made it into your team's model-selection deck this week, you copied a number its own author disowned the next day.
By Sept 27, 2026, expect any audited full-113-task DeepSWE score published for the confirmed model to land in the high 50s to mid 60s, not 80%. Hold us to it.
Flips if a different lab reruns the full 113-task set on the confirmed model and reproduces a score above 75%.
RECEIPTS (7) · CONFIDENCE HIGH
every URL below answered a live HTTP check before publish · sweep 2026-08-28
- ▲ x.com ⧉ · "gpt-5.6-sol: 52% fable: 65% whatever the hell this is: 80%"
- ▼ x.com ⧉ · "Actual DeepSWE run on the ox alpha mystery model is done. Ended at ~63% NOT the 80% my first subset test got"
- ▼ stealthmodelwatch.online ⧉ · "The rumored ~80% pass rate is completely incorrect"
- ● together.ai ⧉ · "Sol edges ahead: 72.7% pass@1 to GLM-5.3's 69.0%"
- ▲ x.com ⧉ · "BREAKING: A mysterious new AI model called "Ox Alpha" is reportedly BEATING Claude Fable 5 and GPT-5.6 Sol at coding, and nobody knows who built it"
- ● vktr.com ⧉ · "On August 26, Z.ai confirmed authorship of the model, vindicating the fingerprinting. The model's official name is GLM-5.3-Flash"
- ● together.ai ⧉ · "It ties Sol at pass@2 (81.1 vs. 81.0) and leads pass@4 (87.6% vs. 85.8%)"


