Z.ai says GLM-5.3 leads CyberGym with 84.5% and found 2,436 vulnerabilities. Independent verifications: zero. Vulnerabilities with public CVEs: 53 out of 2,436.
Every number is self-reported by the company that made the model. The weights are withheld for two weeks, making independent verification impossible until August 28.
"GLM-5.3 leads CyberGym at 84.5%, beating Anthropic Mythos 5 and GPT-5.6 Sol, and found 2,436 confirmed vulnerabilities; weights delayed two weeks for safety." [SOURCE ↗]
THE CLAIM. Z.ai launched GLM-5.3 on August 14 claiming CyberGym 84.5%, beating Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%), plus 2,436 confirmed vulnerabilities across 269 open-source projects.
THE CHECK. searched for independent CyberGym replication of GLM-5.3 on August 16, 2026: zero results. Every score is Z.ai-reported, Z.ai-tested, on Z.ai's own harness configuration. Of the 2,436 "confirmed" vulnerabilities, only 53 have public CVE assignments. The remaining 2,383 are under embargo. Confirmed is doing work it has not earned.
THE TWIST. the weights are withheld for two weeks "for safety testing." That doubles as a perfect scarcity launch window. No one can replicate the benchmark until after the press cycle ends.
Z.ai launched GLM-5.3 on August 14 with a familiar playbook: announce a model, cite a benchmark, claim the top spot, and withhold the weights that would let anyone check.
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 4 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.