GET THE AUTOPSY ➔

Alibaba's new small model tops a benchmark called QwenSWEBench. Read the name again.

Qwen3.8-27B ships real open weights and plausible gains. Every launch score is self-graded, several tests are in-house, the rival's hardest rows are blank, and independent reproductions stand at zero.

01THE CLAIM
"Alibaba's Qwen3.8-27B, released August 14 with Apache 2.0 open weights, delivers frontier-tier agentic coding and computer use for its size: 61.7 on SWE-bench Pro, 73.0 on Terminal-Bench 2.1, 84.3 on OSWorld-Verified, beating Meta's Muse Glimmer 30B across the published rows." [SOURCE ↗]
TRUE, BUT9 SOURCES · LIVE 2026-08-25
ALIBABA QWEN TEAM TRACK RECORD2 CLAIMS · 40/100 BS RATE →
79.0QWENSWEBENCH, THE BENCHMARK QWEN NAMED AFTER ITSELF
0INDEPENDENT REPRODUCTIONS AT LAUNCH
80GBWHAT THE OFFICIAL BF16 BUILD ACTUALLY WANTS
Alibaba's new small model tops a benchmark called QwenSWEBench. Read the name again.
02THE CHECK

THE CLAIM. Alibaba's Qwen3.8-27B, open weights under Apache 2.0, posts frontier-tier agentic scores for a 27B model (61.7 SWE-bench Pro, 73.0 Terminal-Bench 2.1, 84.3 OSWorld-Verified) and beats Meta's Muse Glimmer 30B on the published comparison rows.

THE CHECK. every number on the card was measured by Qwen. The widest margin, 79.0, lands on QwenSWEBench, a benchmark with the vendor's name in the title, and Muse Glimmer's results are missing entirely from several of the harder benchmarks Alibaba ran. Launch-day reviewers found no independent reproduction of any Qwen3.8-27B score. The one-gaming-GPU framing also needs fine print: the official BF16 build wants about 80GB with the KV cache at native context; the 24GB story is a third-party quant at moderate context. The weights are genuinely downloadable, so this one is checkable. It just has not been checked.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"The weights are open and that is real. The scores are Qwen grading Qwen, the widest win is on Qwen's own benchmark, and nobody has reproduced a single number yet."

Alibaba's Qwen team released Qwen3.8-27B on August 14: a dense 27-billion-parameter multimodal model, open weights under Apache 2.0, 262,144 tokens of native context. The model card calls the 3.8 generation the most capable in the Qwen open-model family to date and posts the launch scores: 61.7 on S

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 9 sources with quotes and screenshots, and our on-record call.

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.