Alibaba's new small model tops a benchmark called QwenSWEBench. Read the name again.
Qwen3.8-27B ships real open weights and plausible gains. Every launch score is self-graded, several tests are in-house, the rival's hardest rows are blank, and independent reproductions stand at zero.
"Alibaba's Qwen3.8-27B, released August 14 with Apache 2.0 open weights, delivers frontier-tier agentic coding and computer use for its size: 61.7 on SWE-bench Pro, 73.0 on Terminal-Bench 2.1, 84.3 on OSWorld-Verified, beating Meta's Muse Glimmer 30B across the published rows." [SOURCE ↗]

THE CLAIM. Alibaba's Qwen3.8-27B, open weights under Apache 2.0, posts frontier-tier agentic scores for a 27B model (61.7 SWE-bench Pro, 73.0 Terminal-Bench 2.1, 84.3 OSWorld-Verified) and beats Meta's Muse Glimmer 30B on the published comparison rows.
THE CHECK. every number on the card was measured by Qwen. The widest margin, 79.0, lands on QwenSWEBench, a benchmark with the vendor's name in the title, and Muse Glimmer's results are missing entirely from several of the harder benchmarks Alibaba ran. Launch-day reviewers found no independent reproduction of any Qwen3.8-27B score. The one-gaming-GPU framing also needs fine print: the official BF16 build wants about 80GB with the KV cache at native context; the 24GB story is a third-party quant at moderate context. The weights are genuinely downloadable, so this one is checkable. It just has not been checked.
Alibaba's Qwen team released Qwen3.8-27B on August 14: a dense 27-billion-parameter multimodal model, open weights under Apache 2.0, 262,144 tokens of native context. The model card calls the 3.8 generation the most capable in the Qwen open-model family to date and posts the launch scores: 61.7 on S
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 9 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.