GET THE AUTOPSY ➔

Microsoft's new model goes toe-to-toe with the Claude that was champion in June. It is August.

MAI-Thinking-1's numbers are self-scored and aimed at last season: level with Opus 4.6 on one benchmark, preferred over mid-tier Sonnet 4.6 in an eval Microsoft commissioned.

01THE CLAIM
"Microsoft's MAI-Thinking-1, now rolling out in Foundry, is toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro, hits 97.0% on AIME 2025, and was preferred over Claude Sonnet 4.6 in blind human evaluations." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-25
MICROSOFT AI TRACK RECORD1 CLAIM · 40/100 BS RATE →
97.0%AIME 2025, SELF-REPORTED, UNCONFIRMED
4.6THE CLAUDE GENERATION IT COMPARES TO, ONE BEHIND
1,276TASKS IN THE PREFERENCE EVAL MICROSOFT COMMISSIONED
Microsoft's new model goes toe-to-toe with the Claude that was champion in June. It is August.
02THE CHECK

THE CLAIM. Microsoft's first in-house reasoning model matches Claude Opus 4.6 on SWE-Bench Pro, reaches 97.0% on AIME 2025, and beat Claude Sonnet 4.6 in blind human preference evaluations, all while running 35B active parameters at a mid-weight price.

THE CHECK. every comparison targets the previous Claude generation. Opus 4.6 was the frontier in spring; Opus 5 shipped July 24 and Fable 5 sits above it, and the model Microsoft beat on preference, Sonnet 4.6, is the mid-tier of that older line. The scores are self-reported from a vendor preprint, the human eval was commissioned by Microsoft from its rating partner Surge, an independent aggregator has not confirmed the flagship AIME figure, and Artificial Analysis lists no page for the model at all. The claims were minted at Build in June; the Foundry rollout re-airs them unchanged, two Claude generations later.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"It matches the previous Claude generation on one self-scored benchmark. No independent leaderboard lists it yet."

Microsoft announced MAI-Thinking-1 at Build on June 2, confirming the model previously reported as Project Polaris: a sparse mixture-of-experts design with 35 billion active parameters out of roughly a trillion, a 256K context window, and a pointed provenance pitch, trained on clean, commercially li

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.