Microsoft's new model goes toe-to-toe with the Claude that was champion in June. It is August.
MAI-Thinking-1's numbers are self-scored and aimed at last season: level with Opus 4.6 on one benchmark, preferred over mid-tier Sonnet 4.6 in an eval Microsoft commissioned.
"Microsoft's MAI-Thinking-1, now rolling out in Foundry, is toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro, hits 97.0% on AIME 2025, and was preferred over Claude Sonnet 4.6 in blind human evaluations." [SOURCE ↗]

THE CLAIM. Microsoft's first in-house reasoning model matches Claude Opus 4.6 on SWE-Bench Pro, reaches 97.0% on AIME 2025, and beat Claude Sonnet 4.6 in blind human preference evaluations, all while running 35B active parameters at a mid-weight price.
THE CHECK. every comparison targets the previous Claude generation. Opus 4.6 was the frontier in spring; Opus 5 shipped July 24 and Fable 5 sits above it, and the model Microsoft beat on preference, Sonnet 4.6, is the mid-tier of that older line. The scores are self-reported from a vendor preprint, the human eval was commissioned by Microsoft from its rating partner Surge, an independent aggregator has not confirmed the flagship AIME figure, and Artificial Analysis lists no page for the model at all. The claims were minted at Build in June; the Foundry rollout re-airs them unchanged, two Claude generations later.
Microsoft announced MAI-Thinking-1 at Build on June 2, confirming the model previously reported as Project Polaris: a sparse mixture-of-experts design with 35 billion active parameters out of roughly a trillion, a 256K context window, and a pointed provenance pitch, trained on clean, commercially li
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.