Subscribe

The trick: Cherry-Picked Slice

Google says Gemini 4 Argon beat OpenAI and Anthropic across benchmarks.

The independent index has it tied with GPT-6 Astra and five points behind Claude Opus 5.5.

Issue 311 October 202612 receipts3 min

Google says Gemini 4 Argon scored significantly higher than GPT-6 Astra and Anthropic's Fable and Opus across a variety of benchmarks, with a new top score on DeepSWE v1.1 (77.9%).

Before you read on. Your call?

Google picked the tests; some are run by third parties such as Vals. On Artificial Analysis's independent index, Argon scores 53, matching GPT-6 Astra, and The Decoder reports Claude Opus 5.5 at 58.

The twist

Argon is cheaper per task at launch prices, but it uses 62k output tokens per task against Astra's 27k, and Google's own post says rates double after the introductory period.

53Gemini 4 Argon
58Claude Opus 5.5 on the same index
62k vs 27kaverage output tokens per task

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

The trick has a name

We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“Google's own table has Argon winning. The independent index has it tied with GPT-6 Astra at 53 and behind Opus 5.5 at 58.”

Receipts

  1. Supports techcrunch.com: Google claims that Argon scored significantly higher than OpenAI’s GPT-6 Astra and Anthropic’s Fable and Opus models across a variety of AI benchmarks.
  2. Supports blog.google: It sets a new state of the art on DeepSWE v1.1 (77.9%), which measures a model’s performance in real-world long-horizon software engineering tasks.
  3. Supports blog.google: Argon ties for first place with a top score of 68%
  4. Context techcrunch.com: to show that Argon is currently the leading model on the company’s AI model index.
  5. Refutes artificialanalysis.ai: scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53)
  6. Refutes the-decoder.com: Anthropic's models still lead. Claude Opus 5.5 sits at 58 points and Claude Sonnet 5.5 at 56.
  7. Context the-decoder.com: Google's own benchmark results paint a rosier picture. Argon leads in most of those benchmarks, sometimes by wide margins.
  8. Context artificialanalysis.ai: averaging 62k output tokens per task, compared with 27k for GPT-6 Astra (max)
  9. Context blog.google: After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
  10. Context artificialanalysis.ai: costs $1.99 per Intelligence Index task, 60% of GPT-6 Astra (max, $3.26)
  11. Context techcrunch.com: being rolled out to a select group of the company’s cyber partners through its Fairwind Program
  12. Context the-decoder.com: The price advantage comes from lower token rates, not from efficiency.

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.