Google says Gemini 4 Argon beat OpenAI and Anthropic across benchmarks. The independent index has it tied with GPT-6 Astra and five points behind Claude Opus 5.5.
Google says Gemini 4 Argon scored significantly higher than GPT-6 Astra and Anthropic's Fable and Opus across a variety of benchmarks, with a new top score on DeepSWE v1.1 (77.9%).
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Open the evidence4 source pages +
The claim we checked
Google's Gemini 4 Argon scored significantly higher than OpenAI's GPT-6 Astra and Anthropic's Fable and Opus models across a variety of AI benchmarks.
These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.
Google claims that Argon scored significantly higher than OpenAI’s GPT-6 Astra and Anthropic’s Fable and Opus models across a variety of AI benchmarks.
It sets a new state of the art on DeepSWE v1.1 (77.9%), which measures a model’s performance in real-world long-horizon software engineering tasks.
Argon ties for first place with a top score of 68%
to show that Argon is currently the leading model on the company’s AI model index.
scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53)
Anthropic's models still lead. Claude Opus 5.5 sits at 58 points and Claude Sonnet 5.5 at 56.
Google's own benchmark results paint a rosier picture. Argon leads in most of those benchmarks, sometimes by wide margins.
averaging 62k output tokens per task, compared with 27k for GPT-6 Astra (max)
After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
costs $1.99 per Intelligence Index task, 60% of GPT-6 Astra (max, $3.26)
being rolled out to a select group of the company’s cyber partners through its Fairwind Program
The price advantage comes from lower token rates, not from efficiency.