Start free trial

The trick: Cherry-Picked Slice

Google says Gemini 4 Argon beat OpenAI and Anthropic across benchmarks.

The independent index has it tied with GPT-6 Astra and five points behind Claude Opus 5.5.

Issue 311 October 202612 receipts3 min

Google says Gemini 4 Argon scored significantly higher than GPT-6 Astra and Anthropic's Fable and Opus across a variety of benchmarks, with a new top score on DeepSWE v1.1 (77.9%).

Before you read on. Your call?

Google picked the tests; some are run by third parties such as Vals. On Artificial Analysis's independent index, Argon scores 53, matching GPT-6 Astra, and The Decoder reports Claude Opus 5.5 at 58.

The twist

Argon is cheaper per task at launch prices, but it uses 62k output tokens per task against Astra's 27k, and Google's own post says rates double after the introductory period.

53Gemini 4 Argon
58Claude Opus 5.5 on the same index
62k vs 27kaverage output tokens per task

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Review my 30-day free trial →

30 days free. Payment card required. One introductory trial per customer.

Your trial ends on 30 days after you start. Unless you cancel before then in Account → Manage membership, we charge A$89 for the first year. It renews automatically at A$89/yr until cancelled.

You can ask for a full refund within 14 days after any annual payment, renewals included, with no reason needed, by emailing [email protected]. This voluntary refund does not limit your rights under the Australian Consumer Law.

By starting your trial, you agree to the Terms.

BS Killer is published by Inferno Tech Pty Ltd, ABN 27 647 413 474.

Already a member? Sign in

The trick has a name

We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“Google's own table has Argon winning. The independent index has it tied with GPT-6 Astra at 53 and behind Opus 5.5 at 58.”

Receipts

  1. Supports techcrunch.com: Google claims that Argon scored significantly higher than OpenAI’s GPT-6 Astra and Anthropic’s Fable and Opus models across a variety of AI benchmarks.
  2. Supports blog.google: It sets a new state of the art on DeepSWE v1.1 (77.9%), which measures a model’s performance in real-world long-horizon software engineering tasks.
  3. Supports blog.google: Argon ties for first place with a top score of 68%
  4. Context techcrunch.com: to show that Argon is currently the leading model on the company’s AI model index.
  5. Refutes artificialanalysis.ai: scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53)
  6. Refutes the-decoder.com: Anthropic's models still lead. Claude Opus 5.5 sits at 58 points and Claude Sonnet 5.5 at 56.
  7. Context the-decoder.com: Google's own benchmark results paint a rosier picture. Argon leads in most of those benchmarks, sometimes by wide margins.
  8. Context artificialanalysis.ai: averaging 62k output tokens per task, compared with 27k for GPT-6 Astra (max)
  9. Context blog.google: After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
  10. Context artificialanalysis.ai: costs $1.99 per Intelligence Index task, 60% of GPT-6 Astra (max, $3.26)
  11. Context techcrunch.com: being rolled out to a select group of the company’s cyber partners through its Fairwind Program
  12. Context the-decoder.com: The price advantage comes from lower token rates, not from efficiency.

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

Get tomorrow's check in your inbox.

One AI claim a night, checked against independent sources, receipts attached. Free.

Free email updates. Unsubscribe any time.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.