SUBSCRIBE

OpenAI's president says we're in the AGI era. The benchmark's own inventor scored the same model 37 points lower.

Astra's 99.9% only exists inside a harness OpenAI built for itself.

01THE CLAIM
"OpenAI launched GPT-6 Astra billing it as its most capable and aligned model yet, with president Greg Brockman saying 'it's not unreasonable to feel that we are now in the AGI era,' citing a 99.9% score on the ARC-AGI-3 benchmark." [SOURCE ↗]

THE MOVE: SELF-MARKED, graded by the party that benefits from the grade

TRUE, BUT5 SOURCES · LIVE 2026-09-05
OPENAI TRACK RECORD37 CLAIMS · 39/100 BS RATE →
99.9%GPT-6 Astra's ARC-AGI-3 score OpenAI leads with, scored via its own 'Provider Adapter' harness which carries hidden reasoning state between calls, cost ~$19K to run
62.7%Astra's ARC-AGI-3 score on ARC Prize's neutral 'Standard' harness, the apples-to-apples comparison point, cost ~$26K to run, a 37-point drop
61.2 / 60.9Artificial Analysis Intelligence Index: Astra vs its own six-month-old predecessor GPT-5.6 Sol, effectively a tie
65.7Claude Fable 5.1's Artificial Analysis Intelligence Index score, ahead of Astra
57.2% vs 65.0%Humanity's Last Exam: Astra vs Claude Fable 5.1, Astra loses outright
2.5xAstra's API price per token vs its own predecessor ($10/$50 vs $4/$20 per million tokens)
02THE CHECK

THE CLAIM. OpenAI president Greg Brockman said it's not unreasonable to feel that we are now in the AGI era, pointing to Astra's 99.9% on ARC-AGI-3.

THE CHECK. ARC Prize, the outfit that built the benchmark, ran the same model on its neutral harness and got 62.7%, a 37-point drop, and said outright it is not claiming AGI.

THE TWIST. on the Artificial Analysis Intelligence Index, a broader aggregate less prone to a single-harness trick, Astra ties its own six-month-old predecessor and loses to Claude Fable 5.1, while charging 2.5x more per token to do it.

03SAY THIS IN THE MEETING
"Ask which harness the number came from before you repeat it. 99.9% is OpenAI's homework, graded by OpenAI."
DEEP DIVE · THE FULL AUTOPSY

What actually happened

On September 3, 2026, OpenAI launched GPT-6 Astra, billing it as its most capable and aligned model yet. President Greg Brockman told reporters it's not unreasonable to feel that we are now in the AGI era, leaning on a 99.9% score on ARC-AGI-3.

The number is real. It is also achieved through OpenAI's own "Provider Adapter" harness, a setup that preserves hidden reasoning state between calls. Run the identical model through ARC Prize's neutral "Standard" harness, the version ARC Prize itself ran without OpenAI's scaffolding, and Astra scores 62.7%. That is a 37-point swing that tracks the change in scaffolding, not any claimed change in the model itself. ARC Prize's own writeup says outright: we are not claiming that it is AGI.

Why we rate this NEEDS CONTEXT, not FAILED

Astra is not a bad model. On Artificial Analysis's independent Intelligence Index, a broader aggregate less prone to a single-harness trick, it scores 61.2, effectively level with its own six-month-old predecessor GPT-5.6 Sol (60.9), and it trails Claude Fable 5.1 (65.7). On Humanity's Last Exam it loses outright to Fable 5.1, 57.2 to 65.0. Where Astra likely does better is agentic and computer-use work, precisely the areas OpenAI's own launch materials emphasized alongside the AGI framing. That part of the pitch is plausible on its own terms. The AGI framing riding on top of the benchmark number is not.

The steelman, and why it still fails

OpenAI's defenders will point out that every lab tunes its own serving stack, and that a harness which preserves reasoning state between calls is legitimate engineering, not a trick. Fair, as far as it goes. But the whole point of testing under a second, neutral harness is to check whether a score survives a change of scaffolding, and the record here shows this one does not: the same model lands 37 points lower the moment the benchmark's own inventor runs it outside OpenAI's harness. Whatever the Provider Adapter is to customers, on this test it is the difference between the headline and the result, and a number that needs the vendor's own scaffolding to exist is evidence of a well-optimized pipeline, not of general intelligence.

The mechanism

"AGI era" is a marketing phrase with no agreed technical threshold, which makes it cheap to claim and expensive to challenge. Pairing it with a saturating benchmark number gives the claim the appearance of a hard data point. The trick works because most readers will see "99.9% on ARC-AGI-3" and "AGI era" in the same sentence and assume the number backs the phrase. It backs a specific, non-reproducible test condition instead. Astra also costs 2.5x more per token than its own predecessor, which the launch framing conveniently leaves out of the AGI conversation.

What to do with this

  • Before repeating any benchmark headline, find out whether the score came from a neutral, provider-agnostic harness or a vendor-built one.
  • Evaluate Astra on the tasks it is genuinely strong at (agentic workflows, computer use), not on the AGI framing.
  • Watch the Artificial Analysis Intelligence Index and Humanity's Last Exam over the next model cycle. Those are the numbers Astra is actually behind on.
  • If a vendor's headline number requires their own custom harness to reproduce, treat the number as a demo, not a measurement.
04YOUR MOVE · WHAT IGNORING THIS COSTS

If your team is budgeting on Astra having crossed into AGI, you are budgeting on a number OpenAI cannot reproduce outside its own harness. Buy it, if at all, for the agentic work it's plausibly strong at. Don't buy the headline.

05OUR CALL · ON THE RECORD 2026-09-05

By December 1, 2026, no neutral, provider-agnostic harness will reproduce anything within 15 points of 99.9% on ARC-AGI-3 for Astra. Hold us to it.

Flips if ARC Prize or a second independent lab reproduces a score above 85% on the Standard harness without OpenAI's Provider Adapter, or if OpenAI publishes the adapter's methodology and it survives replication.

RECEIPTS (5) · CONFIDENCE HIGH · every URL below answered a live HTTP check before publish · sweep 2026-09-05

  • SUPPORTS THE CLAIM fortune.com · "It's not unreasonable to feel that we are now in the AGI era"
  • REFUTES IT arcprize.org · "GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness"
  • REFUTES IT arcprize.org · "we are not claiming that it is AGI"
  • ADDS CONTEXT requesty.ai · "Astra is good, but maybe we should calm down with the hype"
  • ADDS CONTEXT emergent.sh · "the Artificial Analysis Intelligence Index, is exactly where Astra looks incremental and lands behind Fable 5.1."

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.