OpenAI's president says we're in the AGI era. The benchmark's own inventor scored the same model 37 points lower.
READ THE FULL STORY PAGE →Astra's 99.9% only exists inside a harness OpenAI built for itself.
"OpenAI launched GPT-6 Astra billing it as its most capable and aligned model yet, with president Greg Brockman saying 'it's not unreasonable to feel that we are now in the AGI era,' citing a 99.9% score on the ARC-AGI-3 benchmark." [SOURCE ↗]
THE MOVE: SELF-MARKED, graded by the party that benefits from the grade
THE CLAIM. OpenAI president Greg Brockman said it's not unreasonable to feel that we are now in the AGI era, pointing to Astra's 99.9% on ARC-AGI-3.
THE CHECK. ARC Prize, the outfit that built the benchmark, ran the same model on its neutral harness and got 62.7%, a 37-point drop, and said outright it is not claiming AGI.
THE TWIST. on the Artificial Analysis Intelligence Index, a broader aggregate less prone to a single-harness trick, Astra ties its own six-month-old predecessor and loses to Claude Fable 5.1, while charging 2.5x more per token to do it.
What actually happened
On September 3, 2026, OpenAI launched GPT-6 Astra, billing it as its most capable and aligned model yet. President Greg Brockman told reporters it's not unreasonable to feel that we are now in the AGI era, leaning on a 99.9% score on ARC-AGI-3.
The number is real. It is also achieved through OpenAI's own "Provider Adapter" harness, a setup that preserves hidden reasoning state between calls. Run the identical model through ARC Prize's neutral "Standard" harness, the version ARC Prize itself ran without OpenAI's scaffolding, and Astra scores 62.7%. That is a 37-point swing that tracks the change in scaffolding, not any claimed change in the model itself. ARC Prize's own writeup says outright: we are not claiming that it is AGI.
Why we rate this NEEDS CONTEXT, not FAILED
Astra is not a bad model. On Artificial Analysis's independent Intelligence Index, a broader aggregate less prone to a single-harness trick, it scores 61.2, effectively level with its own six-month-old predecessor GPT-5.6 Sol (60.9), and it trails Claude Fable 5.1 (65.7). On Humanity's Last Exam it loses outright to Fable 5.1, 57.2 to 65.0. Where Astra likely does better is agentic and computer-use work, precisely the areas OpenAI's own launch materials emphasized alongside the AGI framing. That part of the pitch is plausible on its own terms. The AGI framing riding on top of the benchmark number is not.
The steelman, and why it still fails
OpenAI's defenders will point out that every lab tunes its own serving stack, and that a harness which preserves reasoning state between calls is legitimate engineering, not a trick. Fair, as far as it goes. But the whole point of testing under a second, neutral harness is to check whether a score survives a change of scaffolding, and the record here shows this one does not: the same model lands 37 points lower the moment the benchmark's own inventor runs it outside OpenAI's harness. Whatever the Provider Adapter is to customers, on this test it is the difference between the headline and the result, and a number that needs the vendor's own scaffolding to exist is evidence of a well-optimized pipeline, not of general intelligence.
The mechanism
"AGI era" is a marketing phrase with no agreed technical threshold, which makes it cheap to claim and expensive to challenge. Pairing it with a saturating benchmark number gives the claim the appearance of a hard data point. The trick works because most readers will see "99.9% on ARC-AGI-3" and "AGI era" in the same sentence and assume the number backs the phrase. It backs a specific, non-reproducible test condition instead. Astra also costs 2.5x more per token than its own predecessor, which the launch framing conveniently leaves out of the AGI conversation.
What to do with this
- Before repeating any benchmark headline, find out whether the score came from a neutral, provider-agnostic harness or a vendor-built one.
- Evaluate Astra on the tasks it is genuinely strong at (agentic workflows, computer use), not on the AGI framing.
- Watch the Artificial Analysis Intelligence Index and Humanity's Last Exam over the next model cycle. Those are the numbers Astra is actually behind on.
- If a vendor's headline number requires their own custom harness to reproduce, treat the number as a demo, not a measurement.
If your team is budgeting on Astra having crossed into AGI, you are budgeting on a number OpenAI cannot reproduce outside its own harness. Buy it, if at all, for the agentic work it's plausibly strong at. Don't buy the headline.
By December 1, 2026, no neutral, provider-agnostic harness will reproduce anything within 15 points of 99.9% on ARC-AGI-3 for Astra. Hold us to it.
Flips if ARC Prize or a second independent lab reproduces a score above 85% on the Standard harness without OpenAI's Provider Adapter, or if OpenAI publishes the adapter's methodology and it survives replication.
RECEIPTS (5) · CONFIDENCE HIGH
every URL below answered a live HTTP check before publish · sweep 2026-09-05
- SUPPORTS THE CLAIM fortune.com ⧉ · "It's not unreasonable to feel that we are now in the AGI era"
- REFUTES IT arcprize.org ⧉ · "GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness"
- REFUTES IT arcprize.org ⧉ · "we are not claiming that it is AGI"
- ADDS CONTEXT requesty.ai ⧉ · "Astra is good, but maybe we should calm down with the hype"
- ADDS CONTEXT emergent.sh ⧉ · "the Artificial Analysis Intelligence Index, is exactly where Astra looks incremental and lands behind Fable 5.1."



