Nvidia's agent just went 100% on a benchmark built to resist that. The brain doing the reasoning is Anthropic's, and it scores 30% alone.
A perfect score on the public half of a test that has never been beaten on the half that counts.
"NVIDIA AVO achieved a 100.00 RHAE score across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels, demonstrating a frontier-level general-purpose architecture for long-horizon autonomous agents." [SOURCE ↗]

THE PITCH. NVIDIA AVO hit a perfect 100.00 RHAE score across all 25 environments and 183 levels of the ARC-AGI-3 public benchmark, pitched by Nvidia as proof of "a frontier-level general-purpose architecture."
THE CATCH. Nvidia's own post says the private and semi-private sets, the parts of ARC-AGI-3 nobody can rehearse against, remain unsolved by AVO and every other system. The reasoning inside AVO is Anthropic's Claude Opus 5, which scores 30.2% completely alone.
THE NUMBER THAT EXPLAINS EVERYTHING. 100.00 on the set you can practice against, 0 systems have ever cracked the set you can't.
WHAT NOBODY SAYS OUT LOUD. Nvidia's own post admits its comparison to the prior leaderboard entry "should not be interpreted as a controlled ablation." The architecture claim rests on a model Nvidia did not train.
"Frontier architecture" claims that credit a harness, not the underlying model, should be read as harness engineering, real work, but a different claim than general intelligence.
No system, including AVO, clears the ARC-AGI-3 private set before 2027. Hold us to it.
Flips if AVO or any system posts a verified private-set score above 50%, or if Nvidia publishes a true controlled ablation isolating the harness from the underlying model's contribution.
RECEIPTS (5) · CONFIDENCE HIGH · every URL below answered a live HTTP check before publish · sweep 2026-08-25
- ▼ developer.nvidia.com ⧉ · "They are not results on the semi-private or fully private competition sets."
- ▼ developer.nvidia.com ⧉ · "This should not be interpreted as a controlled ablation: the two systems differ in agent backend, observation representation, memory, context management"
- ▼ eneralabs.com ⧉ · "The private-set benchmark, which withholds environments not available in the public set, remains unsolved."
- ▼ cryptobriefing.com ⧉ · "Without private set results, the 100% score is impressive but incomplete as a measure of general reasoning ability."
- ● cryptobriefing.com ⧉ · "AVO exists as a research demonstration for now, not a product you can buy or integrate."