GET THE AUTOPSY ➔

GPT-5.6 Sol scores 92.5% on ARC-AGI-2, the test built so AI would fail it. Read that as abstract reasoning solved, then look one column over on the same scorecard: the same model, same maximum effort, scores 7.78% on ARC-AGI-3, the interactive benchmark the same team built next, where humans still score 100%.

ARC-AGI-2 was designed to resist pattern-matching, and a frontier model just cleared it well above the average human's 66%. Genuine milestone. But it needs Sol's most expensive reasoning setting; turned down it falls to 42.5%, and the ARC Prize's cost-capped grand prize stays unclaimed. Then ARC-AGI-3, a new interactive benchmark from the same team, drops the same model to 7.78% while humans still score 100%. Each version gets saturated and the next reopens the gap. That is a treadmill, not a mind acquiring general reasoning.

01THE CLAIM
"GPT-5.6 Sol leads the ARC-AGI-2 leaderboard at 92.5%, clearing the abstract-reasoning benchmark built specifically to resist AI and beating the average human, a result read as frontier models cracking fluid reasoning." [SOURCE ↗]
TRUE, BUT7 SOURCES · LIVE 2026-08-25
ARC PRIZE VERIFIED LEADERBOARD AND COVERAGE TRACK RECORD1 CLAIM · 40/100 BS RATE →
92.5%GPT-5.6 SOL ON ARC-AGI-2 AT MAX REASONING; ABOVE THE 66% AVERAGE HUMAN
7.78%SAME MODEL, SAME SETTING, ON ARC-AGI-3, THE NEW INTERACTIVE BENCHMARK
100%WHAT HUMAN TESTERS STILL SOLVE ON ARC-AGI-3
42.5%WHERE ARC-AGI-2 FALLS WHEN THE REASONING EFFORT IS TURNED DOWN
GPT-5.6 Sol scores 92.5% on ARC-AGI-2, the test built so AI would fail it. Read that as abstract reasoning solved, then look one column over on the same scorecard: the same model, same maximum effort, scores 7.78% on ARC-AGI-3, the interactive benchmark the same team built next, where humans still score 100%.
02THE CHECK

THE CLAIM. GPT-5.6 Sol tops the ARC-AGI-2 leaderboard at 92.5%, clearing the benchmark built to isolate fluid reasoning and resist memorization, and beating the average human's 66%. The read everywhere is that frontier models have essentially cracked abstract reasoning.

THE CHECK. the score is real and ARC Prize verified, and on its own terms it is a genuine jump. Two things temper the headline. First, the 92.5% depends on Sol's most expensive reasoning mode; dial the effort down and ARC-AGI-2 collapses to 42.5%, and the ARC Prize's own cost-capped grand prize, which needs above 85% cheaply, is still unclaimed. Second, ARC-AGI is a series that keeps raising the bar, and the newest rung is a different kind of test. ARC-AGI-3 is an interactive, agentic benchmark, where a model must explore an unfamiliar environment, infer the goal, and plan, and there the same model at the same setting scores 7.78% while human testers solve 100%. That does not prove the ARC-AGI-2 result was hollow. It shows the fluid-reasoning progress does not yet extend to agentic novelty, and that every time this team builds a new test, today's models fail it until they catch up. Saturating a benchmark is not the same as acquiring the ability it was built to isolate.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"GPT-5.6 Sol really does hit 92.5% on ARC-AGI-2, above the 66% average human, and that is a real jump. But the same model at the same setting scores 7.78% on ARC-AGI-3, which humans still solve 100% of, and the 92.5% needs its most expensive reasoning mode. That is a benchmark being saturated, not reasoning being solved."

GPT-5.6 Sol sits on top of the ARC-AGI-2 leaderboard at 92.5%. ARC-AGI-2 is not a trivia test. It is the benchmark François Chollet's team built specifically to resist the pattern-matching that let models saturate earlier evaluations, a set of visual puzzles designed to isolate fluid reasoning on ta

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 7 sources with quotes and screenshots, and our on-record call.

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.