GPT-5.6 Sol scores 92.5% on ARC-AGI-2, the test built so AI would fail it. Read that as abstract reasoning solved, then look one column over on the same scorecard: the same model, same maximum effort, scores 7.78% on ARC-AGI-3, the interactive benchmark the same team built next, where humans still score 100%.
ARC-AGI-2 was designed to resist pattern-matching, and a frontier model just cleared it well above the average human's 66%. Genuine milestone. But it needs Sol's most expensive reasoning setting; turned down it falls to 42.5%, and the ARC Prize's cost-capped grand prize stays unclaimed. Then ARC-AGI-3, a new interactive benchmark from the same team, drops the same model to 7.78% while humans still score 100%. Each version gets saturated and the next reopens the gap. That is a treadmill, not a mind acquiring general reasoning.
"GPT-5.6 Sol leads the ARC-AGI-2 leaderboard at 92.5%, clearing the abstract-reasoning benchmark built specifically to resist AI and beating the average human, a result read as frontier models cracking fluid reasoning." [SOURCE ↗]

THE CLAIM. GPT-5.6 Sol tops the ARC-AGI-2 leaderboard at 92.5%, clearing the benchmark built to isolate fluid reasoning and resist memorization, and beating the average human's 66%. The read everywhere is that frontier models have essentially cracked abstract reasoning.
THE CHECK. the score is real and ARC Prize verified, and on its own terms it is a genuine jump. Two things temper the headline. First, the 92.5% depends on Sol's most expensive reasoning mode; dial the effort down and ARC-AGI-2 collapses to 42.5%, and the ARC Prize's own cost-capped grand prize, which needs above 85% cheaply, is still unclaimed. Second, ARC-AGI is a series that keeps raising the bar, and the newest rung is a different kind of test. ARC-AGI-3 is an interactive, agentic benchmark, where a model must explore an unfamiliar environment, infer the goal, and plan, and there the same model at the same setting scores 7.78% while human testers solve 100%. That does not prove the ARC-AGI-2 result was hollow. It shows the fluid-reasoning progress does not yet extend to agentic novelty, and that every time this team builds a new test, today's models fail it until they catch up. Saturating a benchmark is not the same as acquiring the ability it was built to isolate.
GPT-5.6 Sol sits on top of the ARC-AGI-2 leaderboard at 92.5%. ARC-AGI-2 is not a trivia test. It is the benchmark François Chollet's team built specifically to resist the pattern-matching that let models saturate earlier evaluations, a set of visual puzzles designed to isolate fluid reasoning on ta
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 7 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.