GET THE AUTOPSY ➔

GPT-5.6 Terra scored 69.6 and 64.8 on the same benchmark this week. Nothing changed except whose chart it was.

Every score is real and every score is a different lab privately running a suite that shares only a name. The two available cross-checks disagree by about five points each, and both disagreements point in the house's favor.

01THE CLAIM
"Weekend coverage ranks the week's launches on DeepSWE v1.1 as one leaderboard: GPT-5.6 Terra 69.6, Grok 4.6 High 65.9, Gemini 3.7 Flash 65.3, Muse Spark 1.2 59.3, with each lab's launch chart cited as if the numbers share a scale." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-25
GOOGLE TRACK RECORD17 CLAIMS · 41/100 BS RATE →
69.6%GPT-5.6 TERRA ON GOOGLE'S DEEPSWE CHART
64.8%THE SAME TERRA ON META'S DEEPSWE CHART
59.3%MUSE SPARK PER META. GOOGLE'S CHART SAYS 54.9%
GPT-5.6 Terra scored 69.6 and 64.8 on the same benchmark this week. Nothing changed except whose chart it was.
02THE CHECK

THE CLAIM. the week's coding-agent launches stack neatly on DeepSWE v1.1: GPT-5.6 Terra 69.6, Grok 4.6 High 65.9, Gemini 3.7 Flash 65.3, Muse Spark 1.2 59.3. Aggregators and social posts quote these as one ranking.

THE CHECK. no shared harness produced those numbers. Google's launch chart, Meta's launch chart, and xAI's launch table each ran a same-named suite independently, and the same models land in both of the two charts that overlap. That overlap is the tell. GPT-5.6 Terra: 69.6 on Google's chart, 64.8 on Meta's. Muse Spark 1.2: 59.3 on Meta's chart, 54.9 on Google's. Roughly five points of daylight per model, and in each case the chart owner's rival scores lower on the owner's chart. As the one careful comparison in circulation puts it, cross-chart readings are 'suggestive but not proof, since it's two different labs running the same-named suite independently', or shorter: 'a shared benchmark name is not a shared benchmark'. The individual numbers are ordinary first-party launch stats. The leaderboard assembled from them is fiction with a spreadsheet aesthetic.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Terra is 69.6 on Google's chart and 64.8 on Meta's. Same model, same benchmark name, five points apart. Every cross-lab gap smaller than that is noise wearing a ranking."

Three coding-model launches landed within four days: xAI's Grok 4.6 on August 12, Google's Gemini 3.7 Flash on August 13, and the ongoing rollout of Meta's Muse Code agent on Muse Spark 1.2. Each launch shipped a chart, and each chart included a benchmark called DeepSWE v1.1, a long-horizon software

🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNT

You just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.