FRIDAY 28 AUGUST 2026THE DAILY AUTOPSY · ISSUE № 17LEDGER CURRENT 2026-08-28
The BS Killer
We read the paper so you don't have to. Receipts included.
+++ 80% TO 58% THE VIRAL 'MYSTERY MODEL BEATS GPT' SCORE FELL 22 POINTS ON THE FULL BENCHMARK +++ $71B THE FULL CAMPUS COSTS MORE TO BUILD THAN ANTHROPIC'S 6-YEAR RENTAL DEAL FOR PART OF IT +++ $0 NVIDIA DISCLOSED ZERO DOLLARS FOR 2 MILLION MORE GPUS THE SAME DAY ITS STOCK JUMPED 8.74% +++ ⚡ CAUGHT IN 4K, DAILY +++ 1 WEEK THE 'INDEPENDENT' REVIEW OF OPENAI'S OWN HACK ONLY COVERED THE 7 DAYS OPENAI PICKED +++ BS INDEX 40/100 HOW FULL OF IT THE AI INDUSTRY WAS THIS ISSUE, WEIGHTED AND AUDITABLE +++ WATCH ALIBABA QWEN: DAYS LEFT ON 'OPEN WEIGHTS FOR QWEN3.8-MAX 'NEXT WEEK'' +++ WATCH OPENAI: DAY WAITING ON 'PUBLISH THE ASTRA MATH PAPERS AND FORMAL' +++ WATCH VOLTA: DAYS LEFT ON 'NORWAY SITE FULLY DEPLOYED FOR THE '$10B' +++ WATCH APPLE V OPENAI: DAYS LEFT ON 'INJUNCTION HEARING ON THE TRADE-SECRETS ' +++ WATCH BLACK FOREST LABS: DAYS LEFT ON 'FLUX 3 DEV OPEN WEIGHTS 'LATER IN 2026'' +++ ISSUE № 17 NEW VERDICTS EVERY NIGHT, 9PM SYDNEY. BRING YOUR OWN SKEPTICISM +++ +++ 80% TO 58% THE VIRAL 'MYSTERY MODEL BEATS GPT' SCORE FELL 22 POINTS ON THE FULL BENCHMARK +++ $71B THE FULL CAMPUS COSTS MORE TO BUILD THAN ANTHROPIC'S 6-YEAR RENTAL DEAL FOR PART OF IT +++ $0 NVIDIA DISCLOSED ZERO DOLLARS FOR 2 MILLION MORE GPUS THE SAME DAY ITS STOCK JUMPED 8.74% +++ ⚡ CAUGHT IN 4K, DAILY +++ 1 WEEK THE 'INDEPENDENT' REVIEW OF OPENAI'S OWN HACK ONLY COVERED THE 7 DAYS OPENAI PICKED +++ BS INDEX 40/100 HOW FULL OF IT THE AI INDUSTRY WAS THIS ISSUE, WEIGHTED AND AUDITABLE +++ WATCH ALIBABA QWEN: DAYS LEFT ON 'OPEN WEIGHTS FOR QWEN3.8-MAX 'NEXT WEEK'' +++ WATCH OPENAI: DAY WAITING ON 'PUBLISH THE ASTRA MATH PAPERS AND FORMAL' +++ WATCH VOLTA: DAYS LEFT ON 'NORWAY SITE FULLY DEPLOYED FOR THE '$10B' +++ WATCH APPLE V OPENAI: DAYS LEFT ON 'INJUNCTION HEARING ON THE TRADE-SECRETS ' +++ WATCH BLACK FOREST LABS: DAYS LEFT ON 'FLUX 3 DEV OPEN WEIGHTS 'LATER IN 2026'' +++ ISSUE № 17 NEW VERDICTS EVERY NIGHT, 9PM SYDNEY. BRING YOUR OWN SKEPTICISM +++

THE AI BS REPORT · EVERY CLAIM CHECKED

They said it.
We checked it.

Every big AI claim, checked against independent evidence and stamped with a verdict you can audit. The internet is fluent. Fluent is not true. Receipts or it didn't happen.

765CLAIMS CHECKED · AI + BOOKS + PAPERS
2264RECEIPTS ON FILE
34%SURVIVED THE RECEIPTS
376PAPERS ON THE STAND
LATEST ISSUE · № 17 · 2026-08-28
GUESS THE VERDICT · CALL IT BEFORE YOU SCROLL

“A stealth model beat Claude and GPT on a coding benchmark. The sample size was ten questions.”

A stealth model beat Claude and GPT on a coding benchmark. The sample size was ten questions.

Its own tester ran the other 103 the next day and lost 17 points before an independent run lost more.

01THE CLAIM
"A viral claim on X said the mystery stealth model 'Ox Alpha' beats Claude Fable 5 and GPT-5.6 Sol on the DeepSWE coding benchmark, based on an 8/10-task sample (effectively >80%)." [SOURCE ↗]
TRUE, BUT7 SOURCES · LIVE 2026-08-25
BEN DAVIS (@DAVIS7) TRACK RECORD1 CLAIM · 40/100 BS RATE →
80%Ben Davis's original viral small-sample score for Ox Alpha on DeepSWE (8 of 10 tasks)
63%Ben Davis's own full 113-task re-run score, posted the next day
58.4%independent StealthModelWatch full 113-task run, 95% CI [49.2%, 67.1%]
72.7%Together AI's rigorous 4-trial pass@1 for GPT-5.6 Sol on DeepSWE
69.0%Together AI's rigorous 4-trial pass@1 for GLM-5.3, Ox Alpha's confirmed model family
A stealth model beat Claude and GPT on a coding benchmark. The sample size was ten questions.
02THE CHECK

THE CLAIM. A tester's viral post said the free stealth model Ox Alpha scored 80% on the DeepSWE coding benchmark, beating GPT-5.6 Sol's 52% and Claude Fable 5's 65%, and Coin Bureau broadcast it to a much larger audience as a mystery model beating the leaders. THE CHECK: the tester had run 10 of the benchmark's 113 tasks. He ran the rest the next day and landed at 63%. An independent lab ran all 113 and got 58.4%, calling the 80% figure completely incorrect. THE TWIST: the model's maker, Zhipu AI, has since confirmed its identity as GLM-5.3-Flash, and on a more rigorous multi-trial benchmark the base GLM-5.3 model (not confirmed as the same Flash variant) loses to GPT-5.6 Sol on a single attempt but ties or leads it once retries are allowed.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
"Ten questions is a party trick, not a benchmark."
DEEP DIVE · THE FULL AUTOPSY

What actually happened

On August 21, 2026, a tester posting as Ben Davis (@davis7) ran a free, unlabeled "stealth" model called Ox Alpha through a small slice of the DeepSWE coding benchmark, ten tasks out of a possible 113. The result: eight of ten solved, an effective score above 80%, against 52% for GPT-5.6 Sol and 65% for Claude Fable 5 on the same tiny slice. He posted the numbers with a caption admitting he was confused by them.

That caveat did not survive contact with the internet. The account AGTP repeated the numbers without the sample-size warning, and Coin Bureau, an account with a much larger, less technical audience, turned it into "BREAKING: A mysterious new AI model called Ox Alpha is reportedly BEATING Claude Fable 5 and GPT-5.6 Sol at coding, and nobody knows who built it." That version is what actually went viral.

The correction arrived fast, from the same person. The next day, Davis posted his own full 113-task run: roughly 63%, not 80%. A separate, independent tester (StealthModelWatch) ran the complete benchmark too and landed even lower, 58.4%, explicitly stating the rumored 80% pass rate was completely incorrect. On August 26, Zhipu AI (Z.ai) ended the identity mystery, confirming authorship and naming the model GLM-5.3-Flash.

Why we rate this needs_context

Every piece of the mechanical story checks out: the original 10-task tweet, the self-correction, the independent full run, the identity reveal. Nothing here is fabricated. What breaks is the headline framing, "beats Claude and GPT," which only ever existed in a sample size too small to mean anything, and which the original poster killed himself within 24 hours.

The steelman, and why it still falls short

The fairest pushback: a more rigorous, multi-trial benchmark from Together AI found GPT-5.6 Sol ahead of the base GLM-5.3 model on a single attempt (72.7% to 69.0%), but GLM-5.3 ties Sol at a second attempt (81.1% to 81.0%) and leads by a fourth (87.6% to 85.8%). For agentic workflows that permit retries, that is a real, non-hyped point in the model family's favor, with one caveat: this benchmark tested plain GLM-5.3, not confirmed as the same Flash checkpoint Zhipu named as Ox Alpha. "Ties on retries" and "beats outright, nobody knows who built it" are different claims regardless, and only the second one got 70 million impressions.

The mechanism

This is the standard shape of a benchmark-hype cycle: a small, honest sample produces a startling number, a caveat gets stripped on the first retweet, and the correction, when it comes, reaches a fraction of the audience the hype did. Ben Davis behaved well here, he ran the correction himself within a day. The system around him did not carry that correction as far as the claim.

What to do with this

  • Treat any benchmark score built from fewer than 30-40 tasks as a rumor, not a result, regardless of who posts it.
  • When a claim says "nobody knows who built it," check whether the maker has simply not spoken yet, that gap tends to close within a week, as it did here.
  • If you are choosing a model off a viral benchmark screenshot, wait for the full-suite number. It is usually lower, and sometimes by more than 20 points.
04YOUR MOVE ⚡ WHAT IGNORING THIS COSTS

Ten questions is a coin flip with a leaderboard attached, and if that tweet made it into your team's model-selection deck this week, you copied a number its own author disowned the next day.

05🔮 OUR CALL · ON THE RECORD 2026-08-28

By Sept 27, 2026, expect any audited full-113-task DeepSWE score published for the confirmed model to land in the high 50s to mid 60s, not 80%. Hold us to it.

Flips if a different lab reruns the full 113-task set on the confirmed model and reproduces a score above 75%.

RECEIPTS (7) · CONFIDENCE HIGH

every URL below answered a live HTTP check before publish · sweep 2026-08-25

  • x.com · "gpt-5.6-sol: 52% fable: 65% whatever the hell this is: 80%"
  • x.com · "Actual DeepSWE run on the ox alpha mystery model is done. Ended at ~63% NOT the 80% my first subset test got"
  • stealthmodelwatch.online · "The rumored ~80% pass rate is completely incorrect"
  • together.ai · "Sol edges ahead: 72.7% pass@1 to GLM-5.3's 69.0%"
  • x.com · "BREAKING: A mysterious new AI model called "Ox Alpha" is reportedly BEATING Claude Fable 5 and GPT-5.6 Sol at coding, and nobody knows who built it"
  • vktr.com · "On August 26, Z.ai confirmed authorship of the model, vindicating the fingerprinting. The model's official name is GLM-5.3-Flash"
  • together.ai · "It ties Sol at pass@2 (81.1 vs. 81.0) and leads pass@4 (87.6% vs. 85.8%)"
NEW VERDICTS EVERY NIGHT AT 9. GET THEM IN YOUR INBOX.
TODAY'S AUTOPSIES · 3 MORE CLAIMS ON THE STAND
THAT WAS THE ISSUE. THE NEXT ONE CAN COME TO YOU.
THAT'S THE LATEST ISSUE.
PREVIOUS EDITIONS
SEE ALL 17 ISSUES IN THE ARCHIVE →
EXPLORE THE BS KILLER · SAME METHOD, EVERY FORMAT

THE WITNESS STAND

376 papers have testified in our verdicts. Weekly full paper autopsies begin per the roadmap.

READ THE STAND →

REELS

48 verdicts as 40-second films. True stories, receipts attached.

WATCH →

HOW A VERDICT IS MADE

CLAIM → CHECK → VERDICT → RECEIPTS

THE METHOD →
THE INSTRUMENT · WHY MEMBERS PAY

The receipt drawer for the whole AI industry.

Free gets you today's verdict. Membership gets you the instrument: search every claim we have ever checked, filter by verdict or company, and pull any lab's full track record and BS rate. The receipts, on tap.

Stop reading press releases disguised as news.

One email a day: the AI industry's loudest claims, each checked against independent sources and stamped with a verdict you can audit. Print a claim without a receipt and the correction runs at the top of the next issue. That is the whole deal.

FREE · ONE EMAIL A DAY · UNSUBSCRIBE IN ONE CLICK · RECEIPTS OR IT DIDN'T HAPPEN

Want audio editions, every receipt with screenshots, and unlimited dossiers? Start a free Pro trial →

MEMBERSHIP · NO SURPRISES

Free unlocks the investigation. Pro gives you the evidence toolkit.

FREE
  • The nightly issue, every verdict
  • Full deep-dive autopsies with a free account
  • Search the ledger · 5 dossiers a month
CREATE A FREE ACCOUNT
PRO · A$89/YR
  • Audio editions of every deep dive
  • Every receipt with screenshots
  • Unlimited company & person dossiers
That is A$0.24 a day. One AI claim you repeated and got wrong costs more.
START YOUR FREE TRIAL ★
1 month free · cancel anytime, keep the free daily