GET THE AUTOPSY ➔

OpenAI moved its own biorisk red line by 20 points and called the model that cleared it safe.

A safety threshold that says 50% one quarter and 30% the next is not a stricter bar, it is a bar that stopped holding still.

01THE CLAIM
"OpenAI's GPT-5.6 System Card states GPT-5.6 Sol scores below its indicative biorisk threshold for protein-binding capability, using '30% as an indicative threshold, based on a survey of 20 independent experts.'" [SOURCE ↗]
TRUE, BUT5 SOURCES · LIVE 2026-08-25
OPENAI - GPT-5.6 SYSTEM CARD TRACK RECORD1 CLAIM · 40/100 BS RATE →
50% -> 30%protein-binding biorisk threshold, GPT-5.5 card (Apr) to GPT-5.6 card (Jul) - 20-point move
0times the word 'survey' appears in the April GPT-5.5 card, despite both new thresholds being credited to a 20-expert survey
118 daysa pass@1 value sat mislabeled as pass@4 across two published system cards before an Aug 19 correction
3.7xsize of the Aug 19 correction (0.4% to 1.48%), which shrank the apparent generational jump from 19x to 5.1x
OpenAI moved its own biorisk red line by 20 points and called the model that cleared it safe.
02THE CHECK

THE PITCH. OpenAI's GPT-5.6 System Card says the model "scores below" its indicative biorisk threshold for protein-binding capability, now set at 30%, "based on a survey of 20 independent experts."

THE CATCH. The prior card, published 11 weeks earlier, set that same threshold at 50%. The word "survey" appears zero times in that April document, despite both thresholds now credited to the same 20-expert process.

THE NUMBER THAT EXPLAINS EVERYTHING. 118 days. That is how long a mislabeled score sat published across two system cards before an August 19 correction that shrank the model's apparent generational jump from 19x to 5.1x.

WHAT NOBODY SAYS OUT LOUD. the protein threshold got stricter, which cuts against a bad-faith reading, but a threshold that moves 20 points in 11 weeks on a retrofitted justification is not a fixed line, it is one drawn after the fact.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
""Which threshold, dated when, and what changed the number since the last card?""
04YOUR MOVE ⚡ WHAT IGNORING THIS COSTS

A safety margin measured against a threshold that moves is not a fixed margin. When a lab reports "below threshold," ask which version of the threshold, and when it last changed.

05🔮 OUR CALL · ON THE RECORD 2026-08-25

OpenAI adjusts at least one more Preparedness Framework threshold within two system card cycles, by mid-2027, without a standalone announcement. Hold us to it.

Flips toward "normal calibration" if OpenAI publishes the raw 20-expert survey data and shows the methodology was genuinely unchanged between cards. Flips toward "worse" if a third threshold moves in the permissive direction without disclosure.

RECEIPTS (5) · CONFIDENCE HIGH · every URL below answered a live HTTP check before publish · sweep 2026-08-25

  • deploymentsafety.openai.com · "Accordingly, we propose 50% correctness as the threshold for biorisk concern."
  • deploymentsafety.openai.com · "August 19, 2026: We corrected GPT-5.5's pass@4 score on the hard-negative protein binding prediction evaluation from 0.4% to 1.48%."
  • securebio.substack.com · "identifies an approach that successfully evades certain screening systems, but would be inconvenient for a malicious actor to carry out in practice"
  • arxiv.org · "The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices"
  • kingy.ai · "OpenAI believes the family is capable enough in biological and chemical domains to trigger stronger safeguards."

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.