GET THE AUTOPSY ➔

Inherent says its 27-billion-parameter model beats GPT-5.5. Its own paper says the model calls GPT-5.5 to do the work.

The "AI Scientist" that outperformed two frontier labs turns out to have one of them running inside it.

01THE CLAIM
"Faraday, a 27-billion-parameter 'AI Scientist' from Inherent Labs, outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers." [SOURCE ↗]
TRUE, BUT6 SOURCES · LIVE 2026-08-25
INHERENT LABS TRACK RECORD1 CLAIM · 40/100 BS RATE →
27BFaraday's advertised parameter count (base model: Qwen 3.6)
73%share of in-distribution Replica tasks Faraday beats Opus 4.8 and GPT-5.5 on
60%win rate on out-of-distribution tasks (13-point drop off held-out data)
310total Replica benchmark tasks, drawn from 100 papers - all Inherent's own
$50Mseed round backing the launch
Inherent says its 27-billion-parameter model beats GPT-5.5. Its own paper says the model calls GPT-5.5 to do the work.
02THE CHECK

THE PITCH. Faraday, a 27B model from Inherent Labs, "outperforms Claude Opus 4.8 and GPT-5.5" at replicating research papers, in a launch that landed a TechCrunch feature and a live press cycle.

THE CATCH. Read the arXiv methods section and Faraday hands every coding subtask to GPT-5.5 Codex, "including at evaluation time." The 27B number describes the orchestrator, not the system doing the work being scored.

THE NUMBER THAT EXPLAINS EVERYTHING. 73% in-distribution, but only 60% out-of-distribution, a 13-point drop on held-out science tasks, on Inherent's own 310-task benchmark, judged by Inherent's own rubric, against baselines already a generation out of date by launch day.

WHAT NOBODY SAYS OUT LOUD. this is "GPT-5.5 plus a wrapper" beating "GPT-5.5 alone," scored by the company that built the wrapper.

03SAY THIS IN THE MEETING · 📸 SCREENSHOT IT
""Which parts of the pipeline are your 27B model, and which parts are GPT-5.5 doing the heavy lifting?""
04YOUR MOVE ⚡ WHAT IGNORING THIS COSTS

When a benchmark result includes a competitor's model as a component, the fair comparison is wrapper-vs-no-wrapper, not vendor-vs-vendor. Ask what's inside before you believe what beat what.

05🔮 OUR CALL · ON THE RECORD 2026-08-25

No independent third-party rerun of Replica within 60 days. If one happens, Faraday's edge over bare GPT-5.5 with a simple harness narrows to single digits. Hold us to it.

Flips toward Inherent if an independent lab reruns Replica (or a public benchmark like MLE-Bench) and confirms a comparable gap using open baselines. Flips toward "just a wrapper" if removing GPT-5.5 Codex collapses Faraday's score.

RECEIPTS (6) · CONFIDENCE HIGH · every URL below answered a live HTTP check before publish · sweep 2026-08-25

  • arxiv.org · "Faraday is provided with a frontier coding agent to use as a tool. A wrapper script runs the Codex CLI non-interactively."
  • arxiv.org · "GPT-5.5 in the final stage and for evaluation"
  • aiweekly.co · "Faraday is built on a 27-billion-parameter Qwen base and calls OpenAI's GPT-5.5 Codex for coding subtasks."
  • superpowerdaily.com · "The disclosed evaluation includes no numerical scores, named test papers or methodology, limiting what can be concluded from the comparison."
  • techcrunch.com · "What was most interesting to us about this was not so much the result of beating those frontier agents"
  • inherentlabs.ai · "a 27B-parameter "AI scientist" agent that outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research"

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.