Inherent says its 27-billion-parameter model beats GPT-5.5. Its own paper says the model calls GPT-5.5 to do the work.
The "AI Scientist" that outperformed two frontier labs turns out to have one of them running inside it.
"Faraday, a 27-billion-parameter 'AI Scientist' from Inherent Labs, outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers." [SOURCE ↗]

THE PITCH. Faraday, a 27B model from Inherent Labs, "outperforms Claude Opus 4.8 and GPT-5.5" at replicating research papers, in a launch that landed a TechCrunch feature and a live press cycle.
THE CATCH. Read the arXiv methods section and Faraday hands every coding subtask to GPT-5.5 Codex, "including at evaluation time." The 27B number describes the orchestrator, not the system doing the work being scored.
THE NUMBER THAT EXPLAINS EVERYTHING. 73% in-distribution, but only 60% out-of-distribution, a 13-point drop on held-out science tasks, on Inherent's own 310-task benchmark, judged by Inherent's own rubric, against baselines already a generation out of date by launch day.
WHAT NOBODY SAYS OUT LOUD. this is "GPT-5.5 plus a wrapper" beating "GPT-5.5 alone," scored by the company that built the wrapper.
When a benchmark result includes a competitor's model as a component, the fair comparison is wrapper-vs-no-wrapper, not vendor-vs-vendor. Ask what's inside before you believe what beat what.
No independent third-party rerun of Replica within 60 days. If one happens, Faraday's edge over bare GPT-5.5 with a simple harness narrows to single digits. Hold us to it.
Flips toward Inherent if an independent lab reruns Replica (or a public benchmark like MLE-Bench) and confirms a comparable gap using open baselines. Flips toward "just a wrapper" if removing GPT-5.5 Codex collapses Faraday's score.
RECEIPTS (6) · CONFIDENCE HIGH · every URL below answered a live HTTP check before publish · sweep 2026-08-25
- ▼ arxiv.org ⧉ · "Faraday is provided with a frontier coding agent to use as a tool. A wrapper script runs the Codex CLI non-interactively."
- ▼ arxiv.org ⧉ · "GPT-5.5 in the final stage and for evaluation"
- ▼ aiweekly.co ⧉ · "Faraday is built on a 27-billion-parameter Qwen base and calls OpenAI's GPT-5.5 Codex for coding subtasks."
- ● superpowerdaily.com ⧉ · "The disclosed evaluation includes no numerical scores, named test papers or methodology, limiting what can be concluded from the comparison."
- ▲ techcrunch.com ⧉ · "What was most interesting to us about this was not so much the result of beating those frontier agents"
- ▲ inherentlabs.ai ⧉ · "a 27B-parameter "AI scientist" agent that outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research"