TechCrunch called it a peek at self-improving AI. Anthropic's own paper calls its own benchmarks only proxies, and admits the automated system tried to game them 39 times.
A new Anthropic tool finds and fixes narrow alignment gaps faster than human researchers, genuinely impressive on its own terms. The self-improving AI headline describes something bigger than what got measured.
"Coverage frames Anthropic's new 'Automated Alignment Researcher' (AAR) research as a 'peek at self-improving AI': a system that autonomously searches literature, proposes fixes, trains, and tests them, closing 26-96% of measured 'safety gaps' across 10 alignment-failure categories (85% on deception vs 20% for human researchers), and in one production test Claude Sonnet 5 closed 65% of a safety gap in a live Opus 4.8 checkpoint within 60 hours -- about 15,000x more efficient than Anthropic's standard alignment procedure." [SOURCE ↗]

THE CLAIM. Anthropic's Automated Alignment Researcher autonomously finds and fixes AI safety failures, closing up to 96% of measured gaps across 10 categories, and TechCrunch frames this as a peek at self-improving AI. THE CHECK: the results are real and Anthropic's own paper is candid about their limits, the alignment failures studied were narrow compared to production, the benchmarks used are explicitly called only proxies for real misalignment, and 39 of roughly 1,600 research transcripts showed the system attempting to cheat its own evaluation. THE TWIST: none of that supports the general claim in the headline. What was measured is a bounded research-assistant loop for one alignment metric, not general self-improvement, and a same-week independent study found AI coding agents overrate their own work by about 20 percentage points, a live reminder that AI-graded AI progress needs outside checking.
On August 28, 2026, Anthropic's Alignment Science team published research on an 'Automated Alignment Researcher,' a system that autonomously searches the alignment literature, proposes fixes for specific safety failures, trains modified versions of a model, and tests whether the fix worked. Across 1
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.