The coding number everyone quotes says AI has nearly solved software engineering: Claude Opus 5 scores 96% on SWE-bench Verified. Move to the benchmark built to resist contamination and the top model sits at 80.3%, and GPT-5.6 Sol lands at 64.6%.
SWE-bench Verified, the coding leaderboard in every launch post, is saturating; benchlm itself calls it nearing saturation, and OpenAI stopped using it in February, with a separate audit finding most of the tasks it checked had flawed tests. Its successor, SWE-bench Pro, is both cleaner and harder: it uses actively maintained repositories with no public ground-truth leakage, plus longer, multi-file, enterprise-scale fixes. There the ceiling drops to 80.3% while GPT-5.6 Sol comes in at 64.6%. The 96% is a saturated, leaky ruler; the honest number is lower.
"Frontier models have all but solved software engineering: on SWE-bench Verified, the coding benchmark quoted in launch posts, the top models now sit near 96%, with Claude Opus 5 leading at 96%." [SOURCE ↗]

THE CLAIM. coding is nearly a solved problem for frontier models. The number carrying that story is SWE-bench Verified, where the top of the board is now packed near 96%, Claude Opus 5 in front at 96%, a figure that reads as almost every real bug fixed.
THE CHECK. that benchmark is worn out and leaky. benchlm's own leaderboard note says the score is 'nearing saturation for frontier models,' meaning it can no longer tell the best models apart, and OpenAI publicly stopped using SWE-bench Verified in February 2026 after an audit found, in the auditors' words, that '59.4% of audited problems contain flawed test cases that reject correct solutions.' The honest successor is SWE-bench Pro, which, unlike Verified, 'uses actively maintained repositories with no public ground-truth leakage.' On Pro the ceiling falls to 80.3%, and the spread that Verified had flattened reopens: GPT-5.6 Sol, near the top of every coding conversation, sits at 64.6%. The models are strong. Solved is a word the quoted benchmark can no longer support.
Every coding launch this year leans on one benchmark, SWE-bench Verified, a set of real GitHub issues a model has to fix so its patch passes the repository's tests. As of the August 18 leaderboard, 'Claude Opus 5 leads the SWE-bench Verified leaderboard on BenchLM's August 2026 update with 96%,' and
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 6 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access — this looks like our error, not yours.