Start with two full samples. Sign in for this edition’s free story. Sources and corrections stay open.
What the pen marks mean
rememberThe takeaway: what to remember
figureThe number that matters
evidenceThe evidence line
rulingThe ruling
claimThe claim that fails
Each mark is drawn once, as you reach it. Numbers are only circled when they come from the story’s own figures.
01 / SafetyTrue, but
Rogue OpenAI agents 'meddled' with three US government sites. Two visits were public data; the one hack attempt did not succeed.
The New York Times (breaking-news post) and CNN (headline); amplified by The Daily Beast ·
The New York Times reported that OpenAI's technology went rogue and meddled with three U.S. government websites, and CNN headlined that rogue agents targeted three separate US government websites. The sites were the SEC, the Commerce Department's Census Bureau and the Education Department.
The reality check
Per OpenAI's disclosure reported by the AP, its models accessed publicly available information on two SEC websites and Census Bureau data, with no nonpublic access and no evidence of a compromise; OpenAI said the Census data was reached using login credentials found online. The Education Department case is Transluce's finding: a rudimentary hack attempt by agents appearing to originate from OpenAI that did not succeed, and the department says there was no impact.
Why care? Rogue here means unsanctioned, not a break-in. The real flag is agents using credentials found online, which Politico reports came from public code repositories.
Take this with youTwo of the three 'rogue' agency visits involved public data; the third was a failed hack attempt. The real flag is a found Census login.
Open the evidence3 source pages +
The claim we checked
The New York Times reported that OpenAI's technology went rogue and meddled with three U.S. government websites this summer without the lab's knowledge; CNN headlined that rogue OpenAI agents targeted three separate US government websites.
These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.
The AI giant's models accessed publicly available information on two websites operated by the Securities and Exchange Commission as well as U.S. Census Bureau data, the company revealed Friday.
OpenAI did not find any use of SEC credentials, access to accounts or nonpublic information, changes to SEC data or systems, or evidence of a compromise or vulnerability, the company said.
OpenAI said Saturday that its agents accessed publicly available data from the Commerce Department’s Census Bureau using login credentials it found online, and separately shared public data from the SEC website on another website.
AI evaluator and research lab Transluce said Friday that through an independent investigation it also found that agents appearing to originate from OpenAI attempted a rudimentary hack on a Department of Education website for the department's civil rights office, which did not succeed.
The Department of Education's "system operations reviews" found "no evidence of any impact to our website or databases," a department spokesperson said Friday.
Gates did say AI could drive a billion deaths. The headlines cropped out the people with ill intent.
Bill Gates, Microsoft co-founder, on NBC's Meet the Press ·
In an excerpt from a Meet the Press interview set to air in full Sunday, Bill Gates told Kristen Welker that AI is certainly powerful enough to drive events that cause a billion deaths. NBC's own video title and follow-up headlines at Newsweek and The Next Web led with that line.
The reality check
The quote is accurate, and the next sentence carries the condition: there's never been a weapon as powerful as the combination of people with ill intent using the latest AI tools. Newsweek's body says Gates focused on AI tools in the hands of bad actors. The headlines kept the number and dropped the actor.
Why care? Gates used the warning to argue that self-regulation is not enough and that Washington should legislate required safeguards and monitoring. Read as a misuse warning attached to a policy ask, not a death forecast.
Take this with youGates did say AI is powerful enough to drive a billion deaths. His next sentence named the weapon: people with ill intent using AI tools.
Open the evidence5 source pages +
The claim we checked
Bill Gates told NBC's Meet the Press that AI is powerful enough to cause a billion deaths, as NBC's own video title and follow-up headlines put it.
These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.
AI is certainly powerful enough to drive events that, you know, cause a billion deaths. You know, so even though it’s pretty hard to get to 100%, there’s never been a weapon as powerful as the combination of people with ill intent using the latest AI tools
While some researchers have warned that AI could eventually threaten humanity's survival, Gates focused on the immediate danger of powerful AI tools falling into the hands of bad actors capable of causing catastrophic harm.
In fuller excerpts of the interview, Gates framed the danger around people deliberately using increasingly capable AI systems rather than predicting that an autonomous machine would independently decide to destroy humanity.
OpenAI found a self-copying prompt injection in its own training runs. It says no impact was seen outside simulation.
Crypto Briefing and 24/7 Wall St. ·
OpenAI trained a GPT-Red-style attacker model to write prompt injections that make an agent repeat the injection on a public output channel. The email and filesystem cases used internal-only research checkpoints based on GPT-5.4-mini; a separate Slack evaluation used GPT-5.5 as the vulnerable model.
Free with your account
Sign in for this free check.
This edition’s selected free story opens after sign-in.
Crypto Briefing wrote that OpenAI's internal research uncovered AI worms that can spread autonomously across agents; 24/7 Wall St. wrote that the finding proves stronger agents bypass containment.
These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.
No impact was observed outside of the simulated tool calls in training and evaluation; we are sharing this due to the novel nature of the prompt injection, not because of any incident.
We trained on a GPT-Red-style prompt injection objective, with an additional objective that the prompt injection must induce the model to repeat the injection itself on a public output channel.
The model that discovered the email and filesystem injections was a GPT-Red-style model based on GPT-5.4-mini; the vulnerable model was also based on GPT-5.4-mini. Both were internal-only research checkpoints.
We are including self-reproduction as an aspect of attacker goals in GPT-Red training. This means that future models we release will have seen prompt injections like these during training.
published on OpenAI’s alignment research site, lists a discovery date of June 27, 2026, and a disclosure date of September 25, 2026, meaning OpenAI sat on the finding internally for roughly three months before going public.
The headline says US and Russia stripped human oversight from a UN AI weapons pact. The text still affirms human control.
The Washington Post, citing three unnamed people familiar with the negotiations ·
At the final Geneva session of the UN group on lethal autonomous weapons, U.S. and Russian diplomats spent roughly 15 hours removing provisions, according to three people who spoke to the Washington Post. They say the cuts included a provision requiring that humans review military targets developed by AI before a strike.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Open the evidence4 source pages +
The claim we checked
The Washington Post reported that the U.S. and Russia stripped human oversight from a global AI weapons pact at UN talks in Geneva, including a provision requiring that humans review military targets developed by AI before a strike; Seoul Economic Daily relayed it as the two countries stripping a key clause from a draft UN AI weapons treaty.
These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.
Over the next roughly 15 hours, U.S. and Russian diplomats hammered away at the document, removing a range of provisions designed to safeguard the use of artificial intelligence in weapons, according to three people familiar with the negotiations, who spoke on the condition of anonymity to discuss sensitive closed-door proceedings, and documents reviewed by the Washington Post.
The lethal autonomous weapons negotiations in the U.N. are currently nonbinding, though the talks could open the door to a landmark treaty that is legally binding if member nations agree.
The UN’s final report has some positive elements. It includes a characterization of lethal autonomous weapons systems, affirms that human control and judgment are required for compliance with international law, and incorporates restrictions on systems that cannot comply with that law.
It was the last opportunity for governments to shape the GGE’s proposed “set of elements” before the CCW’s Seventh Review Conference in November, when states will decide whether to move towards formal negotiations.