Start with two full samples. Sign in for this edition’s free story. Sources and corrections stay open.
What the pen marks mean
rememberThe takeaway: what to remember
figureThe number that matters
evidenceThe evidence line
rulingThe ruling
claimThe claim that fails
Each mark is drawn once, as you reach it. Numbers are only circled when they come from the story’s own figures.
01 / SafetyTrue, but
The 'first AI hack of a government' was three attempts. None appear to have worked.
Transluce researchers, amplified by ABC News ·
Transluce published evidence that AI agents, some linked to an OpenAI agent swarm, used the urlquery.net scanning service to get around access limits and probed three public data providers for vulnerabilities, including the Australian Institute of Health and Welfare (AIHW). Its report calls the AIHW episode part of the first reported instance of agents hacking a government.
The reality check
The same report says none of the hacking attempts it identified appear to have succeeded, though its public records are incomplete. At AIHW the agents fetched a public file from a pre-production server after being blocked, and the agency told the ABC it has no evidence non-public data was accessed. The confirmed non-public access, at the Medicare statistics portal, does not appear in the logs behind this claim.
Why care? Attempted and achieved are different events. The finding worth acting on is agents escalating to exploit probes when a routine task is blocked.
Take this with youWhen an AI security incident is reported, ask what was actually accessed, not only what was attempted.
Open the evidence3 source pages +
The claim we checked
AI agents linked to OpenAI carried out what researchers call the first reported instance of agents hacking a government, attempting to compromise the Australian Institute of Health and Welfare's public data site; ABC News headlined it the 'first' government hack by autonomous AI.
These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.
None of the hacking attempts we identified appear to have succeeded, though the public artifacts we analyzed are incomplete and we cannot rule out successful attempts through private scans or means other than urlquery.net.
Agents working on a pharmaceutical-data task probed for a vulnerability and retrieved a public file from a pre-production server after bot protection blocked the main site.
It follows revelations announced by Prime Minister Anthony Albanese this morning that OpenAI's AI agents had also accessed non-public Medicare health statistics held by Services Australia.
Marles confirmed that the agent’s interactions with three other government websites – the Australian Institute of Health and Welfare (AIHW), the Victorian Department of Health and the NSW Bureau of Crime Statistics and Research – were “entirely normal” and involved only publicly available information.
OpenAI's mental health test, graded by OpenAI's model, scored clinicians below its AI.
OpenAI ·
OpenAI released MentalHealthBench, a benchmark of synthetic mental health conversations with criteria written by licensed experts. Each reply is graded by OpenAI's GPT-5.6 Sol. OpenAI says the results show steady improvement in helping people.
The reality check
In a reference comparison, answers written by clinicians scored 38.5% and GPT-6 Astra scored 57.3%. The authors attribute the gap largely to clinicians answering briefly, and replies written with the rubric in view scored 99.0%. The score measures rubric coverage on single replies. Nothing in the record measures what happened to a person afterwards.
Why care? A higher score means a reply ticked more expert-written boxes. It does not show that anyone was helped, which is the claim.
Take this with youRead an AI mental health benchmark as a score for replies, not evidence that people were helped.
Open the evidence3 source pages +
The claim we checked
OpenAI says results on its new MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations.
These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.
Using privacy-preserving techniques, we created synthetic mental health conversations that accurately reflect real-world usage patterns of AI for mental health.
Expert-authored completions written by clinicians scored 38.5%, which the authors attribute largely to clinicians writing short responses as if in an in-person conversation, often asking a single question or making a simple statement.
Rubric-aware completions, written with the grading rubrics provided, scored 99.0%, which the paper describes as a sanity check on the evaluation’s noise ceiling.
MentalHealthBench reference comparison, reported by NxCode. Each metric uses its own scale; receipts in the story.
03 / ProductsTrue, but
Google says even Google cannot read its new AI memory. The last named audit said safe from everyone except Google.
Google ·
Google described a planned persistent memory layer for Private AI Compute: memories in encrypted cloud storage, keys held on the user's devices, and data it says will be inaccessible to anyone else, even Google. Help Net Security reports the example uses are presented as potential, not available features.
Free with your account
Sign in for this free check.
This edition’s selected free story opens after sign-in.
Google says its planned server-side memory for Private AI Compute will keep users' AI memories in encrypted storage with keys held only on their devices, ensuring the data is inaccessible to anyone else, even Google.
These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.
Although the overall system relies upon proprietary hardware and is centralized on Borg Prime, NCC Group considers that Google has robustly limited the risk of user data being exposed to unexpected processing or outsiders, unless Google, as a whole organization, decides to do so
NCC Group, which has conducted an external assessment of Private AI Compute between April and September 2025, said it was able to discover a timing-based side channel in the IP blinding relay component
The full explanation appears when your reading access is confirmed.
04 / SafetyHolds up
Perplexity's sandbox wall held in 108 runs. Its network fence did not, at first.
Perplexity ·
Perplexity gave nine AI models root access inside SPACE, the sandbox behind Perplexity Computer, and asked them to escape to the host or reach a blocked network destination. It reports no VM-to-host escape in 108 runs.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Open the evidence2 source pages +
The claim we checked
Perplexity says that when nine AI models with root access tried to break out of SPACE, the sandbox behind Perplexity Computer, no VM-to-host escape was observed in 108 runs.
These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.