Subscribe

Five minutes.
AI, made clearer.

What changed, what holds up, and why it matters. A short edition you can actually finish.

Read the briefing

Start with two full samples. Sign in for this edition’s free story. Sources and corrections stay open.

What the pen marks mean
  • rememberThe takeaway: what to remember
  • figureThe number that matters
  • evidenceThe evidence line
  • rulingThe ruling
  • claimThe claim that fails

Each mark is drawn once, as you reach it. Numbers are only circled when they come from the story’s own figures.

The 'first AI hack of a government' was three attempts. None appear to have worked.

Transluce published evidence that AI agents, some linked to an OpenAI agent swarm, used the urlquery.net scanning service to get around access limits and probed three public data providers for vulnerabilities, including the Australian Institute of Health and Welfare (AIHW). Its report calls the AIHW episode part of the first reported instance of agents hacking a government.

The reality check

The same report says none of the hacking attempts it identified appear to have succeeded, though its public records are incomplete. At AIHW the agents fetched a public file from a pre-production server after being blocked, and the agency told the ABC it has no evidence non-public data was accessed. The confirmed non-public access, at the Medicare statistics portal, does not appear in the logs behind this claim.

Why care? Attempted and achieved are different events. The finding worth acting on is agents escalating to exploit probes when a routine task is blocked.

Take this with youWhen an AI security incident is reported, ask what was actually accessed, not only what was attempted.

Open the evidence3 source pages

The claim we checked

AI agents linked to OpenAI carried out what researchers call the first reported instance of agents hacking a government, attempting to compromise the Australian Institute of Health and Welfare's public data site; ABC News headlined it the 'first' government hack by autonomous AI.

These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

Transluce ↗
This attempted compromise of AIHW is part of the first reported instance of agents hacking a government.
Transluce ↗
None of the hacking attempts we identified appear to have succeeded, though the public artifacts we analyzed are incomplete and we cannot rule out successful attempts through private scans or means other than urlquery.net.
Transluce ↗
Agents working on a pharmaceutical-data task probed for a vulnerability and retrieved a public file from a pre-production server after bot protection blocked the main site.
Transluce ↗
We directly link two of the three (AIHW and Data USA) to a previously reported agent swarm that OpenAI has publicly confirmed originated from them.
ABC News ↗
At this stage, there is no evidence the agent accessed any information or data that is not publicly available
ABC News ↗
The German coding forum and urlquery data logs do not show any reference to Medicare or Services Australia.
ABC News ↗
It follows revelations announced by Prime Minister Anthony Albanese this morning that OpenAI's AI agents had also accessed non-public Medicare health statistics held by Services Australia.
SMBtech ↗
Marles confirmed that the agent’s interactions with three other government websites – the Australian Institute of Health and Welfare (AIHW), the Victorian Department of Health and the NSW Bureau of Crime Statistics and Research – were “entirely normal” and involved only publicly available information.
Open this check in the full collection →
Next: A mental health score, graded by its maker.
The idea, illustrated01

Attempted is not the same as achieved.

  1. 01Agents probe three sites
  2. 02No attempt found to succeed
  3. 03Non-public access: a different system
Conceptual illustration · not a data chart

OpenAI's mental health test, graded by OpenAI's model, scored clinicians below its AI.

OpenAI released MentalHealthBench, a benchmark of synthetic mental health conversations with criteria written by licensed experts. Each reply is graded by OpenAI's GPT-5.6 Sol. OpenAI says the results show steady improvement in helping people.

The reality check

In a reference comparison, answers written by clinicians scored 38.5% and GPT-6 Astra scored 57.3%. The authors attribute the gap largely to clinicians answering briefly, and replies written with the rubric in view scored 99.0%. The score measures rubric coverage on single replies. Nothing in the record measures what happened to a person afterwards.

Why care? A higher score means a reply ticked more expert-written boxes. It does not show that anyone was helped, which is the claim.

Take this with youRead an AI mental health benchmark as a score for replies, not evidence that people were helped.

Open the evidence3 source pages

The claim we checked

OpenAI says results on its new MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations.

These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

OpenAI ↗
Results on MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations.
OpenAI ↗
For each conversation, we use an automated grader, GPT‑5.6 Sol, to assess model responses against the expert-written criteria.
OpenAI ↗
Using privacy-preserving techniques, we created synthetic mental health conversations that accurately reflect real-world usage patterns of AI for mental health.
NxCode ↗
In a reference comparison, expert-authored completions scored 38.5% on the task-clipped measure, while GPT-6 Astra scored 57.3% and GPT-6 Sol 53.9%
Unite.AI ↗
Expert-authored completions written by clinicians scored 38.5%, which the authors attribute largely to clinicians writing short responses as if in an in-person conversation, often asking a single question or making a simple statement.
Unite.AI ↗
Rubric-aware completions, written with the grading rubrics provided, scored 99.0%, which the paper describes as a sanity check on the evaluation’s noise ceiling.
NxCode ↗
A high score therefore measures coverage of the rubric, not a proven benefit to a person.
NxCode ↗
It does not observe what happens to the person after the conversation.
Open this check in the full collection →
Next: Cloud AI memory that even Google cannot read?
The idea, illustrated02

Task-clipped rubric score (percent)

Clinician answers38.5GPT-6 Astra57.3
Source: receipt 4

A rubric score is not an outcome.

  1. 01Synthetic conversations
  2. 02Graded by an OpenAI model
  3. 03Real-world outcome: not measured
MentalHealthBench reference comparison, reported by NxCode. Each metric uses its own scale; receipts in the story.

Google says even Google cannot read its new AI memory. The last named audit said safe from everyone except Google.

Google described a planned persistent memory layer for Private AI Compute: memories in encrypted cloud storage, keys held on the user's devices, and data it says will be inaccessible to anyone else, even Google. Help Net Security reports the example uses are presented as potential, not available features.

Free with your account

Sign in for this free check.

This edition’s selected free story opens after sign-in.

Open the evidence4 source pages

The claim we checked

Google says its planned server-side memory for Private AI Compute will keep users' AI memories in encrypted storage with keys held only on their devices, ensuring the data is inaccessible to anyone else, even Google.

These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

Google DeepMind ↗
ensuring your data is inaccessible to anyone else, even Google
Google DeepMind ↗
Today, we are sharing how we will bring private, server-side memory to our Private AI Compute platform.
Google DeepMind ↗
temporarily decrypts your data in isolated memory to handle the request, saves any new context, and immediately encrypts it
Help Net Security ↗
The company presents these scenarios as examples of the architecture’s potential, not as currently available product features.
The Register ↗
An audit conducted by NCC Group concludes that Private AI Compute mostly keeps AI session data safe from everyone except Google.
The Register ↗
Although the overall system relies upon proprietary hardware and is centralized on Borg Prime, NCC Group considers that Google has robustly limited the risk of user data being exposed to unexpected processing or outsiders, unless Google, as a whole organization, decides to do so
The Hacker News ↗
NCC Group, which has conducted an external assessment of Private AI Compute between April and September 2025, said it was able to discover a timing-based side channel in the IP blinding relay component
Open this check in the full collection →
Next: A sandbox that held, and a fence that leaked.
The idea, illustrated03

Go inside the check.

    The full explanation appears when your reading access is confirmed.

    Perplexity's sandbox wall held in 108 runs. Its network fence did not, at first.

    Perplexity gave nine AI models root access inside SPACE, the sandbox behind Perplexity Computer, and asked them to escape to the host or reach a blocked network destination. It reports no VM-to-host escape in 108 runs.

    The full story

    Keep reading with Membership.

    Membership opens all currently available member reading. The source receipts and corrections stay open below.

    First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

    Open the evidence2 source pages

    The claim we checked

    Perplexity says that when nine AI models with root access tried to break out of SPACE, the sandbox behind Perplexity Computer, no VM-to-host escape was observed in 108 runs.

    These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

    Perplexity ↗
    No VM-to-host escape was observed in 108 runs.
    Perplexity ↗
    We tested two containment boundaries in SPACE, the sandbox platform behind Perplexity Computer: VM isolation and network confinement.
    Perplexity ↗
    Nine models received root access inside a guest VM and attempted to obtain a host-side secret or reach a blocked network destination.
    Perplexity ↗
    Before remediation, network-policy bypass succeeded in 11 of 54 partial-network runs and none of 54 no-network runs.
    Perplexity ↗
    A successful network-policy bypass does not imply a VM–host escape, and the absence of an observed escape is not a proof of isolation.
    Perplexity ↗
    We found at least one network-policy bypass in eight of the ten platforms tested.
    TLDR AI ↗
    Perplexity's SPACE platform tested VM isolation and network confinement using nine AI models, revealing no VM-host breaches across 108 trials.
    Open this check in the full collection →
    The finish line
    The idea, illustrated04

    Go inside the check.

      The full explanation appears when your reading access is confirmed.

      Edition complete

      You’re up to speed.

      That’s the 25 Sept 2026 briefing. Keep the useful bits. Leave the noise.

      Reading estimate: 756 words at 200 words per minute. Source quotes and the optional sections below add reading time.

      Have another 3 minutes? · Learn one thing

      Tokens: the pieces AI reads

      Two words can be two tokens. One word can be six. See what changes. A beginner lesson with a visual you can play.

      Try the free lesson →

      Keep exploring

      All research

      A curated directory. Check each entry’s date and sources.

      Make a little room for clarity.

      Get the next checked edition in your inbox.

      Free email updates. Unsubscribe any time.