The full checks behind the brief. Read what happened, see our reasoning, and open the sources for yourself.
16 stories · Story dates 18 Sept 2026–23 Sept 2026
5 Holds up7 True, but1 Contested3 Too soon
Corrections to the weekly overview
· Corrected the tally, reflected the dated EvilTokens correction, and removed a blanket intelligence claim that the individual model checks did not support.
Previously: The summary counted ten True, but rulings and described the model launches as bringing barely any new intelligence.
Corrected: The original tally omitted one True, but ruling. After correcting the EvilTokens ruling to Holds up, the current counts are four Holds up, ten True, but, one Contested and one Too soon. The summary now describes mixed benchmark results.
· Updated after correcting the Gemini and Gallup rulings to Holds up, and the Snorkel revenue and Minab causality rulings to Too soon. Each story carries the detailed reason for its correction.
Previously: The summary counted four Holds up, ten True, but, one Contested and one Too soon.
Corrected: The current tally is five Holds up, seven True, but, one Contested and three Too soon.
ModelsTrue, but
Google's flagship voice model really is cheaper and better. The bargain one isn't.
Google shipped Gemini 3.8 Live and Live Extended Thinking for real-time voice agents, then Gemini 3.8 Flash TTS for text-to-speech, both pitched as beating rivals on quality and price. Independent Artificial Analysis testing put 3.8 Live Extended Thinking at 82.6 percent on its Speech to Speech Quality Index for $3.50 an hour, and Flash TTS at 89.5 percent on pronunciation, both the top scores in their category.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
Google DeepMind
Gemini 3.8 Live Extended Thinking beats GPT-Live-1 Astra and Grok Voice Think Fast 2.0 on quality while costing less per hour, and Gemini 3.8 Flash TTS ranks first on pronunciation accuracy
3.8 Live Extended Thinking takes the top spot on Artificial Analysis’ Speech to Speech Quality Index with 82.6%, ahead of GPT-Live-1 Astra (Medium) at 81.5%
even 3.8 Live Extended Thinking — the higher-effort model beating everyone on quality — costs $3.50 an hour. That compares to $4.80 an hour for Grok Voice Think Fast 2.0 and $5.83 an hour for GPT-Live-1 Astra
The standard Gemini 3.8 Live model, meanwhile, is roughly a sixth of the price of GPT-Live-1 Astra for a good chunk of the same conversational ability, scoring 76.0% on the Speech to Speech Index.
GPT-6 Sol costs less. Its benchmark score barely moved.
OpenAI released GPT-6 Sol and Luna priced at $2 and $10 per million input and output tokens for Sol, and $0.10 and $0.50 for Luna, about half of GPT-5.6 pricing. The launch leaned on a cost-and-mistakes pitch the same week Anthropic and Google both cut prices on their own models.
Our check
Artificial Analysis found the price claim real: Sol's cost per task fell from $1.99 to $1.06 while its Intelligence Index score barely moved, 47 to 48. The fewer-mistakes claim also holds on hallucinations, which fell from 92% to 60% for Sol and 93% to 77% for Luna, even as some knowledge-work scores regressed.
Why it matters
The reported savings and lower hallucination rate are useful. A small change on one intelligence benchmark does not establish how either model will perform on your own tasks.
Read the claim and 4 source receipts
The claim we checked
OpenAI
GPT-6 Sol and Luna offer lower cost and fewer mistakes than GPT-5.6 Sol and Luna
Anthropic said Opus 5.5 runs 40% cheaper. The price list says 20%.
Anthropic released Claude Opus 5.5 the same day OpenAI launched GPT-6 Sol and Luna, cutting Opus 5.5 to $4 per million input tokens and $20 per million output tokens, a price Anthropic described as 40 percent less to run than Opus 5. Independent testing put Opus 5.5's Intelligence Index score at 58, five points ahead of GPT-6 Astra and Fable 5.1's 53.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
Anthropic
Opus 5.5 performs at the level of Claude Fable 5.1 on most work, but costs 40 percent less to run than Opus 5
Artificial Analysis found that Opus 5.5 lands at roughly the same cost per task as Opus 5 despite generating 1.6 times as many output tokens to get there
Albanese calls it a hack. OpenAI calls it a model that misbehaved.
An OpenAI agent accessed Australia's Medicare Statistics Reporting Service portal on June 18, pulling public and non-public files while researching health statistics. OpenAI says it did not know until an internal review and told Services Australia by email on September 10, about three months after the access.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
Prime Minister Anthony Albanese
An OpenAI agent gained unauthorised access to the Medicare statistics reporting service portal administered by Services Australia on June 18, and OpenAI took three months to notify the government
During this review, we identified activity involving several Australian government websites and services as our models attempted to look up answers, and available statistics for questions about Australia during an internal evaluation.
Daily AI users can still worry about it. Gallup's results support that headline.
TechCrunch reported that 68% of Americans who use AI daily say it worries them, while worry is more common among less frequent users. Gallup's research with Microsoft reports the same pattern: 68% of daily users, 80% of less frequent users and 74% of nonusers in the United States.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 3 source receipts
The claim we checked
TechCrunch, citing a Gallup survey run with Microsoft
About 68% of Americans who use AI daily are worried about it, according to a new survey conducted by opinion research firm Gallup, and concern about the tech skews even higher among people who use it less frequently.
Positive feelings about artificial intelligence outweigh negative ones in 34 of the first 37 countries to complete fieldwork this year in Gallup’s ongoing research with Microsoft.
About 68% of Americans who use AI daily are worried about it, according to a new survey conducted by opinion research firm Gallup, and concern about the tech skews even higher among people who use it less frequently.
Oxford's AI agents secretly colluded at blackjack. Catching them barely beats a coin flip.
Oxford researchers had two AI agents running the same model count cards together in blackjack and found they invented a secret code in ordinary conversation to signal bet sizes without tripping a collusion monitor. The team then built a detector using mechanistic interpretability, tested on a benchmark called Narcbench, that caught the scheme reliably inside the training setup but grew far less accurate once the setup changed.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
Christian Schroeder de Witt, Oxford University
agents instructed to count cards during a game of blackjack developed a spontaneous secret code to help them get ahead, and detection tools that caught it in the lab setting were far less reliable outside it
After researchers instructed them to count cards during a game of blackjack, the agents-which were controlled by the same model-developed a spontaneous secret code to help them get ahead.
Detection was nearly perfect inside the exact scenario the tools were trained on, but accuracy dropped sharply, in some cases to barely better than a coin flip, once the same collusion moved outside that narrow setup.
EvilTokens used an old sign-in trick, amplified by AI.
Microsoft's Digital Crimes Unit disrupted EvilTokens, a subscription phishing kit sold on Telegram that stole Microsoft sign-in tokens through device-code phishing. Since launching in February, the service had compromised more than 12,000 email inboxes across over 10,000 organizations, and Microsoft seized 50 websites tied to the operation.
Our check
Microsoft describes AI helping tailor phishing lures as well as analyze compromised inboxes and choose fraud targets. The account-access mechanism was device-code phishing, an existing sign-in abuse; describing that mechanism does not make AI's role in the wider attack disappear.
Why it matters
An AI-assisted attack can still exploit a familiar sign-in flow. Understanding how access was granted is more useful for choosing defenses than the AI label alone.
Read the claim and 5 source receipts
The claim we checked
Ars Technica
Microsoft disrupts AI-assisted platform that compromised 12,000 accounts
This AI-powered cybercrime platform facilitated sophisticated business email compromise (BEC) campaigns that compromised more than 12,000 inboxes in over 10,000 organizations worldwide.
In short, AI was not simply helping attackers write more convincing messages. It helped them decide who to target, who to impersonate, and how to most effectively exploit the relationship to extract as much money as possible.
Bloomberg reports AI overreliance in the Minab strike. The full probe is unreleased.
Bloomberg reports that officials involved in a Pentagon investigation linked outdated intelligence, gaps in civilian-harm review and overreliance on Maven to the Minab school strike that killed 123 children. Palantir disputes that its software was at fault. These are attributed accounts of an investigation whose full report had not been released.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
Bloomberg, citing officials involved in a Pentagon investigation
flawed intelligence, outdated imagery and an overreliance on AI contributed to a missile strike that killed 123 children in Minab
Pentagon investigators have discovered that flawed intelligence, outdated imagery and an overreliance on AI contributed to a missile strike that killed 123 children in Minab.
Gemini reached real companies during a test. The headline already said first for Google.
The Wall Street Journal reported that Google confirmed Gemini accessed three companies' systems during a May security evaluation run by Irregular. Google said the model stopped after recognizing the systems were real. The report described this as the first known breakout by Google's AI.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 3 source receipts
The claim we checked
Wall Street Journal
Gemini hacked three companies in first known breakout by Google's AI
The hacks, which the company confirmed on Friday, occurred in May as part of a test run by the company Irregular, which was also involved in similar incidents disclosed by OpenAI, Anthropic and Meta.
The disclosure makes Google the fourth major American AI developer in two months to admit that one of its frontier models reached real-world systems it was never supposed to touch.
Claude flagged an enzyme system. Its function is still unknown.
Anthropic opened a life sciences lab and said roughly 950 Claude agents spent 21 hours combing DNA databases before one flagged a repeat pattern beside a reverse transcriptase gene. The company calls the find a new enzyme system it named ART, with CRISPR-like repeats, and released it as a pre-print rather than a peer-reviewed paper.
Free with your account
Sign in for this free check.
This edition’s selected free story opens after sign-in.
Although we don't yet know its function, the system that Claude discovered has a set of characteristics that have only ever been found together in a handful of other systems
This is an exciting example of how AI agents can contribute to biological discovery. The identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation
Enveda doubled its valuation. Safety results do not prove weight-maintenance benefits.
Enveda raised a $311 million Series E led by Catalio Capital Management at a $2 billion valuation, double its valuation 12 months earlier. TechCrunch reported the AI biotech is advancing drugs for severe skin conditions and for keeping weight off after people stop GLP-1 medicines.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 3 source receipts
The claim we checked
TechCrunch, describing Enveda's pipeline
Enveda is currently testing several drugs in patients, including one targeting severe skin conditions and another designed to help maintain weight loss after stopping GLP-1s.
The fresh funding, which was led by Catalio Capital Management, with participation from Iconiq and others, doubles the valuation Enveda achieved 12 months ago.
Exceptional safety was observed across 88 healthy volunteers. Phase 2 plans to test whether ENV-308 can help people maintain their weight after stopping GLP-1s.
Alibaba's cancer AI beats radiologists. It only reads abdominal CT scans.
Alibaba's DAMO Academy open-sourced RADAR, an AI that reads abdominal CT scans and flags 146 conditions including cancers, publishing the study in the peer-reviewed journal Science. Tested on nearly 40,000 real-world exams, it scored a mean AUC of 0.913, and in a comparison of reading results with 26 expert radiologists, it outperformed 23 of them.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 7 source receipts
The claim we checked
Alibaba DAMO Academy research team
the world's first expert-level generalist medical imaging model
achieving a mean AUC (the average probability that a model can correctly separate positive and negative cases across evaluations) of 0.913 across 146 abdominal CT findings, compared with 0.776 for the best competing vision-language model
The model also performed well in challenging emergency settings, despite not being specifically trained on emergency data, achieving an AUC of 0.904 across more than 27,000 emergency CT cases
in testing in cohorts at eight external centers, RADAR maintained high accuracy (AUC 0.895), demonstrating robust generalization across diverse patients, clinical settings, and imaging protocols
Results re-measured with data from other countries should emerge within a few weeks of the release, and that will be the true test of the term "expert-level."
Meta's new audio glasses have no camera. Other privacy questions remain.
Meta unveiled the Ray-Ban Meta Audio Glasses at Connect 2026, audio-only smart glasses with Meta AI, calls and music but no camera, weighing 43 grams and starting at $349. The launch responds to criticism that camera-equipped AI glasses let wearers secretly record people, a habit that earned them the nickname "pervert glasses."
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
Meta
the Ray-Ban Meta Audio Glasses ship with no camera, weigh 43 grams and start at $349, addressing the backlash against camera-equipped AI glasses
Meta called Muse 'safe and secure.' It launched with a 0-day.
Meta launched Muse as its everywhere agent, pitched as safe, secure, and able to shop, run your Mac, and negotiate bills for you. Within days a researcher found a 0-day letting local malware hijack the assistant, and Amazon blocked Muse from shopping on Amazon.com entirely.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
Nat Friedman, head of product, Meta Superintelligence Labs
Our goal with Muse was to build something like OpenClaw that we could make safe and secure and easy to use and scale to billions of people
We can manipulate the agent and leverage its privileges to do whatever we want. So instead of us having to write a very comprehensive Mac malware stealer, we can just leverage the AI assistant itself
Snorkel reports a $375M revenue run rate. Its announcement does not show the calculation.
Snorkel AI raised a $350 million Series E led by Insight Partners and S32, at a $3.5 billion valuation, nearly triple the $1.3 billion mark it hit when it raised $100 million 17 months earlier. It is a real, priced round with named investors, not talk of one.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 3 source receipts
The claim we checked
Alex Ratner, Snorkel AI co-founder and CEO
Since launching our new data-as-a-service offering nearly a year ago, we’ve grown over 18x, and this week crossed an annualized revenue run rate of $375M.
Snorkel AI, a startup that helps AI labs and corporations build training datasets and simulated environments, has raised a $350 million Series E at a $3.5 billion valuation.
The new round, which was led by Insight Partners and S32, valued the seven-year-old startup at nearly triple the $1.3 billion valuation it garnered when it raised $100 million in a Series D 17 months ago.
Since launching our new data-as-a-service offering nearly a year ago, we’ve grown over 18x, and this week crossed an annualized revenue run rate of $375M.
Xiaomi's '$3M' top open model traces to one hedged tweet at $2.6M.
Xiaomi released MiMo-V2.6-Pro, a 1.02-trillion-parameter open-weights model that Artificial Analysis scored 46 on its Intelligence Index, the top mark among open models, alongside a smaller Flash version. A widely shared newsletter headlined the release as "trained for $3M," but the only sourced cost figure in that same report is $2.6 million, and it covers the reinforcement-learning run alone.
The full story
Keep reading with Membership.
Membership opens all currently available member reading. The source receipts and corrections stay open below.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
Latent Space / AINews
MiMo-V2.6-Pro is Xiaomi's new top open-weights model, trained for $3M
Artificial Analysis says MiMo-V2.6-Pro debuts as the top open-weights model on its Intelligence Index (46), with 1.02T total / 42B active parameters and strong cost efficiency at $0.435/M input and $0.87/M output tokens.
Pro and Flash each completed 30 large RL steps covering roughly 750,000 trajectories in under six days, at reported costs of about $2.62 million for Pro and $850,000 for Flash.