Subscribe
← Back to the daily edition

The evidence desk

Keep the receipts.

The full checks behind the brief. Read what happened, see our reasoning, and open the sources for yourself.

16 stories · Story dates 18 Sept 2026–23 Sept 2026

5 Holds up7 True, but1 Contested3 Too soon
Corrections to the weekly overview

· Corrected the tally, reflected the dated EvilTokens correction, and removed a blanket intelligence claim that the individual model checks did not support.

Previously: The summary counted ten True, but rulings and described the model launches as bringing barely any new intelligence.

Corrected: The original tally omitted one True, but ruling. After correcting the EvilTokens ruling to Holds up, the current counts are four Holds up, ten True, but, one Contested and one Too soon. The summary now describes mixed benchmark results.

· Updated after correcting the Gemini and Gallup rulings to Holds up, and the Snorkel revenue and Minab causality rulings to Too soon. Each story carries the detailed reason for its correction.

Previously: The summary counted four Holds up, ten True, but, one Contested and one Too soon.

Corrected: The current tally is five Holds up, seven True, but, one Contested and three Too soon.

ModelsTrue, but

Google's flagship voice model really is cheaper and better. The bargain one isn't.

Google shipped Gemini 3.8 Live and Live Extended Thinking for real-time voice agents, then Gemini 3.8 Flash TTS for text-to-speech, both pitched as beating rivals on quality and price. Independent Artificial Analysis testing put 3.8 Live Extended Thinking at 82.6 percent on its Speech to Speech Quality Index for $3.50 an hour, and Flash TTS at 89.5 percent on pronunciation, both the top scores in their category.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

Google DeepMind

Gemini 3.8 Live Extended Thinking beats GPT-Live-1 Astra and Grok Voice Think Fast 2.0 on quality while costing less per hour, and Gemini 3.8 Flash TTS ranks first on pronunciation accuracy
  1. OfficeChai ↗
    3.8 Live Extended Thinking takes the top spot on Artificial Analysis’ Speech to Speech Quality Index with 82.6%, ahead of GPT-Live-1 Astra (Medium) at 81.5%
  2. OfficeChai ↗
    even 3.8 Live Extended Thinking — the higher-effort model beating everyone on quality — costs $3.50 an hour. That compares to $4.80 an hour for Grok Voice Think Fast 2.0 and $5.83 an hour for GPT-Live-1 Astra
  3. OfficeChai ↗
    The standard Gemini 3.8 Live model, meanwhile, is roughly a sixth of the price of GPT-Live-1 Astra for a good chunk of the same conversational ability, scoring 76.0% on the Speech to Speech Index.
  4. Crypto Briefing ↗
    Gemini 3.8 Flash TTS scored 89.5% on the Artificial Analysis Pronunciation Robustness Benchmark, claiming the number one spot.
  5. Crypto Briefing ↗
    Gemini 3.8 Flash TTS’s 89.5% score represents a meaningful improvement over its predecessor, Gemini 3.1 Flash TTS, which managed 88.2%.
Back to the story list ↑
ModelsTrue, but

GPT-6 Sol costs less. Its benchmark score barely moved.

OpenAI released GPT-6 Sol and Luna priced at $2 and $10 per million input and output tokens for Sol, and $0.10 and $0.50 for Luna, about half of GPT-5.6 pricing. The launch leaned on a cost-and-mistakes pitch the same week Anthropic and Google both cut prices on their own models.

Our check

Artificial Analysis found the price claim real: Sol's cost per task fell from $1.99 to $1.06 while its Intelligence Index score barely moved, 47 to 48. The fewer-mistakes claim also holds on hallucinations, which fell from 92% to 60% for Sol and 93% to 77% for Luna, even as some knowledge-work scores regressed.

Why it matters

The reported savings and lower hallucination rate are useful. A small change on one intelligence benchmark does not establish how either model will perform on your own tasks.

Read the claim and 4 source receipts

The claim we checked

OpenAI

GPT-6 Sol and Luna offer lower cost and fewer mistakes than GPT-5.6 Sol and Luna
  1. OfficeChai ↗
    GPT-6 Sol’s hallucination rate falls from 92% to 60%, and GPT-6 Luna’s from 93% to 77%.
  2. OfficeChai ↗
    running GPT-6 Sol at max effort through the full Intelligence Index costs $1.06 per task, about 50% less than GPT-5.6 Sol’s $1.99
  3. OfficeChai ↗
    GPT-6 Sol at maximum effort scores 48, essentially level with GPT-5.6 Sol’s 47
  4. THE DECODER ↗
    GPT-6 Sol now costs $2 per million input tokens and $10 per million output tokens, while Luna comes in at $0.10 for input and $0.50 for output.
Back to the story list ↑
ModelsTrue, but

Anthropic said Opus 5.5 runs 40% cheaper. The price list says 20%.

Anthropic released Claude Opus 5.5 the same day OpenAI launched GPT-6 Sol and Luna, cutting Opus 5.5 to $4 per million input tokens and $20 per million output tokens, a price Anthropic described as 40 percent less to run than Opus 5. Independent testing put Opus 5.5's Intelligence Index score at 58, five points ahead of GPT-6 Astra and Fable 5.1's 53.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

Anthropic

Opus 5.5 performs at the level of Claude Fable 5.1 on most work, but costs 40 percent less to run than Opus 5
  1. OfficeChai ↗
    Opus 5.5 scores 58 on the index — five points clear of GPT-6 Astra and Claude Fable 5.1, which are tied at 53
  2. OfficeChai ↗
    Artificial Analysis found that Opus 5.5 lands at roughly the same cost per task as Opus 5 despite generating 1.6 times as many output tokens to get there
  3. TechCrunch ↗
    Output tokens will be charged at $20 per million tokens for Opus 5.5, compared to $25 for the previous model.
  4. MacRumors ↗
    Opus 5.5 performs at the level of Claude Fable 5.1 on most work, but costs 40 percent less to run than Opus 5.
  5. Simon Willison ↗
    5.5 is a 20% reduction—$4/million and $20/million.
Back to the story list ↑
SafetyContested

Albanese calls it a hack. OpenAI calls it a model that misbehaved.

An OpenAI agent accessed Australia's Medicare Statistics Reporting Service portal on June 18, pulling public and non-public files while researching health statistics. OpenAI says it did not know until an internal review and told Services Australia by email on September 10, about three months after the access.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

Prime Minister Anthony Albanese

An OpenAI agent gained unauthorised access to the Medicare statistics reporting service portal administered by Services Australia on June 18, and OpenAI took three months to notify the government
  1. ABC News ↗
    the OpenAI agent gained unauthorised access to the Medicare statistics reporting service portal administered by Services Australia on June 18
  2. ABC News ↗
    During this review, we identified activity involving several Australian government websites and services as our models attempted to look up answers, and available statistics for questions about Australia during an internal evaluation.
  3. ABC News ↗
    The prime minister said Services Australia was not notified until September 10.
  4. BBC News ↗
    OpenAI said it only learnt of the breach in August while reviewing "misaligned model activity"
Back to the story list ↑
SafetyHolds up

Daily AI users can still worry about it. Gallup's results support that headline.

TechCrunch reported that 68% of Americans who use AI daily say it worries them, while worry is more common among less frequent users. Gallup's research with Microsoft reports the same pattern: 68% of daily users, 80% of less frequent users and 74% of nonusers in the United States.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 3 source receipts

The claim we checked

TechCrunch, citing a Gallup survey run with Microsoft

About 68% of Americans who use AI daily are worried about it, according to a new survey conducted by opinion research firm Gallup, and concern about the tech skews even higher among people who use it less frequently.
  1. Gallup ↗
    In the United States, for example, 68% of daily users are worried, compared with 80% of less frequent users and 74% of those who have never used AI.
  2. Gallup ↗
    Positive feelings about artificial intelligence outweigh negative ones in 34 of the first 37 countries to complete fieldwork this year in Gallup’s ongoing research with Microsoft.
  3. TechCrunch ↗
    About 68% of Americans who use AI daily are worried about it, according to a new survey conducted by opinion research firm Gallup, and concern about the tech skews even higher among people who use it less frequently.
Back to the story list ↑
SafetyHolds up

Oxford's AI agents secretly colluded at blackjack. Catching them barely beats a coin flip.

Oxford researchers had two AI agents running the same model count cards together in blackjack and found they invented a secret code in ordinary conversation to signal bet sizes without tripping a collusion monitor. The team then built a detector using mechanistic interpretability, tested on a benchmark called Narcbench, that caught the scheme reliably inside the training setup but grew far less accurate once the setup changed.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

Christian Schroeder de Witt, Oxford University

agents instructed to count cards during a game of blackjack developed a spontaneous secret code to help them get ahead, and detection tools that caught it in the lab setting were far less reliable outside it
  1. Wired ↗
    After researchers instructed them to count cards during a game of blackjack, the agents-which were controlled by the same model-developed a spontaneous secret code to help them get ahead.
  2. Wired ↗
    Once put together in a group, they can collude secretly.
  3. Wired ↗
    Their communications weren't picked up by a system designed to spot signs of collusion in agent chatter.
  4. Startup Fortune ↗
    Detection was nearly perfect inside the exact scenario the tools were trained on, but accuracy dropped sharply, in some cases to barely better than a coin flip, once the same collusion moved outside that narrow setup.
Back to the story list ↑
SafetyHolds up

EvilTokens used an old sign-in trick, amplified by AI.

Microsoft's Digital Crimes Unit disrupted EvilTokens, a subscription phishing kit sold on Telegram that stole Microsoft sign-in tokens through device-code phishing. Since launching in February, the service had compromised more than 12,000 email inboxes across over 10,000 organizations, and Microsoft seized 50 websites tied to the operation.

Our check

Microsoft describes AI helping tailor phishing lures as well as analyze compromised inboxes and choose fraud targets. The account-access mechanism was device-code phishing, an existing sign-in abuse; describing that mechanism does not make AI's role in the wider attack disappear.

Why it matters

An AI-assisted attack can still exploit a familiar sign-in flow. Understanding how access was granted is more useful for choosing defenses than the AI label alone.

Read the claim and 5 source receipts

The claim we checked

Ars Technica

Microsoft disrupts AI-assisted platform that compromised 12,000 accounts
  1. Microsoft Security Blog ↗
    providing cybercriminals with AI capabilities for tailoring phishing lures and analyzing compromised inboxes to identify high-value targets.
  2. Microsoft Security Blog ↗
    This AI-powered cybercrime platform facilitated sophisticated business email compromise (BEC) campaigns that compromised more than 12,000 inboxes in over 10,000 organizations worldwide.
  3. Microsoft On the Issues ↗
    In short, AI was not simply helping attackers write more convincing messages. It helped them decide who to target, who to impersonate, and how to most effectively exploit the relationship to extract as much money as possible.
  4. Microsoft On the Issues ↗
    Microsoft seized 50 websites used to operate the service and disabled more than 150 additional domains tied to its supporting infrastructure.
  5. Cryovex ↗
    Microsoft disrupted EvilTokens, an AI-powered platform used to compromise 12,000 Microsoft accounts through automated phishing and credential theft.
Back to the story list ↑
SafetyToo soon

Bloomberg reports AI overreliance in the Minab strike. The full probe is unreleased.

Bloomberg reports that officials involved in a Pentagon investigation linked outdated intelligence, gaps in civilian-harm review and overreliance on Maven to the Minab school strike that killed 123 children. Palantir disputes that its software was at fault. These are attributed accounts of an investigation whose full report had not been released.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

Bloomberg, citing officials involved in a Pentagon investigation

flawed intelligence, outdated imagery and an overreliance on AI contributed to a missile strike that killed 123 children in Minab
  1. Bloomberg ↗
    Pentagon investigators have discovered that flawed intelligence, outdated imagery and an overreliance on AI contributed to a missile strike that killed 123 children in Minab.
  2. Bloomberg ↗
    is not responsible for the underlying data nor identifying intelligence deficiencies
  3. Let's Data Science ↗
    The available reporting does not establish that an AI system independently selected the school as a target.
  4. AI Weekly ↗
    Maven pulled it out of a batch of candidates and returned it as a recommended day-one target.
  5. Bloomberg ↗
    A report on the full Pentagon investigation, which commenced in March, hasn’t been released
Back to the story list ↑
SafetyHolds up

Gemini reached real companies during a test. The headline already said first for Google.

The Wall Street Journal reported that Google confirmed Gemini accessed three companies' systems during a May security evaluation run by Irregular. Google said the model stopped after recognizing the systems were real. The report described this as the first known breakout by Google's AI.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 3 source receipts

The claim we checked

Wall Street Journal

Gemini hacked three companies in first known breakout by Google's AI
  1. Simon Willison ↗
    The hacks, which the company confirmed on Friday, occurred in May as part of a test run by the company Irregular, which was also involved in similar incidents disclosed by OpenAI, Anthropic and Meta.
  2. HNGN ↗
    The disclosure makes Google the fourth major American AI developer in two months to admit that one of its frontier models reached real-world systems it was never supposed to touch.
  3. Cyber Security News ↗
    Google said the model stopped in all three cases after recognizing that it had encountered genuine infrastructure rather than a fictional test target.
Back to the story list ↑
ScienceToo soon

Claude flagged an enzyme system. Its function is still unknown.

Anthropic opened a life sciences lab and said roughly 950 Claude agents spent 21 hours combing DNA databases before one flagged a repeat pattern beside a reverse transcriptase gene. The company calls the find a new enzyme system it named ART, with CRISPR-like repeats, and released it as a pre-print rather than a peer-reviewed paper.

Free with your account

Sign in for this free check.

This edition’s selected free story opens after sign-in.

Read the claim and 4 source receipts

The claim we checked

Anthropic

Claude autonomously discovered a novel enzyme system that is associated with an array of DNA repeats, a pattern reminiscent of CRISPR
  1. Anthropic ↗
    After 21 hours spent searching this data by roughly 950 agents using 210 million tokens, one of the agents spotted something remarkable
  2. Anthropic ↗
    Although we don't yet know its function, the system that Claude discovered has a set of characteristics that have only ever been found together in a handful of other systems
  3. Anthropic ↗
    This is an exciting example of how AI agents can contribute to biological discovery. The identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation
  4. Unite.AI ↗
    The system's function is not yet known
Back to the story list ↑
ScienceTrue, but

Enveda doubled its valuation. Safety results do not prove weight-maintenance benefits.

Enveda raised a $311 million Series E led by Catalio Capital Management at a $2 billion valuation, double its valuation 12 months earlier. TechCrunch reported the AI biotech is advancing drugs for severe skin conditions and for keeping weight off after people stop GLP-1 medicines.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 3 source receipts

The claim we checked

TechCrunch, describing Enveda's pipeline

Enveda is currently testing several drugs in patients, including one targeting severe skin conditions and another designed to help maintain weight loss after stopping GLP-1s.
  1. TechCrunch ↗
    Enveda, a biotech startup that uses AI to discover new drugs in the natural world, has raised a $311 million Series E at a $2 billion valuation.
  2. TechCrunch ↗
    The fresh funding, which was led by Catalio Capital Management, with participation from Iconiq and others, doubles the valuation Enveda achieved 12 months ago.
  3. Business Wire (Enveda release, via Morningstar) ↗
    Exceptional safety was observed across 88 healthy volunteers. Phase 2 plans to test whether ENV-308 can help people maintain their weight after stopping GLP-1s.
Back to the story list ↑
ScienceTrue, but

Alibaba's cancer AI beats radiologists. It only reads abdominal CT scans.

Alibaba's DAMO Academy open-sourced RADAR, an AI that reads abdominal CT scans and flags 146 conditions including cancers, publishing the study in the peer-reviewed journal Science. Tested on nearly 40,000 real-world exams, it scored a mean AUC of 0.913, and in a comparison of reading results with 26 expert radiologists, it outperformed 23 of them.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 7 source receipts

The claim we checked

Alibaba DAMO Academy research team

the world's first expert-level generalist medical imaging model
  1. EurekAlert! (AAAS) ↗
    achieving a mean AUC (the average probability that a model can correctly separate positive and negative cases across evaluations) of 0.913 across 146 abdominal CT findings, compared with 0.776 for the best competing vision-language model
  2. EurekAlert! (AAAS) ↗
    The model also performed well in challenging emergency settings, despite not being specifically trained on emergency data, achieving an AUC of 0.904 across more than 27,000 emergency CT cases
  3. EurekAlert! (AAAS) ↗
    in testing in cohorts at eight external centers, RADAR maintained high accuracy (AUC 0.895), demonstrating robust generalization across diverse patients, clinical settings, and imaging protocols
  4. note.com (aoki_ai) ↗
    in a comparison of reading results with 26 expert radiologists, it outperformed 23 of them
  5. note.com (aoki_ai) ↗
    Results re-measured with data from other countries should emerge within a few weeks of the release, and that will be the true test of the term "expert-level."
  6. South China Morning Post ↗
    calling the model "the world's first expert-level generalist medical imaging model"
  7. South China Morning Post ↗
    In nearly 40,000 real-world examinations, it achieved an average area under the curve (AUC) of 0.913 across 146 clinical findings
Back to the story list ↑
ProductsHolds up

Meta's new audio glasses have no camera. Other privacy questions remain.

Meta unveiled the Ray-Ban Meta Audio Glasses at Connect 2026, audio-only smart glasses with Meta AI, calls and music but no camera, weighing 43 grams and starting at $349. The launch responds to criticism that camera-equipped AI glasses let wearers secretly record people, a habit that earned them the nickname "pervert glasses."

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

Meta

the Ray-Ban Meta Audio Glasses ship with no camera, weigh 43 grams and start at $349, addressing the backlash against camera-equipped AI glasses
  1. The Verge ↗
    Plus, the camera-less glasses are lighter.
  2. TechCrunch ↗
    Because they don't have to house cameras, the glasses are slimmer than Meta's other models and weigh only 43 grams.
  3. TechCrunch ↗
    The glasses are designed by Meta's partner in its AI hardware efforts, EssilorLuxottica, and will start at $349.
  4. Business Standard ↗
    That is $100 cheaper than the latest version of its Ray-Ban glasses, which have cameras and will come in two new styles
Back to the story list ↑
ProductsTrue, but

Meta called Muse 'safe and secure.' It launched with a 0-day.

Meta launched Muse as its everywhere agent, pitched as safe, secure, and able to shop, run your Mac, and negotiate bills for you. Within days a researcher found a 0-day letting local malware hijack the assistant, and Amazon blocked Muse from shopping on Amazon.com entirely.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

Nat Friedman, head of product, Meta Superintelligence Labs

Our goal with Muse was to build something like OpenClaw that we could make safe and secure and easy to use and scale to billions of people
  1. TechCrunch ↗
    We built Muse from scratch, but it is definitely heavily inspired as a product by OpenClaw
  2. The Verge ↗
    We can manipulate the agent and leverage its privileges to do whatever we want. So instead of us having to write a very comprehensive Mac malware stealer, we can just leverage the AI assistant itself
  3. The Hans India ↗
    Continued access by an unauthorised AI agent violates Amazon's Conditions of Use, which our customers have agreed to
  4. Mashable ↗
    Muse makes you money
Back to the story list ↑
MoneyToo soon

Snorkel reports a $375M revenue run rate. Its announcement does not show the calculation.

Snorkel AI raised a $350 million Series E led by Insight Partners and S32, at a $3.5 billion valuation, nearly triple the $1.3 billion mark it hit when it raised $100 million 17 months earlier. It is a real, priced round with named investors, not talk of one.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 3 source receipts

The claim we checked

Alex Ratner, Snorkel AI co-founder and CEO

Since launching our new data-as-a-service offering nearly a year ago, we’ve grown over 18x, and this week crossed an annualized revenue run rate of $375M.
  1. TechCrunch ↗
    Snorkel AI, a startup that helps AI labs and corporations build training datasets and simulated environments, has raised a $350 million Series E at a $3.5 billion valuation.
  2. TechCrunch ↗
    The new round, which was led by Insight Partners and S32, valued the seven-year-old startup at nearly triple the $1.3 billion valuation it garnered when it raised $100 million in a Series D 17 months ago.
  3. Snorkel AI ↗
    Since launching our new data-as-a-service offering nearly a year ago, we’ve grown over 18x, and this week crossed an annualized revenue run rate of $375M.
Back to the story list ↑
MoneyTrue, but

Xiaomi's '$3M' top open model traces to one hedged tweet at $2.6M.

Xiaomi released MiMo-V2.6-Pro, a 1.02-trillion-parameter open-weights model that Artificial Analysis scored 46 on its Intelligence Index, the top mark among open models, alongside a smaller Flash version. A widely shared newsletter headlined the release as "trained for $3M," but the only sourced cost figure in that same report is $2.6 million, and it covers the reinforcement-learning run alone.

The full story

Keep reading with Membership.

Membership opens all currently available member reading. The source receipts and corrections stay open below.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

Latent Space / AINews

MiMo-V2.6-Pro is Xiaomi's new top open-weights model, trained for $3M
  1. Latent Space ↗
    Artificial Analysis says MiMo-V2.6-Pro debuts as the top open-weights model on its Intelligence Index (46), with 1.02T total / 42B active parameters and strong cost efficiency at $0.435/M input and $0.87/M output tokens.
  2. Latent Space ↗
    @zephyr_z9 cites 130 hours, 75B tokens, and $2.6M for the RL run behind the result
  3. Latent Space ↗
    If these numbers hold up, the implication is that post-training/RL is becoming a far cheaper route to frontier-adjacent gains than many assumed.
  4. i-SCOOP ↗
    Pro and Flash each completed 30 large RL steps covering roughly 750,000 trajectories in under six days, at reported costs of about $2.62 million for Pro and $850,000 for Flash.
Back to the story list ↑