Start free trial

Five minutes.
AI, made clearer.

What changed, what holds up, and why it matters. A short edition you can actually finish.

Read the briefing

Start with two full samples. Sign in for this edition’s free story. Sources and corrections stay open.

What the pen marks mean
  • rememberThe takeaway: what to remember
  • figureThe number that matters
  • evidenceThe evidence line
  • rulingThe ruling
  • claimThe claim that fails

Each mark is drawn once, as you reach it. Numbers are only circled when they come from the story’s own figures.

OpenAI dropped hundreds of AI math papers at once. Its own README says about 42% of the headline results are machine checked.

OpenAI says it is releasing a broad range of new math results from an internal model, with many proofs checked in Lean.

The reality check

Its own README says about 42% of top-line results are formalized and some unformalized results could have issues.

Why care? Until a proof is checked, it is a claim. By OpenAI's own count, most of the headline results have not been machine checked yet.

Take this with youOpenAI's math release is real work, but its own README says only about 42% of the top-line results are Lean checked, picked from about 4,000 problems posed.

Open the evidence4 source pages

The claim we checked

OpenAI says it is releasing a broad range of new mathematical results produced by an internal frontier model, with many proofs formalized in Lean.

These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

web.archive.org ↗
a broad range of new mathematical results produced by an internal frontier model
web.archive.org ↗
we are sharing formalizations of many of the proofs in Lean
web.archive.org ↗
The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking.
github.com ↗
The repository has ~42% top-line results formalized.
github.com ↗
Some of the unformalized results could have issues.
github.com ↗
Over the course of the evaluation, the model was posed approximately 4,000 problems.
github.com ↗
A family groups related papers, which may include a principal result, companion arguments, consequences, or alternative proofs.
github.com ↗
The current catalogue contains 719 manuscripts organized into 372 families.
github.com ↗
the writeup for the Re(s) > 11/12 zero-free region for the Riemann zeta function was human edited for readability.
terrytao.wordpress.com ↗
This is a guest post by the Association for Human Mathematics
terrytao.wordpress.com ↗
Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power.
terrytao.wordpress.com ↗
Mathematicians did not ask for this work to be done.
fortune.com ↗
OpenAI published AI-generated full or partial solutions Tuesday to more than 370 outstanding mathematical problems
fortune.com ↗
made progress on three other Millennium Prize problems but had not fully solved them
fortune.com ↗
It followed some, but not all, of the steps the advisory group had recommended.
fortune.com ↗
do not understand the AI output well enough to answer questions on the result
fortune.com ↗
My view is that this is great for mathematics
Open this check in the full collection →
The idea, illustrated01

OPENAI: Hundreds of proofs, few checked

  1. 01OPENAI: Hundreds of proofs, few checked
  2. 02What the record shows
  3. 03OpenAI math drop: 42% machine checked
Conceptual illustration · not a data chart
Next: Staff seats cut, not Claude demand

Meta and Microsoft are kicking the Claude habit, says the headline. What they cut is staff seats. Claude stays inside Microsoft's own tool.

The Information reports Meta and Microsoft are pushing staff off Claude.

The reality check

Meta's internal Claude Code users fell from about 60,000 to about 30,000, and spring layoffs explain part of that. Microsoft cut an internal projection of at least $1 billion by more than a third.

Why care? A big buyer swapping staff onto its own tools is a cost and product decision, not a verdict on the model. The customer numbers point the other way.

Take this with youMeta and Microsoft cut internal staff seats for Claude Code, partly via layoffs; Claude stays inside Copilot CLI and customer spend on Claude via Azure keeps growing.

Open the evidence4 source pages

The claim we checked

Meta and Microsoft are pushing staff off Anthropic's Claude, with Meta's internal Claude Code users halved and Microsoft's planned Anthropic spend slashed.

These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

aiweekly.co ↗
Claude Code use inside Meta has dropped to roughly 30,000 employees, down from about 60,000 earlier this year
aiweekly.co ↗
The company projected spending at least $1 billion on Anthropic
aiweekly.co ↗
has since been slashed by more than a third
aiweekly.co ↗
Spring layoffs that hit about 10% of Meta
aiweekly.co ↗
account for part of the decline, but not the bulk of it
aiweekly.co ↗
The internal pullback is narrower than the headline numbers imply.
aiweekly.co ↗
customer spending on Claude via Azure and Bedrock continues to grow
letsdatascience.com ↗
Neither company nor Anthropic has publicly announced the reported changes
letsdatascience.com ↗
The $1 billion figure is described as an earlier internal projection, not audited spending.
letsdatascience.com ↗
both to control costs and to use more of Microsoft
dev.ua ↗
users have until June 30, 2026, to completely remove Claude Code from their workflows
dev.ua ↗
Claude models will remain available through the Copilot CLI
dev.ua ↗
Claude Code was a critical part of this learning curve
shacknews.com ↗
Previous models of Claude are still available to Microsoft employees
shacknews.com ↗
This data is then deleted after 30 days.
dev.ua ↗
explained Executive Vice President Rajesh Jha
Open this check in the full collection →
The idea, illustrated02

META AND MICROSOFT: Seats cut, model kept

  1. 01META AND MICROSOFT: Seats cut, model kept
  2. 02What the record shows
  3. 03Staff seats cut, not Claude demand
Conceptual illustration · not a data chart
Next: Proof plan, law not until 2027

Headlines say Australia will make AI companies prove their safety systems work. The source is a speech, and the law is planned for 2027.

Australia will make AI companies prove their safety systems work.

Free with your account

Sign in for this free check.

This edition’s selected free story opens after sign-in.

Open the evidence3 source pages

The claim we checked

Australia will make AI companies prove their safety systems work, under proposed laws modelled on banking and aviation regulation.

These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

minister.industry.gov.au ↗
Today I want to outline how the Albanese Government is approaching the question of frontier AI regulation
minister.industry.gov.au ↗
and then holding them accountable for whether that process works
minister.industry.gov.au ↗
Government sets the standard those processes must meet, and ensures companies have robust processes in place.
minister.industry.gov.au ↗
These types of approaches could ensure that for developers of models with the sharpest, most acute frontier AI risks
minister.industry.gov.au ↗
the Government is developing AI standards legislation
minister.industry.gov.au ↗
voluntary regulation and codes are fast and flexible, but they fail the incentive test.
minister.industry.gov.au ↗
Our rules will raise the floor for firms that operate within our laws.
abc.net.au ↗
Labor is finalising national AI standards due by the end of the year with plans to legislate the new rules in 2027.
abc.net.au ↗
OpenAI has now backed mandatory safety requirements, independent assessments and incident reporting
proactiveinvestors.com ↗
The approach is expected to form the basis of legislation in 2027
proactiveinvestors.com ↗
developers could soon be required to prove it before the most powerful models are widely deployed
Open this check in the full collection →
The idea, illustrated03

Go inside the check.

    The full explanation appears when your reading access is confirmed.
    Next: Cheap per token, not per task

    Anthropic says Haiku 5.5 is its cheapest small model yet. Per token, yes. Per finished task, one index puts it far above a rival at the same price.

    Anthropic calls Haiku 5.5 its cheapest, fastest and most capable small model, about 75% cheaper to run than Haiku 4.5.

    Members · 30 days free

    See what the evidence actually shows.

    Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

    30 days free. Payment card required. One introductory trial per customer.

    Your trial ends on 30 days after you start. Unless you cancel before then in Account → Manage membership, we charge A$89 for the first year. It renews automatically at A$89/yr until cancelled.

    You can ask for a full refund within 14 days after any annual payment, renewals included, with no reason needed, by emailing [email protected]. This voluntary refund does not limit your rights under the Australian Consumer Law.

    By starting your trial, you agree to the Terms.

    BS Killer is published by Inferno Tech Pty Ltd, ABN 27 647 413 474.

    Open the evidence3 source pages

    The claim we checked

    Claude Haiku 5.5 is the cheapest, fastest and most capable small model Anthropic has released, and costs around 75% less to run than Haiku 4.5.

    These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

    anthropic.com ↗
    the cheapest, fastest, and most capable small model
    anthropic.com ↗
    Haiku 5.5 is available at a much lower price than Haiku 4.5.
    anthropic.com ↗
    On average, it now costs around 75% less to run.
    anthropic.com ↗
    50% lower for requests over 100,000 tokens
    anthropic.com ↗
    which means it uses slightly more tokens per task
    artificialanalysis.ai ↗
    the same as GPT-6 Luna and 10% of the previous Haiku model
    artificialanalysis.ai ↗
    However, this pricing rises 5x to $0.50/$2.50 above 100k.
    artificialanalysis.ai ↗
    uses ~162k output tokens per Intelligence Index task, ~3x GPT-6 Luna (max, ~50k)
    artificialanalysis.ai ↗
    At max effort Haiku 5.5 sits slightly ahead of models such as GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38)
    vals.ai ↗
    16 Claude Haiku 5.5 54.31% $2.99 $0.1 / $0.5
    vals.ai ↗
    26 GPT-6 Luna 51.22% $0.43 $0.1 / $0.5
    Open this check in the full collection →
    The idea, illustrated04

    Go inside the check.

      The full explanation appears when your reading access is confirmed.
      The finish line

      Edition complete

      You’re up to speed.

      That’s the 8 Oct 2026 briefing. Keep the useful bits. Leave the noise.

      Reading estimate: 588 words at 200 words per minute. Source quotes and the optional sections below add reading time.

      Have another 3 minutes? · Learn one thing

      Tokens: the pieces AI reads

      Two words can be two tokens. One word can be six. See what changes. A beginner lesson with a visual you can play.

      Try the free lesson →

      Keep exploring

      All research

      Browse every checked claim by topic.

      A curated directory. Check each entry’s date and sources.

      Make a little room for clarity.

      Get the next checked edition in your inbox.

      Free email updates. Unsubscribe any time.