Start free trial

Five minutes.
AI, made clearer.

Graded at home.

Read the briefing

Members read every story in the archive, re-verified. Sources and corrections stay open to everyone.

What the pen marks mean
  • rememberThe takeaway: what to remember
  • figureThe number that matters
  • evidenceThe evidence line
  • rulingThe ruling
  • claimThe claim that fails

Each mark is drawn once, as you reach it. Numbers are only circled when they come from the story’s own figures.

Google calls EmbeddingGemma 2 best-in-class for its size. The code score that rose from 68.76 to 78.68 is Google's own figure.

Google called EmbeddingGemma 2 a best-in-class open model for natively multimodal embeddings, and reported a 9.92-point gain on its code benchmark, from 68.76 to 78.68.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

  1. At sign-up: 30 days free. Payment card required. One introductory trial per customer.
  2. Your trial ends 30 days after you start. Unless you cancel before then in Account → Manage subscription, we charge A$89 for the first year. It renews automatically at A$89/yr until cancelled.
  3. You can ask for a full refund within 14 days after any annual payment, renewals included, with no reason needed, by emailing [email protected]. This voluntary refund does not limit your rights under the Australian Consumer Law.

By starting your trial, you agree to the Terms.

BS Killer is published by Inferno Tech Pty Ltd, ABN 27 647 413 474.

Open the evidence4 source pages

The claim we checked

EmbeddingGemma 2 is a best-in-class open model for natively multimodal embeddings, with leading scores among sub-1B embedders.

These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

blog.google ↗
EmbeddingGemma 2 is a best-in-class open model for natively multimodal embeddings
blog.google ↗
Achieves leading scores among sub-1B multimodal embedders for its size across benchmarks like MTEB (Massive Text Embedding Benchmark) Code and MAEB (Massive Audio Embedding Benchmark), while matching or outperforming many larger models across text, vision, and audio tasks.
blog.google ↗
delivering a significant 9.92-point improvement on code performance (in MTEB Code, from 68.76 to 78.68)
ai.google.dev ↗
among the strongest multimodal embedding models under 1B parameters
developers.googleblog.com ↗
EmbeddingGemma 2 scores 14% higher than EmbeddingGemma 1 on MTEB (Code)
codesota.com ↗
This page has no CodeSOTA corpus-specific run for EmbeddingGemma 2.
codesota.com ↗
No comparable score for this model exists in the historical table below.
codesota.com ↗
These named v2/v1 results remain separate: they do not establish a rank against scores whose suite revision is unspecified.
codesota.com ↗
The card describes 740M total parameters, with a selectively loadable 270M text component
Open this check in the full collection →
The idea, illustrated01

Go inside the check.

    The full explanation appears when your reading access is confirmed.
    Next: New name, same seat

    Ironclad quotes a 20-minute risk review of an 80-page contract. Its own notes say the agent is Jurist under a new name.

    Ironclad's 8 October release quotes one customer who cut risk review of an 80-page agreement from about three hours to about 20 minutes with Contract Review Agent.

    Members · 30 days free

    See what the evidence actually shows.

    Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

    1. At sign-up: 30 days free. Payment card required. One introductory trial per customer.
    2. Your trial ends 30 days after you start. Unless you cancel before then in Account → Manage subscription, we charge A$89 for the first year. It renews automatically at A$89/yr until cancelled.
    3. You can ask for a full refund within 14 days after any annual payment, renewals included, with no reason needed, by emailing [email protected]. This voluntary refund does not limit your rights under the Australian Consumer Law.

    By starting your trial, you agree to the Terms.

    BS Killer is published by Inferno Tech Pty Ltd, ABN 27 647 413 474.

    Open the evidence4 source pages

    The claim we checked

    Ironclad's Contract Review Agent cut one lawyer's risk review of an 80-page agreement from about three hours to about 20 minutes.

    These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

    ironcladapp.com ↗
    I used to spend about three hours reviewing an 80-page master services agreement and identifying the risks before I even started redlining
    ironcladapp.com ↗
    With Ironclad's Contract Review Agent, I can now complete that analysis and identify the most important risks in about 20 minutes.
    prnewswire.com ↗
    I can now complete that analysis and identify the most important risks in about 20 minutes.
    ironcladapp.com ↗
    Existing fleet: Intake Agent, Contract Review Agent, Archive Agent, Renewal Agent
    ironcladapp.com ↗
    Generally available today: Ironclad Agent, Search Agent, Analytics Agent, Workflow Designer Agent, Contract Family Agent, Policy-Based Access Control
    web.archive.org ↗
    On October 8, 2026, Ironclad Jurist will become Contract Review Agent.
    web.archive.org ↗
    Existing Jurist users will continue to have access to the contract review capabilities they use today, with their existing permissions unchanged
    web.archive.org ↗
    The new name reflects its role as Ironclad's specialized agent for contract review, including redlining, risk analysis, and drafting.
    Open this check in the full collection →
    The idea, illustrated02

    Go inside the check.

      The full explanation appears when your reading access is confirmed.
      Next: Public board, private set

      Perplexity's Q2D-Web leaderboard draws on 69,721 queries from its own search traffic. The paper says the set stays private.

      Perplexity introduced Q2D-Web as a public leaderboard built from 190 million web documents and 69,721 agent-reformulated queries sampled from its production search traffic.

      Members · 30 days free

      See what the evidence actually shows.

      Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

      1. At sign-up: 30 days free. Payment card required. One introductory trial per customer.
      2. Your trial ends 30 days after you start. Unless you cancel before then in Account → Manage subscription, we charge A$89 for the first year. It renews automatically at A$89/yr until cancelled.
      3. You can ask for a full refund within 14 days after any annual payment, renewals included, with no reason needed, by emailing [email protected]. This voluntary refund does not limit your rights under the Australian Consumer Law.

      By starting your trial, you agree to the Terms.

      BS Killer is published by Inferno Tech Pty Ltd, ABN 27 647 413 474.

      Open the evidence4 source pages

      The claim we checked

      Perplexity's Q2D-Web is a public leaderboard for agentic web retrieval, built from 190 million documents and 69,721 production queries.

      These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

      community.perplexity.ai ↗
      a benchmark and public leaderboard for evaluating retrieval in agentic RAG systems
      web.archive.org ↗
      It consists of 190 million web documents and 69,721 agent-reformulated queries in ten languages, sampled over nine months of PII-free production search traffic.
      web.archive.org ↗
      because Q2D-Web is derived from Perplexity production traffic, these models may benefit from an in-distribution advantage.
      web.archive.org ↗
      results should be interpreted with this potential advantage in mind.
      web.archive.org ↗
      It has 100.9 million documents, 9,374 test queries, and one click-derived positive label per query.
      arxiv.org ↗
      Q2D-Web remains a private benchmark rather than a released dataset
      arxiv.org ↗
      raising absolute Recall@1000 only by 4 to 7 points
      arxiv.org ↗
      their relative ordering is largely insensitive to the choice of judgment set
      arxiv.org ↗
      We benchmark 13 retrievers including lexical, dense, and late-interaction models
      Open this check in the full collection →
      The idea, illustrated03

      Go inside the check.

        The full explanation appears when your reading access is confirmed.
        The finish line

        Edition complete

        You’re up to speed.

        That’s the 9 Oct 2026 briefing. Keep the useful bits. Leave the noise.

        Reading estimate: 520 words at 200 words per minute. Source quotes and the optional sections below add reading time.

        Get tomorrow's check in your inbox.

        One AI claim a night, checked against independent sources, receipts attached. Free.

        Free email updates. Unsubscribe any time.