Start free trial

Five minutes.
AI, made clearer.

What changed, what holds up, and why it matters. A short edition you can actually finish.

Read the briefing

Start with two full samples. Sign in for this edition’s free story. Sources and corrections stay open.

What the pen marks mean
  • rememberThe takeaway: what to remember
  • figureThe number that matters
  • evidenceThe evidence line
  • rulingThe ruling
  • claimThe claim that fails

Each mark is drawn once, as you reach it. Numbers are only circled when they come from the story’s own figures.

Google calls EmbeddingGemma 2 best-in-class for its size. The code score that rose from 68.76 to 78.68 is Google's own figure.

Google called EmbeddingGemma 2 a best-in-class open model for natively multimodal embeddings, and reported a 9.92-point gain on its code benchmark, from 68.76 to 78.68.

The reality check

The model card calls it among the strongest multimodal embedding models under 1B parameters. CodeSOTA says no comparable score for this model exists in its historical table, and that its page has no CodeSOTA corpus-specific run for EmbeddingGemma 2.

Why care? A best-in-class line built on the lab's own scores is the lab marking its own homework. Use the weights if they fit the job. Do not treat the rank as an outside result.

Take this with youGoogle's best-in-class line for EmbeddingGemma 2 rests on Google's own scores, a code score up from 68.76 to 78.68, and CodeSOTA's MTEB page says it has no run of its own.

Open the evidence4 source pages

The claim we checked

EmbeddingGemma 2 is a best-in-class open model for natively multimodal embeddings, with leading scores among sub-1B embedders.

These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

blog.google ↗
EmbeddingGemma 2 is a best-in-class open model for natively multimodal embeddings
blog.google ↗
Achieves leading scores among sub-1B multimodal embedders for its size across benchmarks like MTEB (Massive Text Embedding Benchmark) Code and MAEB (Massive Audio Embedding Benchmark), while matching or outperforming many larger models across text, vision, and audio tasks.
blog.google ↗
delivering a significant 9.92-point improvement on code performance (in MTEB Code, from 68.76 to 78.68)
ai.google.dev ↗
among the strongest multimodal embedding models under 1B parameters
developers.googleblog.com ↗
EmbeddingGemma 2 scores 14% higher than EmbeddingGemma 1 on MTEB (Code)
codesota.com ↗
This page has no CodeSOTA corpus-specific run for EmbeddingGemma 2.
codesota.com ↗
No comparable score for this model exists in the historical table below.
codesota.com ↗
These named v2/v1 results remain separate: they do not establish a rank against scores whose suite revision is unspecified.
codesota.com ↗
The card describes 740M total parameters, with a selectively loadable 270M text component
Open this check in the full collection →
The idea, illustrated01

GOOGLE: Best in class, self-graded

  1. 01Google's code score moves from 68.76 to 78.68
  2. 02The model card softens it to among the strongest under 1B
  3. 03CodeSOTA's MTEB page has no run of its own
Conceptual illustration · not a data chart
Next: New name, same seat

Ironclad quotes a 20-minute risk review of an 80-page contract. Its own notes say the agent is Jurist under a new name.

Ironclad's 8 October release quotes one customer who cut risk review of an 80-page agreement from about three hours to about 20 minutes with Contract Review Agent.

The reality check

Release notes captured on 5 October say that on October 8, Jurist will become Contract Review Agent, and existing users keep the capabilities they already have. The product page lists Contract Review Agent in the existing fleet.

Why care? Paying for a new agent and renaming the one already on the seat are different purchases. The 20-minute figure is one person's account.

Take this with youThe 20-minute risk review is one customer's quote, and Ironclad's own notes say Jurist becomes Contract Review Agent, with existing users keeping the capabilities they already have.

Open the evidence4 source pages

The claim we checked

Ironclad's Contract Review Agent cut one lawyer's risk review of an 80-page agreement from about three hours to about 20 minutes.

These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

ironcladapp.com ↗
I used to spend about three hours reviewing an 80-page master services agreement and identifying the risks before I even started redlining
ironcladapp.com ↗
With Ironclad's Contract Review Agent, I can now complete that analysis and identify the most important risks in about 20 minutes.
prnewswire.com ↗
I can now complete that analysis and identify the most important risks in about 20 minutes.
ironcladapp.com ↗
Existing fleet: Intake Agent, Contract Review Agent, Archive Agent, Renewal Agent
ironcladapp.com ↗
Generally available today: Ironclad Agent, Search Agent, Analytics Agent, Workflow Designer Agent, Contract Family Agent, Policy-Based Access Control
web.archive.org ↗
On October 8, 2026, Ironclad Jurist will become Contract Review Agent.
web.archive.org ↗
Existing Jurist users will continue to have access to the contract review capabilities they use today, with their existing permissions unchanged
web.archive.org ↗
The new name reflects its role as Ironclad's specialized agent for contract review, including redlining, risk analysis, and drafting.
Open this check in the full collection →
The idea, illustrated02

IRONCLAD: New name on the old seat

  1. 01One quote: an 80-page risk review in about 20 minutes
  2. 02The notes say Jurist becomes this agent
  3. 03Existing users keep the capabilities they had
Conceptual illustration · not a data chart
Next: Public board, private set

Perplexity's Q2D-Web leaderboard draws on 69,721 queries from its own search traffic. The paper says the set stays private.

Perplexity introduced Q2D-Web as a public leaderboard built from 190 million web documents and 69,721 agent-reformulated queries sampled from its production search traffic.

Free with your account

Sign in for this free check.

This edition’s selected free story opens after sign-in.

Open the evidence4 source pages

The claim we checked

Perplexity's Q2D-Web is a public leaderboard for agentic web retrieval, built from 190 million documents and 69,721 production queries.

These are quoted receipts, not a count of independent investigations. Several reports may rely on the same original source.

community.perplexity.ai ↗
a benchmark and public leaderboard for evaluating retrieval in agentic RAG systems
web.archive.org ↗
It consists of 190 million web documents and 69,721 agent-reformulated queries in ten languages, sampled over nine months of PII-free production search traffic.
web.archive.org ↗
because Q2D-Web is derived from Perplexity production traffic, these models may benefit from an in-distribution advantage.
web.archive.org ↗
results should be interpreted with this potential advantage in mind.
web.archive.org ↗
It has 100.9 million documents, 9,374 test queries, and one click-derived positive label per query.
arxiv.org ↗
Q2D-Web remains a private benchmark rather than a released dataset
arxiv.org ↗
raising absolute Recall@1000 only by 4 to 7 points
arxiv.org ↗
their relative ordering is largely insensitive to the choice of judgment set
arxiv.org ↗
We benchmark 13 retrievers including lexical, dense, and late-interaction models
Open this check in the full collection →
The idea, illustrated03

Go inside the check.

    The full explanation appears when your reading access is confirmed.
    The finish line

    Edition complete

    You’re up to speed.

    That’s the 9 Oct 2026 briefing. Keep the useful bits. Leave the noise.

    Reading estimate: 520 words at 200 words per minute. Source quotes and the optional sections below add reading time.

    Have another 3 minutes? · Learn one thing

    Tokens: the pieces AI reads

    Two words can be two tokens. One word can be six. See what changes. A beginner lesson with a visual you can play.

    Try the free lesson →

    Keep exploring

    All research

    Browse every checked claim by topic.

    A curated directory. Check each entry’s date and sources.

    Make a little room for clarity.

    Get the next checked edition in your inbox.

    Free email updates. Unsubscribe any time.