The trick: Self-Marked
Q2D-Web is a public leaderboard of 69,721 queries from Perplexity's own traffic.
The paper says the benchmark stays private.
Q2D-Web is a public leaderboard for agentic web retrieval.
Before you read on. Your call?
TRUE, BUT
69721
Perplexity's post says it holds 190 million documents and 69,721 queries from the company's own production traffic, and that Perplexity models may benefit from an in-distribution advantage.
The twist
The paper says the benchmark remains private rather than a released dataset, that 13 retrievers kept a largely stable order across judgment sets, and that a one-third sample raises Recall@1000 by 4 to 7 points.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Get tomorrow’s check free by email.
30 days free. Payment card required. One introductory trial per customer.
Your trial ends on 30 days after you start. Unless you cancel before then in Account → Manage membership, we charge A$89 for the first year. It renews automatically at A$89/yr until cancelled.
You can ask for a full refund within 14 days after any annual payment, renewals included, with no reason needed, by emailing [email protected]. This voluntary refund does not limit your rights under the Australian Consumer Law.
By starting your trial, you agree to the Terms.
BS Killer is published by Inferno Tech Pty Ltd, ABN 27 647 413 474.
Couldn't check your access. That's on us.
The trick has a name
We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →
Receipts
- Supports community.perplexity.ai:
a benchmark and public leaderboard for evaluating retrieval in agentic RAG systems
- Supports web.archive.org:
It consists of 190 million web documents and 69,721 agent-reformulated queries in ten languages, sampled over nine months of PII-free production search traffic.
- Context web.archive.org:
because Q2D-Web is derived from Perplexity production traffic, these models may benefit from an in-distribution advantage.
- Context web.archive.org:
results should be interpreted with this potential advantage in mind.
- Context web.archive.org:
It has 100.9 million documents, 9,374 test queries, and one click-derived positive label per query.
- Refutes arxiv.org:
Q2D-Web remains a private benchmark rather than a released dataset
- Context arxiv.org:
raising absolute Recall@1000 only by 4 to 7 points
- Context arxiv.org:
their relative ordering is largely insensitive to the choice of judgment set
- Context arxiv.org:
We benchmark 13 retrievers including lexical, dense, and late-interaction models
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.
