Start free trial

The trick: Self-Marked

Q2D-Web is a public leaderboard of 69,721 queries from Perplexity's own traffic.

The paper says the benchmark stays private.

Issue 399 October 20269 receipts3 min

Q2D-Web is a public leaderboard for agentic web retrieval.

Before you read on. Your call?

Perplexity's post says it holds 190 million documents and 69,721 queries from the company's own production traffic, and that Perplexity models may benefit from an in-distribution advantage.

The twist

The paper says the benchmark remains private rather than a released dataset, that 13 retrievers kept a largely stable order across judgment sets, and that a one-third sample raises Recall@1000 by 4 to 7 points.

190 millionweb documents in Q2D-Web
69,721agent-reformulated queries in Q2D-Web
4 to 7points a one-third corpus sample adds to absolute Recall@1000
13retrievers benchmarked in the Q2D-Web paper

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Review my 30-day free trial →

30 days free. Payment card required. One introductory trial per customer.

Your trial ends on 30 days after you start. Unless you cancel before then in Account → Manage membership, we charge A$89 for the first year. It renews automatically at A$89/yr until cancelled.

You can ask for a full refund within 14 days after any annual payment, renewals included, with no reason needed, by emailing hello@bskiller.com. This voluntary refund does not limit your rights under the Australian Consumer Law.

By starting your trial, you agree to the Terms.

BS Killer is published by Inferno Tech Pty Ltd, ABN 27 647 413 474.

Already a member? Sign in

The trick has a name

We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →

Say this in tomorrow's meeting“Perplexity's Q2D-Web leaderboard uses 69,721 queries from its own search traffic. The paper says the benchmark stays private, and Perplexity says its models may have a home-field edge.”

Receipts

  1. Supports community.perplexity.ai: a benchmark and public leaderboard for evaluating retrieval in agentic RAG systems
  2. Supports web.archive.org: It consists of 190 million web documents and 69,721 agent-reformulated queries in ten languages, sampled over nine months of PII-free production search traffic.
  3. Context web.archive.org: because Q2D-Web is derived from Perplexity production traffic, these models may benefit from an in-distribution advantage.
  4. Context web.archive.org: results should be interpreted with this potential advantage in mind.
  5. Context web.archive.org: It has 100.9 million documents, 9,374 test queries, and one click-derived positive label per query.
  6. Refutes arxiv.org: Q2D-Web remains a private benchmark rather than a released dataset
  7. Context arxiv.org: raising absolute Recall@1000 only by 4 to 7 points
  8. Context arxiv.org: their relative ordering is largely insensitive to the choice of judgment set
  9. Context arxiv.org: We benchmark 13 retrievers including lexical, dense, and late-interaction models

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

Get tomorrow's check in your inbox.

One AI claim a night, checked against independent sources, receipts attached. Free.

Free email updates. Unsubscribe any time.

This story is a stable, citable object. If you can falsify a verdict,tell us. Corrections are loud here.