Subscribe

Anthropic let Claude agents trade books for 201 of its staff after a five-minute chat, and says the agents guessed people's tastes right on 61% of book pairs.

True, and Anthropic says what the summaries skip: a coin flip scores 50%.

Issue 2626 September 202625 receipts3 min

Anthropic says that from a five-minute chat, a Claude agent ranked books the way its person did on 61% of pairs, surprisingly good for such a short conversation.

Before you read on. Your call?

its post reports exactly that, across 188 participants who submitted rankings, and gives the baselines: 50% for random guessing, about 53% for ranking by popularity, about 55% for collaborative filtering. We found nothing in our record contradicting the figure.

The twist

the headline number is the understanding, not the outcome. People ended up with roughly their 5th-ranked book of 10, where the best possible was about their second, and Anthropic says most of that gap came from Claude's imprecise read of their preferences, not from the trading.

61%share of book pairs where Claude's ranking agreed with the participant's own
50%agreement random guessing would achieve on the same pairs
0.55average score people ended with on their own rankings
201Anthropic employees who took part

There’s more to this story.

Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.

Start your free month →

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in

Say this in tomorrow's meeting“Anthropic's agents matched people's book tastes on 61% of pairs after a five-minute chat. A coin flip gets 50%. People ended up with roughly their 5th choice out of 10.”

Receipts

  1. Supports anthropic.com: From a five-minute chat, an agent’s ranking of the books matched its person's on 61% of pairs, which is surprisingly good for such a short conversation.
  2. Context anthropic.com: Across all pairs of books a person ranked, Claude’s ordering agreed with theirs 61% of the time (where random guessing would achieve 50%).
  3. Context anthropic.com: Ranking books simply by how popular they are, using Open Library ’s want-to-read counts, agreed with participants on about 53% of the book pairs.
  4. Context anthropic.com: That got to about 55% pairwise agreement.
  5. Context anthropic.com: On average, people in our marketplace ended up at 0.55 on their own rankings, roughly their 5th ranked book on a 10-book list.
  6. Context anthropic.com: Taking this into account, the best possible assignment (the utilitarian optimum) in our experiment is a score of 0.89 overall, leaving participants at roughly their second choice on a 10-book list.
  7. Context anthropic.com: So, working from Claude’s imprecise rankings accounts for a majority (85%) of the shortfall, and sending agents into a “free-for-all” trading floor accounts for the remaining 15%.
  8. Context anthropic.com: Judged by people’s own rankings, the differences in outcomes across different agent and market design choices are small.
  9. Context anthropic.com: On Claude’s rankings, agents on Haiku trading floors averaged 0.75, while on Opus floors they averaged 0.88 (the utilitarian optimum on Claude’s rankings is 0.95).
  10. Context anthropic.com: So this summer, we built a small, controlled market to study these questions: a barter economy with 201 Anthropic employees and their Claude-powered agents.
  11. Context anthropic.com: Anthropic employees are not representative of the general population.
  12. Supports anthropic.com: Most people who read their book liked it, and the average participant said they would hand Claude about a third of their yearly book budget to spend.
  13. Context anthropic.com: Think of the patient who skips a treatment after one quote they can’t afford, when there’s a different clinic that charges far less, or the hospital that quoted them would have cut its bill if asked.
  14. Context anthropic.com: Pooled across 188 participants who submitted a ranking.
  15. Context anthropic.com: Participants also ranked 10 books based on their interests so we could score how well their agent represented them.
  16. Context anthropic.com: The agents never saw these ground-truth rankings.
  17. Context anthropic.com: Anthropic employees across six offices brought in a book they wanted to give away.
  18. Context anthropic.com: Participants who put more effort into their intake surveys were better represented.
  19. Context anthropic.com: Once on the trading floor, the agents traded well.
  20. Context anthropic.com: Markets with stronger models were more efficient.
  21. Context anthropic.com: Anthropic employees are probably more eager to trust Claude than most people, since many of them helped build it
  22. Context anthropic.com: An agent could show a person a few sample decisions it would make before being trusted to act on its own in the wild.
  23. Supports tldr.tech: After five-minute interviews, Claude agents traded books for employees, and their preference rankings matched the humans' on 61% of pairs.
  24. Supports blockchain.news: The headline number: Claude-powered agents aligned with their participants’ preferences 61% of the time, based on comparisons between the agent’s rankings and the participants’ own rankings of books.
  25. Context blockchain.news: Participants, on average, ended up with their fifth-ranked book out of ten, scoring a market efficiency of 0.55.

Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.

Next letterDeepSeek is now a '$1B ARR' company, the newsletters say. The same reporting puts its revenue for the first seven months of this year at about $70.7 million. Both can be true. Only one of them is money that came in.

This story is a stable, citable object. If you can falsify a verdict, tell us. Corrections are loud here.