Anthropic let Claude agents trade books for 201 of its staff after a five-minute chat, and says the agents guessed people's tastes right on 61% of book pairs.
True, and Anthropic says what the summaries skip: a coin flip scores 50%.
Anthropic says that from a five-minute chat, a Claude agent ranked books the way its person did on 61% of pairs, surprisingly good for such a short conversation.
Before you read on. Your call?
VERIFIED
61%
its post reports exactly that, across 188 participants who submitted rankings, and gives the baselines: 50% for random guessing, about 53% for ranking by popularity, about 55% for collaborative filtering. We found nothing in our record contradicting the figure.
The twist
the headline number is the understanding, not the outcome. People ended up with roughly their 5th-ranked book of 10, where the best possible was about their second, and Anthropic says most of that gap came from Claude's imprecise read of their preferences, not from the trading.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
Receipts
- Supports anthropic.com:
From a five-minute chat, an agent’s ranking of the books matched its person's on 61% of pairs, which is surprisingly good for such a short conversation.
- Context anthropic.com:
Across all pairs of books a person ranked, Claude’s ordering agreed with theirs 61% of the time (where random guessing would achieve 50%).
- Context anthropic.com:
Ranking books simply by how popular they are, using Open Library ’s want-to-read counts, agreed with participants on about 53% of the book pairs.
- Context anthropic.com:
That got to about 55% pairwise agreement.
- Context anthropic.com:
On average, people in our marketplace ended up at 0.55 on their own rankings, roughly their 5th ranked book on a 10-book list.
- Context anthropic.com:
Taking this into account, the best possible assignment (the utilitarian optimum) in our experiment is a score of 0.89 overall, leaving participants at roughly their second choice on a 10-book list.
- Context anthropic.com:
So, working from Claude’s imprecise rankings accounts for a majority (85%) of the shortfall, and sending agents into a “free-for-all” trading floor accounts for the remaining 15%.
- Context anthropic.com:
Judged by people’s own rankings, the differences in outcomes across different agent and market design choices are small.
- Context anthropic.com:
On Claude’s rankings, agents on Haiku trading floors averaged 0.75, while on Opus floors they averaged 0.88 (the utilitarian optimum on Claude’s rankings is 0.95).
- Context anthropic.com:
So this summer, we built a small, controlled market to study these questions: a barter economy with 201 Anthropic employees and their Claude-powered agents.
- Context anthropic.com:
Anthropic employees are not representative of the general population.
- Supports anthropic.com:
Most people who read their book liked it, and the average participant said they would hand Claude about a third of their yearly book budget to spend.
- Context anthropic.com:
Think of the patient who skips a treatment after one quote they can’t afford, when there’s a different clinic that charges far less, or the hospital that quoted them would have cut its bill if asked.
- Context anthropic.com:
Pooled across 188 participants who submitted a ranking.
- Context anthropic.com:
Participants also ranked 10 books based on their interests so we could score how well their agent represented them.
- Context anthropic.com:
The agents never saw these ground-truth rankings.
- Context anthropic.com:
Anthropic employees across six offices brought in a book they wanted to give away.
- Context anthropic.com:
Participants who put more effort into their intake surveys were better represented.
- Context anthropic.com:
Once on the trading floor, the agents traded well.
- Context anthropic.com:
Markets with stronger models were more efficient.
- Context anthropic.com:
Anthropic employees are probably more eager to trust Claude than most people, since many of them helped build it
- Context anthropic.com:
An agent could show a person a few sample decisions it would make before being trusted to act on its own in the wild.
- Supports tldr.tech:
After five-minute interviews, Claude agents traded books for employees, and their preference rankings matched the humans' on 61% of pairs.
- Supports blockchain.news:
The headline number: Claude-powered agents aligned with their participants’ preferences 61% of the time, based on comparisons between the agent’s rankings and the participants’ own rankings of books.
- Context blockchain.news:
Participants, on average, ended up with their fifth-ranked book out of ten, scoring a market efficiency of 0.55.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.