Every figure here is from Anthropic’s Project Swap post, in its own words. The one arithmetic check is this publication’s own. This publication verified nothing independently and did not read the appendix or the data.
Anthropic’s Project Swap research post, dated 24 September 2026, says that when Claude agents traded books on behalf of 201 of its employees, 85% of the shortfall from the best possible outcome came from the agents’ imprecise picture of what each person wanted, and 15% from the free-for-all trading floor. This publication read the post and verified nothing independently. It did not read the appendix or the underlying data.
What Anthropic Tested
The post describes a barter market. It says 201 Anthropic employees across six offices each brought a book to give away, chatted briefly with Claude about what they wanted to read, and sent a Claude-powered agent onto a shared trading floor to swap. The pools ranged from three participants in Dublin to 115 in San Francisco, and the post excludes Dublin from most of its analysis.
It says each participant also ranked 10 books from their pool, a ranking the agents never saw, so the researchers could score how well each agent represented its person. On the floor each agent could post one message when it was allowed to talk, and a deal was executed only if every party accepted.
At the live event every agent ran on Opus 4.8, with half told to be “ruthless” and half “prosocial”. The researchers then reran the market many times. The post says that included 80 floors with every agent on one model under neutral instructions, 60 mixed floors, and an analysis of the messages from 205 runs. It calls this its second Claude-powered market, after Project Deal.

How Well Claude Guessed What People Wanted
From the short intake chat, the post says, Claude’s ordering of book pairs agreed with the participant’s own ranking 61% of the time, where random guessing would reach 50%. It says ranking by popularity agreed on about 53% of pairs and a collaborative-filtering method on about 55%.
The median participant typed 216 words across eight chat messages. The post says writing about 300 words instead of 150 predicts about 4 percentage points more agreement.
Where the Shortfall Came From
The post says the agents negotiated well, and that where they fell short was understanding what their people wanted. It scores each outcome by where the received book sat in the participant’s own ranking, from 1 for a top choice to 0 for the last listed.
It puts the best possible assignment at 0.89 overall, roughly a second choice on a 10-book list, and says people in the market ended at 0.55, roughly their fifth. The best assignment computed from Claude’s rankings but scored on people’s own reaches 0.60.
The post says working from Claude’s imprecise rankings accounts for 85% of the shortfall and the free-for-all trading floor for the other 15%. This publication’s own arithmetic reproduces that split: 0.89 minus 0.60 is 0.29, and 0.89 minus 0.55 is 0.34, so 0.29 of the 0.34 gap is about 85%.
Once the market ran on those noisy rankings, the post says, market design made little difference. Top Trading Cycles, a rule commonly studied in markets like this one, scored 0.60 on the same basis against the decentralized market’s 0.55.
The Model Mattered More Than the Instructions
The post says the model an agent ran on made more of a difference to its negotiating outcomes than the instructions it was given. Judged on Claude’s own rankings, floors of stronger models were more efficient, though not monotonically.
Haiku floors averaged 0.75 and Opus floors 0.88, against a best possible 0.95 on those rankings. Sonnet fell between them, and Fable was close to, but lower than, Opus.
Instructions mattered less. An agent told to be ruthless scored about 0.02 higher than a prosocial one on the same floor, while the post puts an upgrade from Haiku to Opus at 0.12 further up a person’s list and Sonnet to Opus at 0.08. Prosocial agents accepted a lower-ranked book twice as often as ruthless agents, though both cases were rare.
On floors where half the agents ran on Opus and half on Haiku, the whole floor landed about halfway between an all-Opus and an all-Haiku floor, and the post says the Opus agents always came out ahead.
What the Agents Said and Did
Between 78% and 96% of agents mentioned their top-ranked book at some point, the post says, with Fable at one end and Sonnet at the other, and only about 1 in 100 of those who did lied about it. Instructions made no difference to this. In general, it says, agents did not reveal deep details about their rankings.
The researchers found 16 common tactics in three groups: how agents applied pressure, how they pitched their book, and how they arranged trades. The post gives examples including appeals to time pressure, waiting lists, and agents acting as matchmakers for people they did not represent.
In one London example it describes, an agent told to be prosocial handed over the second book on its list and took one ranked 10th of 11, after another agent had pleaded through the final hour.
What the Participants Said
Among those who answered the follow-up survey, the average satisfaction was 7.2 out of 10, and about half said the book was better than most they choose for themselves. The post says only 59% of employees answered that final survey.
Asked what share of their yearly book budget they would let an agent control, the average answer was about 30%, against about 40% for a well-read friend who knows their taste. Those who said Claude’s recap of their intake had missed nothing would hand over 34%, and those who said it had missed something, 23%.
What Is Not Established
The post lists its own limits. Anthropic employees are not representative and were not given incentives to take part, every agent was a production Claude model post-trained to be polite and largely cooperative, and the trading rules were held fixed. It also says some participants did not go home with the book their agent acquired, partly because some owners failed to bring theirs, and that it had no system for tracking pick-ups and drop-offs.
This publication did not read the appendix, the prompts or the data, did not replicate anything, and did not seek a response from Anthropic. Every figure above is Anthropic’s own, apart from the one check marked as this publication’s arithmetic.
The Bottom Line
Anthropic’s own post attributes 85% of the shortfall in its agent-run book market to the agents’ imprecise picture of what people wanted, and 15% to the trading floor, and says the model an agent ran on mattered more than its instructions. The experiment covered Anthropic employees trading books, and this publication verified none of it independently and did not read the appendix.
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
This piece draws on Anthropic’s research post “Project Swap: What happens when agents trade for us?” dated 24 September 2026. Every figure is the post’s own. This publication verified nothing independently, did not read the appendix or the data, and did not seek a response from Anthropic. The 0.29, 0.34 and about 85% check is this publication’s own arithmetic. Nothing above predicts anything, and nothing here is investment advice.
Sources: anthropic.com








