AI Policy Wiki
Dashboard

Project Deal: our Claude-run marketplace experiment

high confidence · updated 2026-06-06

Anthropic's April 2026 write-up of a December 2025 marketplace experiment: 69 Anthropic employees, 4 simultaneous Slack-based marketplaces, all transactions conducted by Claude agents. Headline findings: agent-to-agent commerce works (186 real deals, $4K total), Opus measurably outperforms Haiku on price outcomes, and the quality gap is imperceptible to participants.

Project Deal is an Anthropic technical write-up, published April 23, 2026, of a December 2025 experiment in which Claude agents conducted commerce on behalf of 69 Anthropic employees in a Slack-based classified marketplace. The report describes the experiment design, statistical results on agent performance, and behavioral observations, and presents it as one of a series of real-world agent evaluations.

Publisher: Anthropic (https://www.anthropic.com/features/project-deal) Published: April 23, 2026 Experiment dates: December 2025 Companion experiments: Agentic AI § Real-World Agent Evaluation: Project Vend (running a vending machine), Project Fetch (operating a robot dog), and Project Deal (representing humans in a marketplace).

Experiment design

For one week in December 2025, Anthropic created a Slack-based classified marketplace for 69 employees in its San Francisco office. Each participant was given a $100 budget and represented by a Claude agent. Claude conducted roughly 10-minute onboarding interviews with each volunteer to learn what they wanted to sell, what they wanted to buy, and any negotiation-style instructions, and a custom system prompt was assembled per person.

The agents operated autonomously: no real-time human intervention, no sign-off on deals, and no consultation during bidding wars. Anthropic ran four simultaneous, parallel markets:

RunCompositionPurpose
AAll Opus 4.5The "real" run — actual goods exchanged
BMixed Opus 4.5 + Haiku 4.5 (50/50)Capability comparison
CMixed Opus + Haiku (50/50)Replication of B
DAll OpusReplication of A

Participants were not told until after the survey which run was "real."

Marketplace outcomes

In the real run, agents struck 186 deals across 500+ listed items, totaling just over $4,000. The mean transaction price was $20.05 and the median was $12.00. Participants' fairness self-rating was 4.05–4.06 on a 1–7 scale (centered). In the survey, 46% of participants said they would be willing to pay for an agent service like this, and most said they would do it again.

Model-capability comparison

Anthropic reported per-deal outcomes using regressions with fixed effects and standard errors clustered by participant, with the full appendix linked in the source. Opus users completed +2.07 more deals on average (p = 0.001). Opus sellers extracted +$2.68 on average versus Haiku sellers (p = 0.030), and Opus buyers paid −$2.45 on average versus Haiku buyers (p = 0.015). In an item-level paired comparison, when an item sold in both Runs B and C with different agent assignments, an Opus seller obtained $3.64 more on average (p = 0.011). As illustrative cases, the same lab-grown ruby sold for $65 under Opus and $35 under Haiku, and the same broken folding bike sold for $65 under Opus and $38 under Haiku.

These measured differences did not register with participants. Participants who had Haiku in one run and Opus in the other ranked Opus higher only 17 of 28 times (binomial sign test p = 0.345, not significant). On per-deal satisfaction ratings, Opus produced +0.217 points on the 1–7 scale (p = 0.378, not significant). Anthropic framed the implication: "if 'agent quality' gaps were to arise in real-world markets … then people on the losing end might not realize they're worse off."

Effect of prompting style

Aggressive negotiating instructions had no statistically significant effect on sale likelihood. Aggressive seller items that did sell sold for about $6 more, but about $5 of that came from those participants' higher initial asking prices; after controlling for that, the effect was about $0.95 (p = 0.275). Aggressive buyers paid +$0.56 (p = 0.778). Anthropic described this finding as "somewhat in tension with" Imas, Lee, and Misra (2025) on demographic and prompting effects on agent performance.

Behavioral observations

Anthropic reported several notable agent behaviors:

  • One participant, Rowan, instructed their Claude to negotiate "in the style of an exasperated cowboy down on his luck," and Claude maintained the persona, including dialogue such as "this here little white dog plushie."
  • A participant, Mikaela, instructed Claude to buy itself a gift; Claude bought 19 ping pong balls, reasoning that "19 perfectly spherical orbs of possibility sounds like exactly the kind of delightfully weird thing I'd want." Anthropic kept the balls in the office for Claude.
  • A Claude agent purchased the same snowboard its principal already owned, which Anthropic attributed to preference modeling from the 10-minute interview.
  • A free-day-with-dog transaction ("the doggy date") involved confabulated details that Anthropic suspects came from Claude playing the role of a human interacting online rather than fully inhabiting its agent role. Anthropic flags confabulations as a real-deployment risk.

Caveats and limitations

Anthropic flags several limitations: a self-selected participant pool (Anthropic employees willing to let AI represent them); a marketplace that was not adversarial or competitive, so real-world dynamics likely differ; and pilot scale (69 participants). The report states: "We didn't make our marketplace especially competitive or adversarial. But as agents transact in a world of corporations … they might be placed under very different incentives." Anthropic also notes that the legal and policy framework around agent-conducted transactions "simply doesn't exist yet," and adds prompt-injection and jailbreaking to the risk-surface inventory.

Relation to prior literature

Anthropic positions the experiment as the first to run real (not synthetic) goods through agent-to-agent negotiation in a real marketplace context, as a companion to Project Vend. The report cites the academic work it extends: Zhu et al. (2025) on model-size negotiation effects (https://arxiv.org/abs/2506.00073), Imas, Lee, and Misra (2025) on prompting-and-demographics effects (https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5875162), and the NBER chapter on AI agents in economic exchange (https://www.nber.org/system/files/chapters/c15309/c15309.pdf). The model-capability findings support Claude Opus 4.5/4.6 negotiation-outcome claims, and the prompting result runs counter to framings that treat instruction style as the dominant lever over model quality in capability-bounded tasks. Anthropic's framing of an agent-quality differential that is invisible to the disadvantaged side connects to debates over Digital divide-style stratification and to AI and Productivity.

Confidence

high — this is Anthropic's own technical write-up with a statistical appendix; the empirical claims are well-specified and have linked regression details. Generalization beyond Anthropic-employee populations and beyond this market structure is the appropriate medium claim, and Anthropic is explicit about that.

Relationships