An agent can trade for a person only as well as it can represent changing preferences
Anthropic’s controlled book-barter experiment found that short intake chats let Claude rank pairs in line with participants 61% of the time. The scarce control is not bargaining speed but a calibrated, revisable representation of what the principal wants.

What happened
Anthropic ran a controlled barter market with 201 employees and Claude-powered agents across six offices. Participants ranked 10 books themselves; agent rankings derived from short intake chats agreed with participant pairwise orderings 61% of the time across 188 submitted rankings.
Why it matters
A competent market agent can still produce a poor outcome when its preference model is incomplete or stale. Delegated commerce therefore needs confidence, confirmation, reversibility and conflict rules before it needs broader authority.
Anthropic's Project Swap created a small barter market in which 201 employees brought books and Claude-powered agents traded on their behalf across six office pools. A short semi-structured intake conversation produced a ranking over every book in a person's pool. Separately, participants ranked 10 books themselves; agents never saw those rankings.
Across 188 participants who submitted a ranking, Claude's pairwise ordering agreed with the person's ordering 61% of the time, compared with 50% for random choice, about 53% for book popularity and about 55% for a collaborative-filtering baseline. The median participant typed 216 words across eight messages. Anthropic reports that roughly doubling intake length from 150 to 300 words was associated with about four percentage points more agreement. This is a controlled company experiment with employees, books and no money, not evidence about high-stakes procurement or consumer welfare.
Treat preference representation as a safety-critical input
An agent's authority should depend on how well the system knows what the principal wants. For each delegated task, record the preference source, when it was collected, which constraints are hard, which are negotiable and where confidence is low. A single intake conversation should not silently become a durable mandate. Preferences can change with price, timing, context, new information and the person's own learning.
Before action, show a compact preview: intended outcome, material trade-offs, constraints used and uncertainties. Require confirmation when confidence is low, consequences are hard to reverse or the action introduces a new counterparty. For repeated low-risk transactions, sample confirmations and compare accepted outcomes with predicted preferences. That creates calibration data rather than assuming the agent's explanation is evidence of accuracy.
Separate negotiation quality from representation quality
Project Swap's key analytical move is to distinguish the bargaining mechanism from the preference estimate. A market can allocate efficiently relative to an agent's ranking while still disappointing the person whose ranking was misrepresented. Production evaluation should preserve that separation. Measure representation agreement, constraint violations, regret after review, reversal rate and negotiation efficiency as different quantities.
The counterargument is that continual confirmation removes the benefit of delegation. A risk-tiered mandate avoids that trap. Let the agent act within bounded price, category, time and counterparty limits; require consent outside them; and provide an immediate, low-friction cancellation path. Where preferences conflict, such as lower price versus labour or privacy standards, do not let the model invent the priority. Ask or apply a named policy chosen in advance.
There is also a distribution problem. Project Swap pooled mostly company employees in a deliberately simple setting, with office pools ranging from three to 115 participants. Preference elicitation may perform differently across languages, accessibility needs, financial stress or users who communicate briefly. Test calibration by group and interface, but do not infer sensitive traits merely to improve a recommendation.
Set a redress rule before launch. The principal should be able to see which preference or constraint drove the action, correct it and know whether a counterpart has already relied on the transaction. Logs must support dispute resolution without exposing unrelated private conversation. For multi-agent markets, define who bears loss when a proxy misrepresents a preference, a counterparty exploits ambiguity or two automated policies conflict. Accountability cannot be delegated to the same agent whose representation is in question.
The practical decision is to gate agent authority on demonstrated preference calibration, not on fluent bargaining. Start with reversible transactions and a narrow mandate. Preserve the intake, proposed action, confidence, confirmation and outcome so the person can inspect and correct the representation. The Skills Intelligence glossary can help standardise terms, but the operating rule should be concrete: when preferences are uncertain or changing, the agent pauses before the transaction becomes irreversible.