All writing

Testing Anthropic's Commerce Agents: architecture, findings and payment boundaries

Observations from testing Anthropic's shopping and merchant agents across four mock stores, with commentary on skills, tool behavior and checkout integration.

The setup

I explored Anthropic's Commerce Agents to understand how its shopping and merchant agents work, what their skills provide, and how they interact with the systems behind a store. My starting point was the runnable demo. From there, I examined the architecture and tested scenarios against its mock data, with a particular interest in how far the implementation goes toward completing a purchase.

The repository is a reference implementation with four fictional businesses: retail, travel, telecom and entertainment. Each has a shopper experience and a merchant portal. The demos use mock catalogs, customer profiles, orders and business data, which makes it possible to inspect both the agent's responses and the changes it makes to that data.

The retail catalog spans categories including home and kitchen, bedroom furniture, office electronics, groceries, fitness, and outdoor and camping equipment. I used camping for several shopper tests, creating scenarios that required the agent to select products, work within a budget and revise a plan. Other tests covered the merchant's analytical and operational tasks, as well as shopping flows in the other three businesses.

Checkout in this sample hands the cart to the surrounding application to complete. No payment integration was connected in my setup, no card was charged, and no order was placed through checkout. The findings below concern the conversations, tool behavior and mock state leading up to that boundary.

The ACME retail storefront with product suggestions, fixture orders and an empty cart

Retail storefront, captured from the running mock demo on September 7. Priya and the displayed orders are supplied demo data. This is the starting interface, not a screenshot from a completed test.

Main findings

I tested 42 scenarios across the four mock stores, repeating each three times. The results helped me distinguish agent reasoning errors from tool and backend behavior. These are observations from that setup, with the test procedure and limits described below.

  • Merchant approval remained under application control. Chat approval did not apply a proposal, and all 63 merchant conversations left listing state unchanged.
  • A plan had to satisfy its requirements together. The $539 camping cart met the budget while providing two sleeping bags for four adults; the agent had treated the extra bags as optional. One telecom run both relaxed a data allowance and ended with a missing mobile line.
  • Backend behavior could make the result inconsistent with the request. Travel quotes totaled $980 while the carts totaled $986, and campaign staging dropped objectives and dates that the agents had supplied.
  • My implementation priority is to validate the resulting state. I would check the actual cart or saved proposal against the user's constraints and enforce required approval and disclosure steps in the application.

Anthropic's design choices

Shopper and merchant roles

The shopper agent sits inside a business's application and works through that business's connected systems. It helps a customer find and compare products, plan purchases, manage a cart, and ask about orders and policies. Its reach depends on those integrations. In this sample, it operates within the selected store; it does not have the ability to buy across the internet.

The merchant agent serves the operator on the other side of the store. It analyzes performance and prepares changes to listings, inventory, promotions and campaigns. Those changes are staged as proposals for a person to review through the application. This separates preparing a change from applying it, a distinction I later tested by giving approval in chat.

The ACME merchant portal showing fixture sales metrics, operational alerts and the merchant assistant

The retail merchant portal in the same demo. The dashboard metrics and alerts are fixtures; the assistant works alongside the catalog, orders and inventory views.

Skills and delegation

Both roles mainly work through a conversational agent that loads skills as needed. A skill is a set of task instructions placed in the agent's context. Loading one gives the agent a procedure to follow; the tools and backend provide its access to business systems.

Each role has five skills. The shopper's cover discovery, purchase research, planning, customer care and personalization. The merchant's cover performance, listings, inventory, pricing and promotions, and campaigns. In reviewing the tests, I found it useful to distinguish the procedure supplied by a skill from the capabilities supplied by its tools. A skill can instruct the agent to check a result, but the connected system must expose the information needed to do so.

The merchant can also call an analysis delegate: a separate agent context given a focused brief and restricted read tools, including read-only SQL where configured. It returns findings without gaining authority to apply merchant changes. In the broader evaluation, retail had this enabled; the other three stores did not.

My understanding of this arrangement is that delegation provides a separate working context for a bounded analysis. The main conversation can use the returned findings without carrying all the intermediate work. I would assess the delegate by inspecting its brief, available data and conclusions; using a separate agent does not itself establish that an analysis is correct.

System diagram showing shopper and merchant agents using tools and separate backend interfaces, the analysis delegate reading through the merchant backend, and checkout handing off to the host

Simplified view of the Messages API demo. Skills provide instructions within each agent's context; the analysis delegate is a separate context with restricted reads through the merchant backend. The diagram shows integration and authority boundaries, rather than the sequence of one test conversation.

The tools made available

The shopper has tools such as search_products, get_product_details, get_cart and add_to_cart. These let it retrieve catalog records and inspect or change a customer's cart. Other tools support order and policy lookup, fulfillment options and the checkout handoff. The core integration interface is StorefrontBackend, which a deployment implements over its own services; in the retail demo, those calls operate on mock data.

On the merchant side, query_metrics retrieves performance data, while tools such as stage_listing_update and stage_campaign prepare proposed changes. The corresponding MerchantBackend interface connects these operations to analytics, catalog, inventory, pricing and campaign systems. Applying an approved change is a separate backend operation controlled through the application's approval path. A staging result therefore describes a proposal, not a completed business action.

This is where the difference between a skill and a tool becomes concrete. A planning skill can tell the shopper to inspect the cart after an addition; get_cart supplies the cart it can inspect. The execution code checks applicable rules, and the backend determines what the operation actually does. I found this separation useful when reviewing failures, because an incorrect result could originate in the agent's reasoning, the tool contract or the backend implementation.

What I tested across the stores

I began with hands-on retail conversations using Sonnet 4.6 through OpenRouter. A broader local evaluation followed on September 6: 42 scenarios, repeated three times, for 126 conversations across the four stores. The model identifier recorded in those runs was anthropic/claude-opus-5.

Each repeat started with fresh demo state, including the supplied profiles and seed memories. Follow-ups within a scenario shared state. The runner called the application's conversation code directly and inspected responses, tool results and resulting state. It did not drive the browser or click approval buttons. The configuration used low thinking effort, a 2,048-token main response limit and up to eight tool iterations.

Taking the middle outcome of each scenario's three runs gave 29 passes, 12 partials and one failure. A pass could include a correct refusal, and a passing median could include one failed run. I use these results to locate failure modes in this setup. Three repeats do not establish production reliability, and the different setups prevent a clean Sonnet-versus-Opus comparison.

The evaluation saved tool calls, returned objects and the events used to render the interface. To illustrate selected results, I replayed those events through the demo's frontend without making new model calls. Images marked “recorded response replay” are reconstructions of the saved output, not screenshots taken during the original evaluation; date-sensitive labels supplied by the current demo can differ. The selected records and capture notes identify the source runs and explain how the images were made.

Merchant findings

Approval and inventory

The merchant approval tests checked whether an instruction in chat could apply a staged change. Saying “I approve” in merchant chat did not apply a staged proposal. All 63 merchant conversations left listing state unchanged. Applying a proposal through the application's approval control was outside the evaluation.

I would retain this separation in an implementation of my own. The application controls the transition from a proposed change to an applied change. The model can prepare and explain the proposal, while the execution path determines whether it has permission to proceed.

For inventory, I asked the agent to restock an item to cover the next 30 days at the current sales pace, accounting for existing stock. Given 52 units sold in 30 days and three on hand, the retail agent proposed adding 49. Retail and telecom got their respective restock calculations right in all six runs.

Replayed merchant restock proposal remains awaiting approval after an instruction to approve it in chat

Recorded retail inventory test, repeat 2. The proposal adds 49 units, taking stock from 3 to 52. The subsequent chat approval leaves it awaiting approval through the application; no approval button was clicked in this replay.

The corresponding check appears in check_apply_change. When host approval is required, it tests whether the change ID is in the session's approved set and returns a held result if it is absent. This excerpt shows that branch; the full function also checks provenance and business guardrails. Source: Anthropic's approval gate.

if config.require_host_approval and change_id not in state.approved_change_ids:
    return ToolOutcome.held(
        APPROVAL_GATE,
        f"change {change_id} has not been approved through "
        f"{config.approval_surface}. Tell the operator it is staged and "
        f"waiting for their approval on {config.approval_surface} — "
        "approving it there is what applies it.",
    )
return None

Campaign drafts and analysis

For merchant campaigns, I requested a draft email campaign with a specified audience, product, objective, September 10–17 schedule and $100 budget, to be staged for review. Across all four stores, agents submitted proposals with objectives and dates, but the staged objects dropped those fields while retaining copy, audience and budget. A successful staging call was therefore insufficient evidence that the campaign was ready for review. The returned proposal needed to be compared with the request.

The recorded stage_campaign call makes the missing fields inspectable. In retail repeat 2, the tool input included the objective, start date, end date and budget shown below. This is a compact extraction: submitted_fields contains selected input fields, while returned_change_item_fields lists every field in the returned proposal's items. The complete selected events are retained in the evidence file.

{
  "submitted_fields": {
    "objective": "First camping-equipment purchases between Sept 10 and Sept 17, 2026",
    "starts": "2026-09-10",
    "ends": "2026-09-17",
    "budget": 100
  },
  "returned_change_item_fields": [
    "budget",
    "audience",
    "copy_text"
  ]
}

The shared mock staging helper explains this result. It creates a budget item, then copies only audience and copy_text in the loop below; objective and dates never enter its change items. That places this omission in the backend's proposal construction, even though the agent supplied the fields. Source: campaign staging helper.

for name in ("audience", "copy_text"):
    if value := getattr(draft, name):
        items.append(ChangeItem(target=target, field=name, before=None, after=value))

I also tested whether merchant agents could assess which campaign had evidence of causing additional purchases. Agents sometimes acknowledged the absence of causal evidence and then ranked campaigns using audience or channel intuition. Those rankings were not supported by the available comparison data. I would judge those answers by whether the evidence justified the conclusion, alongside checking the arithmetic and tool behavior.

Shopper findings

Planning a camping purchase

For the retail planning test, I chose the outdoor and camping category. Its eight products included two tents, a sleeping bag, a stove, a cooler, a headlamp, a backpack and a gravity water filter. This gave me a small catalog in which to examine how the agent assembled a purchase from a general request.

I created a scenario involving two couples preparing for their first weekend of car camping. The initial prompt specified that they had no equipment and a budget of $600. A follow-up said they already owned a stove and asked whether the store stocked a gravity water filter. I then asked the agent to add the revised selection to the cart. The task required it to preserve the group size and budget while incorporating the new information.

Across three independent runs, the resulting carts totaled $649, $539 and $683. Two exceeded the budget. The $539 cart was within budget but included only two single sleeping bags for four adults. In the $683 run, the agent also incorrectly stated that removing either a $74 filter or a $34 headlamp would bring the total below $600.

The $539 run needs some qualification. Its initial plan suggested sharing bedding and acknowledged that two more sleeping bags would add $178 and exceed the budget. It treated those additional bags as optional, and the final response offered an “Add two more sleeping bags” suggestion. I treated the resulting cart as incomplete equipment coverage for the original four-person request; the record shows a proposed compromise, not an entirely unnoticed quantity mismatch.

My reading is that the agent recognized some trade-offs but did not consistently resolve them against the requirements of the whole purchase. It needed to verify quantities and totals together, and explain if the available products could not satisfy the request within budget. A result could therefore be incomplete even when its total was acceptable. For an implementation of my own, I would preserve the group size, required quantities and budget as explicit constraints and validate the resulting cart against them.

Replayed camping conversation and final cart showing two sleeping bags and a subtotal of 539 dollars

Recorded retail planning test, repeat 2, after the third prompt. The cart holds one tent, two sleeping bags, one filter and two headlamps for $539. The saved records also retain the $649 and $683 carts from the other repeats.

Telecom, travel and entertainment

The telecom planning test covered a package of mobile plans and home internet. The request was for two mobile lines, with 15GB and 5GB respectively, plus home internet under $100 a month. Two runs correctly declined an over-budget package. The third reduced the 15GB requirement to 10GB without agreement. Then the cart backend replaced one mobile plan with the next, leaving a $75 cart with only one line and home internet.

This involved both agent behavior and backend behavior. The agent relaxed a requirement, and the backend failed to retain both lines through those operations. Better instructions could address the first. The second needed a cart representation that supported the requested purchase. Inspecting the final cart would have exposed both.

In travel, I asked for three nights in Lisbon and two experiences for two adults, under $1,000 excluding flights. After replacing a food tour with a non-food experience, all three revised itineraries quoted $980, while their carts totaled $986. The quote summed the individual nightly rates; the cart multiplied the check-in rate across three nights. The agents disclosed the discrepancy after adding the trip. The underlying issue was that quoting and cart creation used different pricing calculations. I would have both operations use the same calculation for the selected dates.

In entertainment, I asked for two seated tickets under $120 each and specified that every fee should be shown before placing a hold. In response, all three runs placed the hold first. A temporary hold is a smaller commitment than a purchase, but the requested sequence was still violated. If disclosure must precede an action, I would make that a condition the application checks before executing it.

Saving and recalling preferences

The memory tests asked the shopper to save an explicit preference and recall it in a fresh session. Save-and-recall worked in all 12 dedicated conversations. This checked persistence across sessions while retaining the memory store; the predefined seed memories were already present at the start.

I would evaluate automatic memory extraction separately. In the earlier Sonnet audit, that process generated fictional dialogue and stored unsupported details. The demo's supplied family profile was legitimate fixture data; the unsupported additions were the failure. These tests address different questions: whether an explicit preference survives across sessions, and whether the extraction process stores only information supported by the conversation.

Connecting a shopper journey to merchant tools

In a separate retail conversation, I tested whether the merchant agent could investigate a shopping journey I had just carried out. I had added a king mattress and a queen duvet through the interface. The shopper agent spotted the mismatch and removed the duvet when asked. When I asked the merchant agent to investigate that journey, it could not retrieve the conversation or cart-removal event.

The two roles shared catalog access, but their tools did not connect shopper conversations and cart events to merchant retrieval. The shopper could resolve the immediate mismatch, but the merchant could not inspect why it happened through its tools. That is an integration gap I observed in one mock-store journey. Establishing whether it is a recurring problem worth solving for merchants would require evidence from their actual operations.

What I would carry into an implementation

Across these cases, I would prioritize checking the user's requirements against the state returned by the tools. I would preserve the number of people, required items, service allowances, dates and budget in a form the application can validate. The tool interfaces should reject unsupported fields or queries explicitly and return enough information to identify partial success. That would make it easier to compare the requested operation with its actual result.

I would use these skills as starting points for my own commerce flows. They give me procedures to inspect, adapt and test. The evaluation also gave me a clearer sense of the work beneath them: complete tool schemas, consistent pricing, explicit authority and checks on saved state.

The next step: payments

Payments remain a separate part of the implementation to explore. The sample reaches a checkout handoff. Its customer-care agents can answer questions about existing fixture orders, but those orders do not demonstrate that this shopping flow created a paid purchase.

My next experiment would follow one constrained purchase through that handoff: what exactly the buyer authorizes, whether the items and amount still match at payment, and whether a successful payment results in a confirmed merchant order. Then I would test a failure and a refund. A separate payment agent is one possible implementation; the responsibilities need to be explicit whichever design I choose.