Commerce agents rarely fail on language. Every current frontier model can hold a shopping conversation. They fail on what the agent was allowed to fetch, mutate, and hand off, which makes a model benchmark the wrong place to open an architecture review.
Three decisions determine whether the agent ships: the channel it runs on, the tools it can call, and what happens at the edge of what it may answer. Model selection is a fourth, and the only one that is cheap to reverse.
The build runs in that order, ending on the two things that decide whether it survives real shoppers, which are what the agent is allowed to say and what happens when it cannot.
Key takeaways:
Conversational commerce spans on-site chat, messaging channels, voice, and agent-to-agent surfaces, and the channel sets the per-message cost and the compliance burden before any model does.
AI Gateway is the default provider in AI SDK 7, so one model string buys routing, ordered fallback, and spend budgets with no markup on tokens.
A commerce agent should never state a price, stock level, or returns rule it did not fetch, because a wrong answer is a liability a tribunal has already enforced.
Consumers expect a route to a human and rarely get a clean one, which makes handoff design more consequential than model choice.
Automated resolution rate and cost per resolution measure the agent honestly, while engagement counts a shopper who left frustrated as a success.
Copy link to headingWhat conversational commerce covers, and what it rules out
Conversational commerce is the model where AI advises inside the merchant's own surface and the buyer completes checkout in a browser. Two adjacent models get filed under the same name. In agentic commerce, the AI buys on the buyer's behalf through an API. In autonomous commerce, AI runs merchant-side pricing and inventory. The build below targets the first.
The channel covers chat-led support, in-chat discovery, purchase completion in a messaging thread, and voice. Agent-to-agent surfaces count too, under two competing protocols: the Agentic Commerce Protocol from OpenAI and Stripe, and the Universal Commerce Protocol from Shopify and Google.
Stateless FAQ bots and decision trees sit outside the category. They answer from a script, never from your catalog, which is why they don't carry the risks this guide is about, and don't carry the upside either. The line matters because it decides who owns the transaction. Keep checkout in the browser and you keep a liability you already understand.
Copy link to headingPick the conversational commerce channel before the model
Per-message cost and compliance obligations decide the channel, and both belong in the architecture review ahead of the system prompt. The constraint column is the one to read first, since a channel stops being an engineering decision there and becomes an approval or legal one.
On-site web chat is the only row with no per-message fee and no template approval, and one Next.js route handler serves the whole conversation. The cost is reach, since it only meets shoppers already on the storefront.
Copy link to headingThe architecture behind a conversational commerce agent
A commerce agent isn't a chatbot with the catalog pasted into its system prompt. Typed tools and retrieval supply every fact that can be wrong: price, inventory, discount validity, returns eligibility. Draw that line badly and every later decision inherits the problem, because a fact the model invented is one no downstream guardrail can catch.
Copy link to headingThe tool layer decides what the agent can say
The ToolLoopAgent class runs the tool execution loop and stops after 20 steps by default, through isStepCount(20). Commerce work maps cleanly onto tools: getProduct, checkInventory, addToCart, removeFromCart, lookupOrder, applyDiscount, and fetchPolicy.
Scope every one of them to the shopper's own session. lookupOrder should close over the authenticated customer and accept only an order that session already owns. A tool taking a customer identifier as an argument is a tool the model can be talked into pointing somewhere else. Vercel's agent security boundaries guidance puts it directly: an agent working for one customer gets a tool scoped to that customer's data, because a customer ID passed as an argument is subject to prompt injection. Cart tools bind to the session's cart on the same reasoning, and authorization lives in each tool's execute, never in the instruction string.
Define the agent with the tools it may call and an instruction that forbids unfetched facts:
import { ToolLoopAgent } from 'ai';
const agent = new ToolLoopAgent({ model: 'anthropic/claude-sonnet-5', instructions: 'You are a shopping assistant. Never state a price or stock level you did not fetch.', tools: { checkInventory: inventoryTool, applyDiscount: discountTool }, toolApproval: { applyDiscount: 'user-approval', },});
const result = await agent.generate({ prompt: shopperMessage });Define inventoryTool and discountTool with tool() and a schema over their inputs, and read shopperMessage from the request. The route handler wrapping this agent is a public POST endpoint. Authenticate and rate-limit it before the first tool runs, or cart mutation and order lookup are open to anyone who finds the URL.
Tool traffic is where the token spend has moved. Vercel's AI Gateway production index, published in May 2026 and covering traffic through April, put 58.9% of all tokens inside tool-call requests, up from 31.6% six months earlier. An agent built without a tool budget isn't budgeting for the agent it runs.
Some actions need a person. The toolApproval setting pauses execution until a decision returns, and it takes a per-tool policy: a status such as 'user-approval', or a function over the tool's parsed input. The function form is what lets small discounts pass while anything above your threshold waits.
Copy link to headingRetrieval for catalog and policy data
Tools cover facts you can look up by key. Retrieval covers the ones written in prose: return windows, warranty terms, sizing guidance, shipping exceptions. Use the language model middleware pattern with transformParams to inject retrieved chunks before each model call, and keep the embeddings in a vector store. A retrieval-augmented generation baseline starts at 500-token chunks with 50-token overlap, which holds for policy documents and needs tuning for product copy.
Validate what comes back before it reaches the shopper. Output.object() type-validates against your schema and rejects with AI_NoObjectGeneratedError when the model cannot produce a match, so the failure surfaces as an exception you can catch. Output.json() only confirms the text parses, which is the weaker guarantee and the wrong one for anything that quotes money.
Copy link to headingAI Gateway as the model layer
In AI SDK 7, AI Gateway is the default provider. A plain model string routes through it on Node.js 22 or later, with no provider client to configure. A storefront agent gets three things from that on day one.
Ordered fallback comes from a models array under providerOptions.gateway, which the gateway walks in order when the primary fails. The May 2026 index measured roughly 3.5% of requests on AI Gateway completing only after a fallback, which rescued 5.1% of token volume from errors, rate limits, and timeouts. On a conversation holding a cart, a dropped request is a lost sale.
Spend budgets apply at team, project, API key, and user scope, and they stack, so a request has to clear every budget in scope. They are soft caps: the request that crosses the limit still completes, and the next one returns HTTP 402. Per-request model selection comes from callOptionsSchema with prepareCall, which is how cart mutations reach a stronger model than product descriptions without running the whole conversation on the expensive tier.
Wiring all three into one agent loop is a longer build than this section covers, and third-party tool servers join that loop through createMCPClient from @ai-sdk/mcp.
Copy link to headingDeploying conversational commerce on storefront traffic
An agent turn is mostly waiting. The function sends a prompt, idles while tokens stream back, calls a tool, and idles again. Fluid compute fits that shape, because multiple invocations share a single instance and Active CPU pricing bills while your code executes and pauses while it waits on I/O. Traffic over the 2025 holiday peak reached 115.8 billion requests, up 33.6% year over year, on the same compute platform.
Copy link to headingStreaming and function duration
Function duration defaults to 300 seconds on every plan and reaches 800 seconds on Pro and Enterprise. An extended maximum of 1,800 seconds is in beta on those plans. Set it per function, never as a project default, on nodejs20.x, nodejs22.x, nodejs24.x, Bun 1.x or 1.4.x, or python3.12, python3.13, and python3.14.
Run the conversation on Node.js. Vercel Functions using the Edge Runtime have to begin responding within 25 seconds and stop streaming at 300. A chain of catalog and inventory calls can breach that before the cart exists. Vercel now recommends migrating off it, and Next.js 16.3 stopped supporting runtime = 'edge'.
streamText streams text and tool-call events from an App Router route handler, and useChat manages message state, status, stop, and regenerate on the client. Streaming custom data alongside the text delivers product cards and cart state as typed parts, so the client never parses them out of prose.
Long turns need a plan for the client dropping, and resumable streams split that work three ways. Your persistence layer records which stream belongs to which conversation, resumable-stream keeps the body in Redis, and a GET handler returns 204 when nothing is active. Without the first of those, the stream survives but the conversation loses track of it.
Copy link to headingSession and conversation state
A commerce conversation needs three stores, all connected through the Marketplace:
Each integration writes its connection credentials into the project as environment variables. Set those to the Secret type, which Vercel introduced in August 2026 for passwords, API keys, and tokens, so they stay unreadable after creation. A store credential has no reason to appear in a client bundle, and a NEXT_PUBLIC_ prefix on one ships it to every visitor who loads the page.
Persist history in the onEnd callback on toUIMessageStream, and validate anything arriving from a client with validateUIMessages or safeValidateUIMessages first.
Copy link to headingRegion and latency
Vercel Functions run in iad1 by default. A grounded answer makes several round trips before the first token reaches the shopper, each crossing the gap between function and data twice. Put Redis and Postgres in the function's region and that gap stops compounding. Pro supports up to five regions, and Enterprise adds functionFailoverRegions for multi-region failover.
Copy link to headingGrounding, guardrails, and the conversational commerce liability problem
A wrong answer from a commerce agent carries a real cost, and the case law is no longer hypothetical.
In Moffatt v. Air Canada, decided in February 2024, British Columbia's Civil Resolution Tribunal awarded 812.02 Canadian dollars after the airline's chatbot misstated bereavement fare policy. It rejected both defenses. The tribunal treated the chatbot as part of the airline's website and bound the airline to what it said. Correct policy text elsewhere on that site didn't excuse the answer the shopper got.
Two months earlier, a Chevrolet dealership's chatbot agreed to sell a 2024 Tahoe for $1 after a shopper reframed the conversation. These incidents illustrate different risks: misleading policy advice and manipulated price statements. Both require controls beyond a prompt.
Disclosure stopped being a courtesy in the EU when Article 50 of the AI Act applied from August 2, 2026. It requires telling people they are interacting with an AI unless that is obvious from context, an exception the European Commission reads narrowly. A chat widget styled to pass for a person is the case the rule was written for. Penalties for transparency breaches reach 15 million euros or 3% of worldwide annual turnover, which prices one line of disclosure well below the cost of omitting it.
Every one of these failures has a mechanism, which is the useful part, because a mechanism can be engineered against where a bad answer can only be apologized for:
Every row enforces outside the model, which is the only place enforcement holds. Vercel's own agent security guidance is blunt about why. You can't guarantee isolation between user input and the system prompt, and you can't expect the model to always follow the rules. An instruction telling the agent to refuse an override is worth writing. It's a preference, and the boundary has to sit somewhere the model can't reach.
None of this removes the need for a route out of the conversation. That is the guardrail teams most consistently underbuild, and the one that decides what a failure costs once the other six have missed it.
Copy link to headingDesign the handoff before you pick the model
Shoppers want an exit and rarely get a clean one. In SurveyMonkey's research on customer service expectations, 89% say companies should always offer the option to speak with a human. Twilio research cited in the same roundup puts 78% on wanting to switch from an AI agent to a person, and only 15% on having experienced a handoff that carried their context across. No benchmark score closes that gap.
Five triggers belong in the tool loop as conditions:
Intent-match failure: The agent cannot map the request to any tool it holds, or the same question returns a second time.
Explicit request: The shopper asks for a person, which ends the loop immediately.
Authority boundary: The action exceeds what the agent may approve alone.
Negative sentiment: Frustration signals appear, or billing disputes and account security come up.
Unresolved cart: Two or three exchanges pass without the cart moving.
Six things have to transfer when one of those fires: conversation history, verified identity, cart and checkout stage, prior agent actions, any expectation the agent set, and where in the flow the human picks up.
The WorkflowAgent from @ai-sdk/workflow is what keeps that state alive across a restart. Each tool call becomes a durable step that retries automatically, up to three attempts by default, without replaying work that already completed. A tool marked needsApproval suspends the workflow until a person responds, so a shopper who returns an hour later meets the state they left.
Copy link to headingWhat to measure once conversational commerce is live
Engagement counts activity. A shopper who leaves after eight frustrated messages still counts as an engaged session, which is why the number flatters the agent that earned it least. High engagement beside low resolution means the agent is adding effort while the dashboard looks healthy.
Five measures carry more signal:
Automated resolution rate: The share of issues fully solved by AI with no human involvement, which excludes shoppers who gave up.
Cost per resolution: Token, gateway, and compute spend divided by verified resolutions, so $250,000 across 50,000 resolutions reads as $5.00.
Assisted conversion rate: Sessions that used the agent, measured against sessions that did not.
Repeat contact rate: How often the same shopper returns with the same issue.
Abandonment at escalation: The share of shoppers who drop out at the handoff itself.
Cost per resolution is the number that decides whether the deployment survives its budget review. Gartner forecasts that generative AI cost per resolution will exceed $3 by 2030, above many offshore human agents.
Klarna's assistant is the cautionary case. OpenAI's case study credits it with an estimated $40 million profit improvement in 2024 and the equivalent work of 700 full-time agents. Klarna still reported $58 million in customer service and operations costs in the second quarter of 2026, against $51 million a year earlier, and doesn't say what drove the rise. Model price cuts alone won't carry these economics. Tool routing, model tiering, and prompt caching will.
AI Gateway request logs carry model, provider, tokens, and cost per request, and usage now aggregates across every step of a turn. Cost per resolved interaction becomes a query. Attributing it to one conversation needs custom reporting tags, billed separately from the zero-markup token pricing.
Copy link to headingHow Vercel supports conversational commerce at storefront scale
Copy link to headingTraffic that arrives in bursts
A storefront's conversation volume isn't a flat line. It spikes on a promotion, a stockout, or a shipping delay, so the capacity you provision for is the peak and the traffic you pay for should be the average. Fluid compute and Active CPU pricing close that gap, and the saving scales with how much of each turn the agent spends waiting.
Copy link to headingOne model string across every provider
Hardcoding a provider client looks free until an outage arrives mid-conversation with a cart attached, at which point recovery means shipping a deploy. AI Gateway moves that decision out of the codebase, so which provider a conversation runs on becomes a configuration question.
Copy link to headingConversations that outlive a single function
A shopper closing a laptop shouldn't cost you the cart. Resumable streams and WorkflowAgent between them keep the conversation's state outside the process running it, so the failure that ends a turn stops being the failure that ends the sale.
Copy link to headingData the agent is allowed to touch
Grounding needs somewhere to ground, and every retrieval hop is latency a shopper feels before the first token appears. Connecting the three stores through the Marketplace puts them under the same project and in the same region as the functions reading them, which is where that budget is spent or saved.
Copy link to headingShip an agent that survives its first wrong answer
The builds that reach production settle early what the agent may say, then wire the tools that enforce it. Define the tool layer first, so the agent can only state what it can fetch. Wire the gateway and the streaming path next. Set the escalation triggers before writing the first instruction. Swapping the model afterward is a one-line change, which is exactly why it's the wrong decision to start with.
Here is what Vercel handles, so the build stays on the parts specific to your catalog:
AI Gateway: Provider choice, failover, and spend limits become configuration, and per-request cost arrives with no markup on tokens.
Fluid compute and Active CPU pricing: Storefront spikes are absorbed without provisioning for the peak.
AI SDK:
ToolLoopAgentandWorkflowAgentcover the two loops a storefront needs, one fast and one durable, with human approval built into the second.Vercel Marketplace: The three stores a grounded conversation depends on, provisioned against the project and placed where the functions run.
Preview deployments: Every pull request gets a production-grade URL, so a change to the agent's instructions gets reviewed against real catalog data before it reaches shoppers.
Start a new project and put the tool layer in first, or browse templates for a storefront to build the conversation into.
Copy link to headingFrequently asked questions about conversational commerce
Copy link to headingWhen does a commerce agent need more than 800 seconds?
Rarely, and the beta carries a catch worth knowing first. While it runs, Secure Compute and static IPs do not support durations above 800 seconds. A conversation reaching private backends behind a firewall has to pick one.
Copy link to headingDoes EU AI Act Article 50 apply if the merchant is outside the EU?
Yes. The Act reaches providers and deployers established outside the EU whenever the system's output is used inside it. A US merchant shipping to EU shoppers is in scope on the strength of who reads the answer, wherever the function runs.
Copy link to headingWhat changes for WhatsApp commerce agents on October 1, 2026?
The first 1,000 service messages per business phone number each month stay free. Further service messages bill at that market's service rate, which Meta sets to match its utility and authentication rate. The allowance resets monthly and does not roll over.
Copy link to headingDoes Next.js Commerce include a conversational commerce agent?
No. Next.js Commerce is an App Router storefront for Shopify with no AI SDK integration, embeddings, or cart tools. A closer starting point bundles Shopify analytics, consent, and WebMCP tool registration as an opt-in.
Copy link to headingHow do you attribute AI spend to a single conversation?
Attach a user identifier and tags to each request through AI Gateway's custom reporting, then query the reporting endpoint grouped by user or tag. Reporting is billed separately from tokens, so budget for the writes and queries alongside inference.