---
title: Making your catalog discoverable to AI agents
description: Agentic commerce readiness starts with the initial HTML response. Server-render your product catalog, ship JSON-LD, and verify AI agents can read it.
url: "https://vercel.com/kb/guide/agentic-commerce-readiness"
published: 2026-10-01
last_updated: 2026-10-01
authors: Vercel
install_vercel_plugin: npx plugins add vercel/vercel-plugin
---

AI shopping agents read your product pages before any JavaScript runs. The crawlers behind them fetch your HTML and never execute it, so a storefront that assembles its catalog in the browser returns a shell with no product name, price, or availability in it.

Agentic commerce readiness means the first HTML response already states what the product is, what it costs, whether it's in stock, and who sells it. Here's how to get there, starting with the rendering decision everything else depends on.

## **What agentic commerce readiness means for your product catalog**

For a product catalog, readiness is decided by the first HTTP response. Only one of the two kinds of AI traffic hitting your storefront is limited to it.

Background crawlers like GPTBot, OAI-SearchBot, ClaudeBot, and PerplexityBot fetch HTML without executing it. Browser-based agents run a real browser, so they see what a person sees. Build for the crawlers, because anything they can read, the browser-based agents can read too.

Across Vercel's network in December 2024, GPTBot generated [569 million requests](https://vercel.com/blog/the-rise-of-the-ai-crawler) and Anthropic's Claude crawlers generated 370 million. Combined with AppleBot and PerplexityBot, those crawlers accounted for roughly 28% of Googlebot's volume over the same period.

Once a crawler reaches the page, it looks for structure in three places:

- [**Schema.org**](http://Schema.org) **Product markup:** A JSON-LD (JavaScript Object Notation for Linked Data) script tag in the initial HTML that names the product, price, availability, and seller.
  
- **Open Graph tags:** Metadata that agents and platforms extract to build previews.
  
- **A product feed:** Catalog data you push to a shopping surface rather than something the crawler finds on your site.
  

An `llms.txt` file is a lower priority than all three. Across [137,000 sites](https://ahrefs.com/blog/llmstxt-study), 97% of `llms.txt` files were never fetched at all. Agents that did reach one got there through links from other pages in [86% of attributed fetches](https://vercel.com/kb/guide/make-your-site-readable-by-ai-agents). Ship the HTML and schema layers first, then add a linked `llms.txt` if you want it.

## **Why rendering decides whether AI agents can read your catalog**

Structured data only reaches a crawler if it's in the response body. What a non-rendering crawler receives depends on your rendering strategy.

Here's what each strategy sends a crawler:

| Strategy                                                    | What the crawler receives                              | Agent-readable? |
| ----------------------------------------------------------- | ------------------------------------------------------ | --------------- |
| Static generation and Incremental Static Regeneration (ISR) | Complete pre-generated HTML, served from the CDN cache | Yes             |
| Server-side rendering (SSR)                                 | Complete server-rendered HTML, built per request       | Yes             |
| Partial Prerendering (PPR)                                  | A static shell, followed by streamed dynamic content   | Not by default  |
| Client-side rendering (CSR)                                 | An HTML shell and script tags, with no product data    | No              |

Under PPR, Next.js [checks the user agent](https://nextjs.org/docs/app/getting-started/caching) and serves a complete document to bots it recognizes. Its default list was built for search engines and social unfurlers, so GPTBot, ClaudeBot, PerplexityBot, and OAI-SearchBot don't match it and instead receive the streaming shell.

To bring AI crawlers into that blocking render, extend the list in `next.config.ts`:

```tsx
import type { NextConfig } from 'next'

const config: NextConfig = {
  htmlLimitedBots:
    /GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|PerplexityBot|[\w-]+-Google|Google-[\w-]+|Bingbot|applebot|Twitterbot|Slackbot|Discordbot|LinkedInBot|facebookexternalhit/,
}

export default config
```

Setting `htmlLimitedBots` replaces the default list rather than adding to it, so carry over the entries you still rely on. Because the shell is rebuilt instead of reused, a shell component that depends on build-time-only data can fail at request time, and a page that loads for a person can error for a crawler.

The same rendering change affects human search traffic. Ruggable rebuilt its frontend on Next.js and Vercel and saw a [300% increase](https://vercel.com/customers/how-ruggable-saw-more-organic-clicks-by-optimizing-their-frontend) in unbranded search clicks, and Plex refactored to SSR and on-demand ISR for a [6× increase in impressions](https://vercel.com/customers/how-plex-6x-their-impressions-deploying-next-js-on-vercel). Both came from the same mechanism that feeds AI crawlers.

## **An agentic commerce readiness checklist for Next.js storefronts**

With rendering settled, work through these five in order.

### **1\. Allow AI crawlers in robots.txt and confirm your firewall**

A `robots.txt` file that predates AI crawlers can block them through a catch-all disallow rule. The App Router's [robots.ts convention](https://nextjs.org/docs/app/api-reference/file-conventions/metadata/robots) makes the policy explicit and reviewable in code.

Declare the agents you want on your discovery routes:

```tsx
import type { MetadataRoute } from 'next'

export default function robots(): MetadataRoute.Robots {
  return {
    rules: [
      {
        userAgent: ['OAI-SearchBot', 'GPTBot', 'ClaudeBot', 'PerplexityBot'],
        allow: ['/products/', '/collections/'],
        disallow: ['/checkout', '/account', '/admin'],
      },
    ],
    sitemap: '<https://your-store.com/sitemap.xml>',
  }
}
```

`robots.txt` covers only one layer of access. Vercel's [AI bots managed ruleset](https://vercel.com/docs/bot-management) identifies known AI crawlers and can deny them at the firewall regardless of what `robots.txt` says. The ruleset is inactive by default, shown as **Allow** in the dashboard. Check that it isn't blocking crawlers from your catalog routes.

### **2\. Add JSON-LD structured data to every product page**

The JSON-LD script tag has to ship in the initial HTML. Injecting it after load keeps it from reaching any non-rendering crawler.

Render it inside the page component, following the [App Router pattern](https://nextjs.org/docs/app/guides/json-ld):

```tsx
export default async function Page({ params }) {
  const { id } = await params
  const product = await getProduct(id)

  const jsonLd = {
    '@context': 'https://schema.org',
    '@type': 'Product',
    name: product.name,
    description: product.description,
    image: product.image,
    sku: product.sku,
    brand: {
      '@type': 'Brand',
      name: product.brand,
    },
    offers: {
      '@type': 'Offer',
      price: product.price,
      priceCurrency: 'USD',
      availability: 'https://schema.org/InStock',
      priceValidUntil: product.priceValidUntil,
      seller: {
        '@type': 'Organization',
        name: 'Your Store',
      },
    },
  }

  return (
    <>
      <script
        type="application/ld+json"
        dangerouslySetInnerHTML={{
          __html: JSON.stringify(jsonLd).replace(/</g, '\\u003c'),
        }}
      />
      <ProductView product={product} />
    </>
  )
}
```

Type the object with `schema-dts`, and add `aggregateRating` where you have review data. The price in your markup has to match the price on the landing page and at checkout, because a mismatch costs you [merchant listing eligibility](https://support.google.com/merchants/answer/7052112).

### **3\. Configure metadata and Open Graph tags dynamically**

Product pages on parameterized routes should build their metadata from the same product record that feeds the schema.

Generate it with [generateMetadata](https://nextjs.org/docs/app/api-reference/functions/generate-metadata):

```tsx
export async function generateMetadata({ params }) {
  const { id } = await params
  const product = await getProduct(id)
  return {
    title: product.name,
    openGraph: { images: [product.image] },
  }
}
```

On parameterized routes, [metadata can stream in](https://nextjs.org/docs/app/getting-started/metadata-and-og-images) after the initial response. Prerendered pages avoid that because metadata resolves at build time, which is another argument for ISR on crawler-heavy paths. Emit the [required Open Graph properties](https://ogp.me/) (`og:title`, `og:type`, `og:image`, and `og:url`) plus `og:description`.

### **4\. Maintain a product feed for AI shopping surfaces**

ChatGPT Shopping reads a feed you push, not a crawl of your site. Feed onboarding is currently limited to approved merchants, so treat it as an application rather than a configuration change.

The [product feed spec](https://developers.openai.com/commerce/specs/file-upload/products) sets these requirements:

- **File format:** A UTF-8 `.txt`, `.tsv`, or `.csv` file, with Gzip-compressed versions also accepted. A feed already in Google's product data format uploads without renaming its columns.
  
- **Refresh cadence:** A full upload once a day, with changes sent through the API in between. Daily is the recommended baseline, not a hard limit.
  
- **Required fields:** Ten in total, including `item_id`, `title`, `price`, `availability`, `brand`, and `seller_name`. Setting `is_eligible_checkout` has no effect unless `is_eligible_search` is also `true`.
  

Google AI Mode draws on your [Merchant Center feed](https://support.google.com/merchants/answer/13889434?hl=en-GB), and the [Schema.org](http://Schema.org) markup from step 2 corroborates what the feed claims. For a custom XML feed, generate it from a Route Handler at request time with `Content-Type: text/xml`.

### **5\. Keep your sitemap synchronized with live catalog state**

A sitemap that lags your catalog sends crawlers to products you no longer sell, and AI crawlers already waste a large share of their fetches on 404s.

Generate it programmatically with `sitemap.ts`, and split large catalogs using [generateSitemaps](https://nextjs.org/docs/app/api-reference/functions/generate-sitemaps). For Shopify-backed stores, hook `revalidateTag` to the `products/create`, `products/update`, and `products/delete` webhooks so ISR pages track live inventory instead of the last deploy.

## **What blocks AI agents from reading your catalog**

Most catalogs that fail agent readability look correct in a browser, so the failure shows up only when you fetch the page as a crawler.

### **The AI bots ruleset set to Deny**

With the AI bots managed ruleset set to **Deny**, traffic identified as coming from AI bots is blocked, no matter what `robots.txt` allows. Leave it on **Allow**, or add a [bypass rule](https://vercel.com/docs/vercel-firewall/vercel-waf/custom-rules) so discovery routes stay reachable while the rest of the ruleset stays in force.

### **Client-side rendering reintroduced by personalization**

Personalization and A/B testing reintroduce client rendering long after the initial build. Helly Hansen's previous stack rendered personalization client-side, which blocked rendering and caused layout shifts. After migrating to ISR with the Vercel Data Cache, Cumulative Layout Shift spikes above 2.8 were virtually eliminated and Largest Contentful Paint [improved by 30%](https://vercel.com/customers/how-helly-hansen-migrated-to-vercel-and-drove-80-black-friday-growth) at launch.

### **Content behind accordions and read-more toggles**

A crawler that doesn't execute JavaScript can't reach text revealed by a click handler. Server-render the full description and collapse it with CSS instead.

### **A PPR shell with build-time-only dependencies**

The crawler-path re-render can throw while the same page renders correctly for a person. Audit shell components for anything that only exists during prerendering before you rely on PPR for catalog pages.

### **Feeds and schema drifting apart**

A price that differs between your markup, your feed, and Merchant Center costs you merchant listing eligibility. Trigger a feed refresh and ISR revalidation from the same webhook so the two can't diverge.

## **How to test agentic commerce readiness before you ship**

Run this pass before relying on any of the configuration.

Work through these seven checks in order:

1. **Confirm JSON-LD ships in the initial HTML:** Open `view-source:<https://your-store.com/products/example>` and search for `application/ld+json`. Absent means the markup is JavaScript-injected and invisible to crawlers.
   
2. **Validate the schema:** Run the URL through the [Schema.org](http://Schema.org) [Markup Validator](https://validator.schema.org/) for vocabulary errors, then the [Rich Results Test](https://search.google.com/test/rich-results) for merchant listing eligibility.
   
3. **Fetch as GPTBot:** Run `curl -A "GPTBot/1.2 (+https://openai.com/gptbot)" -I <https://your-store.com/products/example>` and expect a 200. A 403 or 429 points at a firewall or CDN rule, not `robots.txt`.
   
4. **Read the full body:** Run `curl -A "GPTBot" <https://your-store.com/products/example> > bot_view.html`, then confirm the name, price, availability, and description appear as readable text.
   
5. **Repeat for the other crawlers.** Use the official [Anthropic](https://privacy.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) and [Perplexity](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) user agent strings.
   
6. **Check Observability:** Open **Observability** and then **CDN Requests** to see whether GPTBot, OAI-SearchBot, ClaudeBot, and PerplexityBot are returning 200s on product routes. Bot and crawler breakdowns are available on all plans.
   
7. **Query the assistants:** Search ChatGPT and Perplexity for your brand plus a product name. Passing steps 1 through 6 and still being absent usually points to the feed rather than the page, and models also answer from cached data.
   

Steps 3 and 4 catch failures that don't appear in a browser, so run them on a schedule.

## **Shopify storefronts and AI agent discoverability on Vercel**

[AI-referred orders on Shopify](https://www.shopify.com/enterprise/blog/ai-search-insights) grew nearly 13× year over year in Q1 2026. A Shopify-backed storefront on Vercel follows the same five steps, with the product webhooks keeping ISR pages in sync with the catalog.

For a new build, `npx create-vercel-shop@latest my-store` scaffolds a Shopify storefront on Vercel. Non-Shopify storefronts can use the Next.js Commerce template as a reference architecture for the same work.

Browser-based agents like Perplexity's Comet and OpenAI's Atlas render pages the way a person's browser does. The server-rendered HTML that makes your catalog readable to GPTBot is the same HTML that those agents read.

## **Next steps**

Once your product pages render on the server and carry valid JSON-LD, the remaining work is keeping the feed and the cache in sync.

[Start a new project](https://vercel.com/new) for your storefront, or [browse the templates](https://vercel.com/templates) for a commerce-ready starting point.

## **Read more**

- [Make your site readable by AI agents](https://vercel.com/kb/guide/make-your-site-readable-by-ai-agents)
  
- [How AI is changing SEO](https://vercel.com/i/how-ai-is-changing-seo)
  
- [Bot Management](https://vercel.com/docs/bot-management)
  
- [Incremental Static Regeneration](https://vercel.com/docs/incremental-static-regeneration)
  
- [Firewall Observability](https://vercel.com/docs/vercel-firewall/firewall-observability)
  
- [How to prepare your storefront for Black Friday traffic](https://vercel.com/kb/guide/black-friday-preparation)
  

## Frequently asked questions

### Do AI crawlers respect robots.txt the same way Googlebot does?

GPTBot respects `robots.txt`, and ChatGPT search opt-outs use the OAI-SearchBot token. ChatGPT-User fires when a person asks ChatGPT to visit a specific page, so it behaves differently from an automated crawl. A firewall or CDN rule can block any of them, regardless of `robots.txt`, so check both layers.

### Does Partial Prerendering break agentic commerce readiness?

By default, yes, for AI crawlers. Next.js gives a blocking render only to user agents on its `htmlLimitedBots` list, which was built for search engines and social unfurlers rather than AI crawlers. Add the AI crawler tokens yourself, remembering that the setting replaces the default list rather than extending it.

### Does enabling Vercel's bot protection block AI crawlers from product pages?

No. The bot protection managed ruleset in **Challenge mode** automatically excludes verified bots, so crawlers in Vercel's verified bot directory aren't challenged. The separate **AI bots** managed ruleset is the one that blocks and is inactive by default. Check the status and action of both in your firewall settings.

### Is llms.txt worth adding to a Next.js storefront?

It's a low-priority layer. Across 137,000 sites, 97% of `llms.txt` files were never fetched, though agents reached the file reliably when other pages linked to it. Ship server-rendered HTML and JSON-LD first, then add a linked `llms.txt` if you want the extra surface.

### How do I confirm a crawler claiming to be GPTBot is really OpenAI's?

A user agent string proves nothing on its own, since it’s easy to tamper with. Verify the request against OpenAI's published crawler IP ranges, or rely on Vercel's verified bot directory, which checks IP ownership, reverse DNS, and cryptographic signatures before treating a bot as legitimate.

## More Next.js guides

- [Rendering strategies for retail: ISR vs SSR vs static](/kb/guide/ssr-vs-ssg-vs-isr-vs-csr-for-ecommerce): Choose between SSR, SSG, ISR, and CSR for each ecommerce surface. Compare rendering modes on freshness, personalization, SEO, and cost.
- [A/B testing on ecommerce storefronts without hurting Core Web Vitals](/kb/guide/ecommerce-ab-testing-core-web-vitals): Client-side A/B tools can worsen LCP and CLS on the pages you're testing. Run ecommerce A/B tests at the edge with Routing Middleware and Vercel Flags.
- [How to optimize Next.js and Sitecore Content SDK apps](/kb/guide/how-to-optimize-next.js-sitecore-content-sdk): Learn how to choose a rendering strategy, configure proxy middleware, and optimize costs for Next.js and Sitecore Content SDK apps on Vercel.