Skip to content
Dashboard

Build a ChatGPT product feed for Shopify

ChatGPT sends a shopper to your product page. The agent reads the HTML, looks for a price, and finds a loading state instead. The answer it hands back has no price in it, and the sale goes to whichever store rendered its page on the server. Losing that sale costs more than most, because AI-referred shoppers convert 60% better than non-AI traffic on US retail sites.

The usual fix is to build a product feed. Most Shopify stores don't need one. If your products are eligible for Shopify Catalog and you sell to US buyers, Shopify already sends your catalog to ChatGPT, and no pipeline you write changes that. What nobody handles for you is how the product page renders when an agent arrives.

This guide covers what the ChatGPT product feed spec requires, why server-rendered product data decides whether GPTBot can read a page, and how to run a feed pipeline on Vercel when you do need one.

Key takeaways:

  • Shopify Catalog sends eligible products to ChatGPT for any store selling to US buyers, so most merchants need no feed at all.

  • The OpenAI product feed spec requires nine fields per row, and the barcode rule most summaries cite applies only to its Google-compatible upload format.

  • GPTBot downloads JavaScript without running it, so a price that loads in the browser never reaches ChatGPT's crawler.

  • Cron Jobs allow one run per day on Hobby and one per minute on Pro, so a feed refreshing faster than daily needs Pro or Enterprise.

  • Vercel verifies ChatGPT's signed requests, so its shopping agent passes the same firewall rule that blocks scrapers.

Copy link to headingWhat is a ChatGPT product feed?

A ChatGPT product feed is a structured file of product records that a merchant pushes to OpenAI over Secure File Transfer Protocol (SFTP). OpenAI ingests and indexes it so ChatGPT can surface those products in shopping conversations. It carries titles, descriptions, images, prices, availability, and seller context as data, in a fixed schema.

The feed has nothing to do with crawling. It's a push from merchant to OpenAI, so it works whether or not GPTBot ever fetches a product page. The separation cuts both ways. A complete feed gets your products into ChatGPT's answers, and it can't do anything about the page a shopper opens when they click one.

Copy link to headingWhat is the difference between Shopify Catalog and a direct ChatGPT product feed?

Shopify Catalog syndicates eligible products to ChatGPT with no app, no custom integration, and no fees beyond standard payment processing, while a direct feed is something a merchant applies for.

The two paths differ on setup, control, and what they cover:

Dimension

Shopify Catalog

Direct ChatGPT product feed

Coverage

Merchants with Catalog-eligible products selling to US buyers

Merchants approved by OpenAI after applying

Setup

Active by default for eligible stores

SFTP endpoint issued after approval

Engineering work

None

Feed generation, scheduling, and validation

Field control

Shopify maps the fields

Every field in the spec is yours to populate

Products outside Shopify

Not covered

Covered

A direct feed is worth applying for when part of your catalog lives outside Shopify, or when you need to set fields yourself rather than take Shopify's mapping. Every other store gets the same ChatGPT placement without writing anything.

Copy link to headingWhy most Shopify merchants don't need a product feed

The Agentic storefronts channel is active by default for eligible stores, and merchants manage it from Sales channels > Agentic in the Shopify admin. Four conditions govern eligibility for the ChatGPT channel:

  • US buyers: The store sells to customers in the United States, though the store itself can be based anywhere.

  • Catalog eligibility: Products qualify for Shopify Catalog, which is what supplies the product data to ChatGPT.

  • Completed store policies: Terms of service, privacy policy, and return and refund policy are filled in under Settings > Policies.

  • Agreed supplemental terms: The merchant accepts the Shopify Agentic Storefronts Supplemental Terms of Service.

Checkout stays on the merchant's side, since ChatGPT shoppers complete the purchase in the store's own checkout, opened in an in-app browser or a new tab. A headless Next.js storefront on Vercel reaches ChatGPT on exactly these terms, with the product page left for you to render.

Copy link to headingWhat the ChatGPT product feed spec requires

Nine fields are required per row, and every other field in the spec is optional or conditional. The nine required fields and the strictness of their validation determine how much of your catalog survives the first upload.

Copy link to headingWhat the nine required fields validate

Each required field carries a validation rule, and availability rejects rows without raising an error when a pipeline treats it casually:

Field

Constraint

item_id

Stable, unique per item or variant, never reused for a different item

title

Product name including the variant; aim for 150 characters or fewer.

description

Plain text, aim for 5,000 characters or fewer

url

Product detail page, publicly accessible, variant selected where possible

brand

Brand as shown on the product page

seller_name

Name of the seller supplying the offer

image_url

Direct image URL showing this variant, HTTP or HTTPS

availability

One of in_stock, out_of_stock, pre_order, backorder, or unknown

price

Amount and uppercase ISO 4217 code, for example 79.99 USD

An omitted, empty, or unrecognized availability value rejects the row outright, and unknown is a real value that has to be sent explicitly rather than left blank. Prices use major units with a decimal point, so 79.99 USD means 79 dollars and 99 cents.

Beyond the required nine, star_rating and review_count are the cheapest trust signal in the spec for a store already surfacing reviews. Variant grouping through group_id, listing_has_variations, and variant_dict costs more and matters more, because each variant ships as its own row with a distinct item_id and a shared group_id. Skipping it sends individual variant stock keeping units (SKUs) with no parent listing.

Copy link to headingWhere the Google-compatible path changes the rules

OpenAI accepts two upload formats and they validate differently. In the OpenAI format, the barcode fields gtin (Global Trade Item Number) and mpn are both optional, and identifier_exists isn't a field at all. The rule that every row needs one of those three belongs to the Google-compatible format, which accepts a feed already built to Google's schema.

Which format you pick changes what the pipeline has to validate. On the Google-compatible path a valid GTIN runs 8, 12, 13, or 14 digits with a valid check digit. Values beginning 02, 04, 2, 05, 98, or 99 are dropped rather than stored, and availability_date becomes required for preorder and backorder rows. Google-compatible feeds also use the registered merchant display name as the seller identity on every row, where the OpenAI format requires and uses per-row seller_name. One format applies to the entire upload rather than to individual rows, so the choice is made once.

Copy link to headingHow feeds are delivered and retained

Feed delivery runs over SFTP, to an endpoint OpenAI issues once it approves your application. Each upload replaces the whole catalog rather than sending only what changed, and daily is the recommended minimum. Parquet with zstd compression is the preferred format, with gzip-compressed JSONL, CSV, and TSV also accepted and XML excluded. Files stay UTF-8, cap at 500,000 items per shard, and target roughly 500 MB per shard file.

Retention is the constraint that catches pipelines out. A product missing from a processed snapshot isn't removed immediately, because OpenAI keeps its most recently processed record for up to 14 days. The window protects a catalog from a delayed shard, and it also means a discontinued product keeps appearing for two weeks unless the pipeline sets is_eligible_search=false instead of relying on omission. Use is_eligible_search=false in the OpenAI format. For Google-compatible feeds, confirm removal handling with OpenAI.

Copy link to headingWhy headless product pages go invisible to ChatGPT's crawler

If price and inventory are fetched only in the browser after the initial render, a crawler that does not execute JavaScript may see only the loading state.

ChatGPT's crawlers fetch JavaScript files and never execute them. Measured with MERJ across our network, GPTBot generated 569 million requests in one month, with JavaScript accounting for 11.50% of ChatGPT's fetches despite none of it running. AppleBot and Gemini render JavaScript through browser-based crawling, and ChatGPT's crawler does not. A Next.js product page pulling price or inventory from the Shopify Storefront API inside a Client Component therefore hands GPTBot the loading state and nothing else.

Copy link to headingWhat belongs in the server-rendered HTML

Anything an agent needs to evaluate or quote has to arrive in the initial server response, and content in that initial response can be read even as JSON data or delayed React Server Components. Four categories of product data belong there:

  • Price: The current price and any sale price, in the same currency the product page displays.

  • Availability: Stock status for the specific variant on the page, not a store-wide in-stock badge.

  • Variants: Each option, the value currently selected, and the price attached to it.

  • Page metadata: The title, description, category, and internal links a crawler follows to reach related products.

Everything a shopper reads before deciding belongs in that first response. View counters, live chat, and social feeds can still render in the browser. Adobe's retail sector visibility benchmark puts individual product pages at 66% machine-readable, the lowest score of any page type it measured, against 75% for homepages.

Copy link to headingWhy deploy-scoped asset URLs compound the 404 rate

ChatGPT's crawler spends 34.82% of its fetches on 404 pages against 8.22% for Googlebot, and most of those requests chase outdated assets under /static/. A crawler re-requesting asset paths from an earlier build wastes its budget on a site that redeploys frequently.

Version skew is the mechanic underneath. Chunk URLs change between deployments, so paths a crawler cached from a previous fetch stop resolving. Maintained redirects, current sitemaps, and consistent URL patterns bring that rate down, and the effect is larger on a headless storefront shipping several times a week than on a site deploying monthly.

Copy link to headingCore components of a ChatGPT product feed pipeline on Vercel

For merchants who do need a direct feed, the pipeline has three parts. Shopify is the source, a Route Handler generates the file, and Cron Jobs plus webhooks keep it current.

Copy link to headingChoosing the Shopify API for your catalog size

The Shopify Storefront API is the simplest source for a catalog under roughly 25,000 objects, and buyer traffic on it isn't rate-limited, though bot and crawler limits and a checkout throttle apply. Above that size the decision is made for you. Shopify caps pagination at 25,000 objects across its GraphQL APIs, and the cap applies to count queries too.

Larger catalogs move to the Admin API's bulkOperationRunQuery, which runs asynchronously and returns a JSONL file whose download URL expires after 7 days. The cost-based limit restores 100 points per second on standard plans and 1,000 on Plus. Variant modeling is what pushes a catalog over the line, since a 6,000-product store with four sizes and three colors per product produces a 72,000-row feed.

Both credentials belong on the server. Shopify's Storefront public access token is browser-safe by design, but nothing in the browser reads it here, and the Admin API token can edit orders and delete products. Neither takes a NEXT_PUBLIC_ prefix, which Next.js inlines into the JavaScript it sends to the browser.

Copy link to headingGenerating and scheduling the file

Generate each snapshot in a Route Handler using fresh Shopify data. Current Next.js Route Handlers are not cached by default, but explicitly cached Shopify reads can still return stale data. Check the route’s data-fetching and caching configuration rather than adding dynamic = 'force-dynamic' on every feed handler.

That route is a public URL. Vercel sends CRON_SECRET as a bearer token on every cron invocation, and the handler has to check it, or anyone who finds the path can trigger a full catalog export on demand:

import type { NextRequest } from 'next/server';
export async function GET(request: NextRequest) {
const authHeader = request.headers.get('authorization');
const cronSecret = process.env.CRON_SECRET;
if (!cronSecret || authHeader !== `Bearer ${cronSecret}`) {
return new Response('Unauthorized', { status: 401 });
}
// Build and write the snapshot here
return Response.json({ success: true });
}

The vercel-cron/1.0 user agent and the x-vercel-cron-schedule header are not a substitute, since any client can send both. Feed cadence then becomes a plan decision:

Plan

Minimum interval

Scheduling precision

Hobby

Once per day

Hourly, within a 59-minute window

Pro

Once per minute

Per minute

Enterprise

Once per minute

Per minute

On Hobby, a cron expression running more than once a day fails at deployment rather than at runtime. Three further behaviors shape the handler, since cron jobs trigger only against the production deployment URL, schedules are always in Coordinated Universal Time (UTC), and Vercel doesn't retry a failed invocation. Idempotency is therefore a requirement rather than a nicety, because recovery means rerunning the job by hand.

Copy link to headingKeeping the feed fresh between runs

Daily snapshots leave a gap that Shopify webhooks close. Next.js Commerce ships with six topics already wired to a revalidation Route Handler that calls revalidateTag() on the products and collections tags. The products/update topic carries more than its name suggests, including variant changes and the inventory movement behind purchases, so a separate inventory_levels/update subscription adds a second stock signal rather than filling a gap.

Every delivery needs its signature checked first. Shopify signs each one with HMAC-SHA256 over the raw request body, keyed on the app's client secret, and sends the digest base64-encoded in the X-Shopify-Hmac-Sha256 header. Skip that check and any POST reaching the URL triggers a rebuild, which lets an anonymous caller run up Incremental Static Regeneration (ISR) spend. The digest has to be computed over the untouched body, before any JSON parsing, and compared in constant time.

Copy link to headingFour failure modes to design around

Four things break a ChatGPT product feed without raising an error. Rows get dropped or stale data stays in place, and the run reports success either way.

Copy link to headingGuard against the webhook race condition

A products/update webhook can arrive before the changed data has finished propagating, so an immediate revalidateTag() locks in the old product. This is a documented issue in vercel/commerce, not an edge case developers invent for themselves. A short delay or a re-check interval before regenerating beats revalidating the moment the webhook lands.

Copy link to headingValidate identifiers before upload

On the Google-compatible path, a row with an invalid GTIN is dropped from the snapshot without a warning. The check digit computes during generation by multiplying alternate digits by 3 and 1 from the rightmost digit before the check digit, summing, and confirming the total divides by 10. Products with no assigned identifier take identifier_exists=no rather than an invented mpn to satisfy the validator.

Copy link to headingDrop sale_price when it doesn't undercut price

A sale price equal to, above, or in a different currency from the regular price goes unused, and on the Google-compatible path an invalid relationship rejects the row. Comparing the two values at generation time and omitting sale_price when the comparison fails handles the common case, which is a promotion whose discount expires while the field doesn't clear.

Copy link to headingLock concurrent runs on slow feeds

When a generation run outlasts its own interval, a second instance starts while the first is still writing. A distributed lock or a wider interval prevents the overlap, and a duration ceiling on the function stops a hung run from holding the lock indefinitely.

Copy link to headingHow Vercel supports a ChatGPT product feed and the pages agents land on

Four Vercel primitives cover most of the surface a headless commerce team ends up owning here.

Copy link to headingRendering product data that survives a no-JavaScript fetch

You can populate every field in the spec and still lose the sale to a price that hydrates client-side.

React Server Components and ISR put product facts in the initial HTML response without giving up dynamic pricing. Product data renders on the server, the crawler reads it as text, and cache tags keep it current when Shopify changes. To confirm this, fetch one of your own product pages with JavaScript disabled and look for price and availability in the source.

Copy link to headingRunning long feed generation without paying for wall-clock time

A full catalog export is mostly waiting. Bulk operations run asynchronously by design, so the function sits idle until the export is ready while the meter runs. That's what makes scheduled export jobs disproportionately expensive on conventional serverless pricing.

Fluid compute bills active CPU rather than wall-clock duration, so a 20-minute run spent mostly on input and output costs far less than its elapsed time suggests. For catalogs outgrowing the 300-second default, maxDuration extends to 800 seconds on Pro and Enterprise, and up to 1800 seconds through extended durations, currently in beta and configured per function.

Copy link to headingAdmitting the shopping agent while denying scrapers

The firewall rule keeping scrapers off product pages can also lock out the agent trying to buy, and that failure stays invisible until revenue moves. ChatGPT's shopping agent identifies as chatgpt-operator, and Vercel recognizes the traffic and verifies its signed requests automatically, so basic pass-through needs no configuration.

For selective control, checkBotId() returns verified-bot fields alongside isBot that make the allowlist explicit:

import { checkBotId } from 'botid/server';
import { NextResponse } from 'next/server';
export async function POST(request: Request) {
const { isBot, isVerifiedBot, verifiedBotName } = await checkBotId();
const isOperator = isVerifiedBot && verifiedBotName === 'chatgpt-operator';
if (isBot && !isOperator) {
return NextResponse.json({ error: 'Access denied' }, { status: 403 });
}
// Your handler continues here
}

The AI bots managed ruleset is free on all plans and off by default. Turning on deny mode blocks GPTBot and OAI-SearchBot, which is the right call if you want training crawlers out and the wrong one if you want ChatGPT to find your products. Either way, rate-limit the routes crawlers hit separately from the paths the shopping agent uses.

Copy link to headingTracing a stale feed to the cache serving it

When a snapshot contains yesterday’s catalog, trace the data through the pipeline: the Shopify API response, any cached reads, the generated file, and the latest successful upload. If you also serve a public feed endpoint, inspect its x-vercel-cache response header to check whether Vercel served a cached response.

The authenticated cron requests shown above include an Authorization header, which makes their function responses ineligible for Vercel CDN caching. Invalidate cached Shopify data when necessary. A separate CDN purge is relevant only if you also cache a public feed response and need to remove that cached copy.

Copy link to headingGet the crawl right before you build the feed

Scoping a feed pipeline first means solving the second problem. Catalog syndication is largely handled, which leaves the crawl as the work that's genuinely yours. The loading state an agent reads instead of a price is a rendering decision, and no platform makes it for you.

The payoff for getting the page right reaches past agents. PAIGE moved to headless Shopify with Next.js and Vercel and increased Black Friday revenue 22%, with a 76% increase in conversion rates. The lift came from the same server-rendering work that makes a product page readable to a crawler. How much of the surrounding pipeline you build yourself depends on what the platform underneath already handles:

  • Fluid compute: Active CPU billing and extended function durations, so full-catalog feed generation that waits on Shopify doesn't bill for idle time.

  • BotID: Verified-bot identification through checkBotId(), so chatgpt-operator passes the same rule that denies scrapers.

  • AI bots managed ruleset: A maintained list of AI crawlers with log or deny actions, free on all plans and updated as new crawlers appear.

  • Cron Jobs: Scheduled feed generation down to once per minute on Pro and Enterprise, running against the production deployment.

  • CDN cache controls: Inspect cache-status headers and invalidate cached public responses when needed, separately from the authenticated feed-generation job.

Start a new project to put these patterns in place, or browse the templates for a headless Shopify storefront already wired up for server-rendered product data.

Copy link to headingFrequently asked questions about ChatGPT product feeds

Copy link to headingDo Shopify merchants need to build a ChatGPT product feed?

No. Merchants with Shopify Catalog-eligible products selling to US buyers are syndicated automatically through agentic storefronts, with no application or pipeline required. A headless storefront still needs server-rendered product pages so GPTBot can read them after a shopper clicks through.

Copy link to headingWhat fields does the ChatGPT product feed spec require?

Nine per row: item_id, title, description, url, brand, seller_name, image_url, availability, and price. Everything else is optional or conditional. Ratings, review counts, and variant grouping are the optional fields most worth populating once the required nine validate cleanly.

Copy link to headingWhat file format should a ChatGPT product feed use?

Parquet with zstd compression is OpenAI's preferred format, with gzip-compressed JSONL, CSV, and TSV also supported and XML excluded. Filenames should stay identical between runs and be overwritten in place, since OpenAI matches products by item_id rather than by filename or shard position.

Copy link to headingWhy would a product page be invisible to GPTBot even if the URL is indexed?

GPTBot fetches JavaScript but doesn't execute it, so a URL can be crawled successfully and still yield an empty product record. Indexing confirms the crawler reached the page, not that it read anything usable there.

More Retail articles

Ready to deploy?