Skip to content
Dashboard

Open-weight models take 56% of token volume, Astra doubles Fable 5.1 spend

AI Gateway Production Index — September 2026

Every month, AI Gateway routes tens of trillions of tokens between production applications and AI labs. That traffic gives us a view of what AI usage actually looks like in today's enterprise, and we publish it here monthly. See the Production Index reports from June, July, and August.

Copy link to headingSeptember 2026 summary

The September index reports on AI Gateway data collected through August 2026.

  • Open-weight models ran the majority of gateway tokens for the first time, up from 7% in December to 56% in August.

  • The average token costs less than half what it did five months ago. Price per token fell 23.2% in August, the third straight monthly drop, and the median team paid 7.6% less.

  • Fable 5, Anthropic's most capable model, lost two-thirds of its share of gateway spend in one month. Opus 5, at half the price, tripled its share. Anthropic kept 64% of all spend.

  • Gemini 3 Flash has lost 95% of its share of gateway tokens since May, and more than three-quarters of the volume it lost went to models from other labs.

Special report: OpenAI launches Astra on September 3

  • GPT-6 Astra took a third of OpenAI's spend within two days of launch and twice Fable 5.1's share of gateway spend. Introduced two days apart at the same price, Astra took 7.7% of all gateway spend in its first twelve days, while Fable 5.1 took 3.7%.

Copy link to headingOpen-weight models take a majority of token volume for the first time

In August, open-weight models ran 56% of all tokens on AI Gateway, marking the first month they took the majority of volume.

In December 2025, they processed fewer than one in ten tokens, and only eight months later, they ran more token volume than all closed-weight models combined.

Open-weight model token share rose every month from April through August, rising from 13% to 56% of total volume.

Though the frontier kept the majority of spend, open-weight dollar share is accelerating. As open-weight models become more capable, customers are moving more production workloads over to them.

Open-weight models processed 56% of August’s gateway tokens and accounted for 14% of spend.

Growth in open-weight model adoption helped push the average price per token across the gateway down 23.2% in August, its third consecutive monthly drop and the steepest since April. Among teams running more than ten million tokens in both months, the median team paid 7.6% less per token, more than double July's 2.9% decline.

Teams can now get more inference from the same budget and reserve frontier models only for the tasks that justify the premium.

Copy link to headingFrontier plateaus as Fable spend goes to Opus 5

Production workloads that justify a frontier model don't always need the most expensive one. They need one that’s good enough.

Fable is the most capable model Anthropic sells. Opus is the tier below it and costs roughly half of Fable’s price per token. When the US export control on Fable 5 was lifted and its access restored on July 1, its gateway spend share surged to 13.2%. At the end of that same month, Opus 5 came online.

In August, Fable 5’s share of gateway spend fell to 4.9%, and Opus 5's share rose to 22.5%. Nine in ten of the teams that ran Fable cut their usage, and more of them moved their workloads to Opus 5 than any other model. Fable’s extra capability wasn’t worth double the price.

Fable 5 fell from 13.2% of gateway spend in July to 4.9% in August as Opus 5's spend share rose to 22.5%.

Teams left Fable, Anthropic's most expensive and capable model, but the lab retained the lion’s share of gateway spend because those workloads stepped down to Opus 5, not a different lab.

Anthropic has taken at least 61 cents of every dollar spent through AI Gateway every month since December, and 64 cents in August. Its models have held the top two spots by spend every month since December, even as the models in those spots changed.

Anthropic has held the top two spots by spend every month since December. Google and OpenAI have each cracked third twice.

Copy link to headingCustomer loyalty follows the model profile, not the lab

Lab loyalty doesn’t follow brand, it follows model profile, and consistency wins.

When a new model preserves what users valued in its predecessor, the lab retains its customers. When it doesn’t, those customers fill the need through other providers.

When Claude Opus 5 launched, it gained almost twice what Fable lost, because it handled the same workloads at half the price. And within five days of Z.ai launching GLM-5.3-Flash, it was running three times GLM-5.2's daily volume.

GLM-5.3-Flash overtook GLM-5.2 one day after appearing on AI Gateway and processed two-thirds of Z.ai’s tokens by August 31.

Google struggled to retain customers with its new models. Because the new offerings didn’t provide a relative advantage on capability or price, a majority of Gemini 3 Flash’s workloads moved to OpenAI, Anthropic, and DeepSeek.

More than three-quarters of the volume Gemini 3 Flash lost went to other labs, and Google's own full-size Flash successors took under a tenth of it.

About half of the volume that left Gemini 3 Flash went to cheaper models, led by GPT-5.6 Luna, which costs less than half as much per token. Most of the other half went to higher-priced models, led by Claude Opus 5 and Sonnet 5, which cost roughly nine and three times as much as Gemini 3 Flash, respectively.

The flight to better-fit models meant that over the same period, Google’s share of gateway token volume fell from 30% to 5%, with Gemini 3 Flash accounting for 22 of the 25 percentage points lost.

Copy link to headingSpecial report: Astra took a third of OpenAI spend within 48 hours and outpaced Fable 5.1 two to one at launch

GPT-6 Astra launched on the AI Gateway on September 3 at the same price as Fable 5.1 and two and a half times the price of GPT-5.6 Sol. Two days later, it accounted for one in every three dollars spent on OpenAI models through the gateway. Its share of spend has held, hovering between 28% and 39% since.

Within OpenAI’s model lineup, Astra and Sol processed 27% of OpenAI’s tokens but accounted for 71% of its spending from September 4 through 16. Luna and Nano processed more than twice as many tokens for about one-ninth as much spending.

Luna processed more than eight times as many tokens as Astra, but Astra accounted for more than four times as much spend.

Anthropic launched Fable 5.1 on September 1, two days before Astra. Over each model's first twelve days on the gateway, Astra took 7.7% of all gateway spend, more than twice Fable 5.1's share of 3.7%, and was used by twice as many teams.

Astra passed Fable 5.1’s cumulative gateway spend on day four and reached twice Fable’s 12-day total by day 12.

OpenAI’s cheaper models carry its volume, while Astra’s early lead over Fable shows it can also attract teams at the highest price point. Together, they let OpenAI compete with other frontier labs for both scale and premium spend.

Stay tuned for more in next month's report.

Copy link to headingAlso in August’s data

  • Google's Nano Banana took the lead in image spend for the first time, at 50% to GPT Image's 44%, even as GPT Image took back the lead in images generated, 46% to 39%.

  • Google's Veo rose to second in video spend, at 20%, up from 15% in July. Seedance remained in first on both videos generated and video spend.

  • The share of videos generated by xAI’s Grok Imagine has more than halved since June, from 42% to 31% to 19%.

Copy link to headingPrevious AI Gateway Production Index reports

Copy link to headingAbout this report

This report uses anonymized, aggregate traffic routed through Vercel AI Gateway through August 2026.

A few notes on measurement:

  • Token volume includes input, output, reasoning, cached-input, and cache-creation tokens.

  • Spending is estimated using labs’ published list prices; actual bills may differ. Average price per token is estimated spending divided by token volume.

  • Statements about where volume moved compare changes among the same teams. They do not trace individual tokens between models.

  • Open-weight classifications follow the current AI Gateway model list, which is broader than the definition used in earlier reports.

  • All figures use the most recent data available; prior months may be revised as methodology is updated.

Ready to deploy?