Skip to content
Dashboard

OpenAI Ultrafast mode now available on AI Gateway

AI Gateway now supports OpenAI's Ultrafast service tier for GPT-6 Astra, providing faster output for interactive applications and rapid coding iterations.

To use Ultrafast, request it for openai/gpt-6-astra through AI SDK, the Chat Completions API, or Responses API:

import { generateText } from 'ai';
const { text } = await generateText({
model: 'openai/gpt-6-astra',
prompt: 'Investigate the failing tests and propose a fix.',
providerOptions: {
openai: { serviceTier: 'ultrafast' },
gateway: { only: ['openai'] },
},
});
console.log(text);

For workflows with frequent tool calls, OpenAI recommends the Responses API over a persistent WebSocket connection to reduce overhead between turns. See the Ultrafast service-tier examples for persistent connections, including AI SDK over WebSocket.

Ultrafast supports US and global processing. Requests pinned to unsupported regions, such as the EU, run at the standard (default) tier. Standard processing remains the default when no service tier is specified.

Requests served at Ultrafast are billed at 6× the standard per-token rate, while requests that fall back to another tier are billed at the rate for the tier actually served. Check the GPT-6 Astra model page for current rates.