[OpenAI](/ai-gateway/models/labs/openai)

# gpt-realtime-whisper

gpt-realtime-whisper is a streaming speech-to-text model for realtime transcription, returning transcript text while audio is still arriving, with a tunable latency and accuracy tradeoff and pricing based on audio duration rather than tokens. Your use is subject to OpenAI's [Terms](https://openai.com/policies/terms-of-use) & [Privacy](https://openai.com/policies/privacy-policy) Policies.

Websockets

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

```
1import { experimental_streamTranscribe as streamTranscribe } from 'ai';
2import { createGateway, gateway } from '@ai-sdk/gateway';
3import { readFile } from 'node:fs/promises';
4

5const modelId = 'openai/gpt-realtime-whisper';
6

7// Mint this on your server, then send only the short-lived token to the client.
8const { token } = await gateway.experimental_transcription.getToken({
9  model: modelId,
10});
11const clientGateway = createGateway({ apiKey: token });
12

13// Raw 24 kHz, 16-bit signed little-endian mono PCM audio.
14const bytes = await readFile('audio.pcm');
15const audio = new ReadableStream<Uint8Array>({
16  start(controller) {
17    controller.enqueue(new Uint8Array(bytes));
18    controller.close();
19  },
20});
21

22const result = streamTranscribe({
23  model: clientGateway.transcription(modelId),
24  audio,
25  inputAudioFormat: { type: 'audio/pcm', rate: 24000 },
26});
27

28for await (const part of result.fullStream) {
29  if (part.type === 'transcript-delta') {
30    process.stdout.write(part.delta);
31  }
32}
33

34console.log('\nFinal:', await result.text);
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/gpt-realtime-whisper) [About](/ai-gateway/models/gpt-realtime-whisper/about) [Providers](/ai-gateway/models/gpt-realtime-whisper/providers) [Similar](/ai-gateway/models/gpt-realtime-whisper/similar) [FAQ](/ai-gateway/models/gpt-realtime-whisper/faq)

## [Copy link to heading](#playground)Playground

Try out gpt-realtime-whisper by OpenAI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75)gpt-realtime-whisper

### Live transcription

Speak into your microphone and watch the transcript appear in real time.

Idle

![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=96&q=75)

Start the session and begin speaking to see the transcript.

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Input | Capabilities | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- |

| ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) [OpenAI](/ai-gateway/models/providers/openai) Legal:[Terms](https://openai.com/policies/terms-of-use)•[Privacy](https://openai.com/policies/privacy-policy) | $1.02/hr |  |  |  |  | 05/07/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#more-models-by-openai)More models by OpenAI

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) [openai/gpt-5.6-luna](/ai-gateway/models/gpt-5.6-luna) | 1.1M | 1.9s | 158tps | $1/M$0.20/M Fast $0.40/M+1 more | $6/M$1.20/M Fast $2.40/M+1 more | Read: $0.10/M$0.02/M+1 more Write: $1.25/M$0.25/M+1 more | $10/K \+ input costs | +4 | ![azure logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fazure.png&w=48&q=75) ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) |  |  |  | 07/09/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) [openai/gpt\-5.6-sol](/ai-gateway/models/gpt-5.6-sol) | 1.1M | 2.4s | 123tps | $4/M$2/M Fast $4/M+1 more | $20/M$10/M Fast $20/M+1 more | Read: $0.40/M$0.20/M+1 more Write: $5/M$2.50/M+1 more | $10/K \+ input costs | +4 | ![azure logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fazure.png&w=48&q=75) ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) |  |  |  | 07/09/2026 |  |
| ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) [openai/gpt-5.6-terra](/ai-gateway/models/gpt-5.6-terra) | 1.1M | 1.6s | 173tps | $2.50/M$2/M Fast $4/M+1 more | $15/M$12/M Fast $24/M+1 more | Read: $0.25/M$0.20/M+1 more Write: $3.13/M$2.50/M+1 more | $10/K \+ input costs | +4 | ![azure logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fazure.png&w=48&q=75) ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) |  |  |  | 07/09/2026 |  |
| ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) [openai/gpt-5.5](/ai-gateway/models/gpt-5.5) | 1M | 0.5s | 63tps | $5/MFast $12.50/M+1 more | $30/MFast $75/M+1 more | Read: $0.50/M+1 more Write: — | $10/K \+ input costs | +4 | ![azure logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fazure.png&w=48&q=75) ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) |  |  |  | 04/24/2026 |  |
| ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) [openai/gpt-5.4](/ai-gateway/models/gpt-5.4) | 1.1M | 0.8s | 82tps | $2.50/MFast $5/M+1 more | $15/MFast $30/M+1 more | Read: $0.25/M+1 more Write: — | $10/K \+ input costs | +4 | ![azure logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fazure.png&w=48&q=75) ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) |  |  |  | 03/05/2026 |  |
| ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) [openai/gpt-5-nano](/ai-gateway/models/gpt-5-nano) | 400K | 5.0s | 192tps | $0.05/M | $0.40/M | Read:$0.005/M Write:— | $14/K \+ input costs | +3 | ![azure logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fazure.png&w=48&q=75) ![openai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fopenai.png&w=48&q=75) |  |  |  | 08/07/2025 |  |

## [Copy link to heading](#about-gpt-realtime-whisper)About gpt-realtime-whisper

gpt-realtime-whisper arrived on May 7, 2026 as OpenAI's streaming speech-to-text model for realtime transcription. Whisper-generation models transcribe completed audio, which suits recordings and post-session processing. gpt-realtime-whisper is the streaming counterpart, built for live audio where the transcript has to appear during the conversation.

Latency and accuracy are tunable rather than fixed. A delay setting controls how much audio gpt-realtime-whisper hears before emitting text: lower values surface partial text sooner, and higher values give the model more context and improve transcript quality. You can also pass the languages you expect and vocabulary hints, so product names, acronyms, and identifiers come through correctly.

gpt-realtime-whisper accepts audio and text input and returns text, with a context window of 0 tokens and up to varies of transcript output per turn. Pricing is based on audio duration rather than text tokens, so cost tracks minutes of speech and forecasts cleanly from call volume. Current rates are listed on this page.

Audio support on AI Gateway is in beta through AI SDK 7. Your server mints a short-lived token and the client streams audio over a WebSocket connection, so your AI Gateway API key never reaches the browser. You get the same authentication, observability, and spend controls as your text models, with no markup on provider pricing.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: gpt-realtime-whisper runs in a persistent streaming session rather than a file upload, so plan for a connection that stays open for the length of the audio. Your interface also has to handle revision, because early partial text can change as more audio arrives.
- Configuration: A delay setting controls the tradeoff. Lower delay produces earlier text, and higher delay gives the model more audio context before it emits, which improves transcript quality. Tune it against production-like audio, including telephony, accents, background noise, and your own domain vocabulary, rather than clean samples.
- Zero Data Retention: Zero Data Retention is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-gpt-realtime-whisper)When to Use gpt-realtime-whisper

### Best for

- Live Captions: Conferences, webinars, and broadcasts that need text on screen as people speak
- Call Monitoring: Support and sales calls transcribed during the call for supervision or analytics
- Meeting Documentation: Notes captured while the conversation happens instead of afterward
- Voice Input Feedback: Interfaces that show users their words appearing as they speak
- Duration-Based Budgeting: Cost that tracks minutes of audio rather than token counts

### Consider alternatives when

- Recorded File Transcription: `whisper-1`, `gpt-4o-transcribe`, and `gpt-4o-mini-transcribe` handle completed audio
- Spoken Replies Required: The gpt-realtime family answers in speech rather than returning text only
- Speech Generation Needed: `tts-1` and `tts-1-hd` turn written text into spoken audio

## [Copy link to heading](#conclusion)Conclusion

gpt-realtime-whisper covers the transcription half of live audio: text as the words are spoken, a delay setting you tune to your product, and billing by the minute. Use it for captions, monitoring, and live documentation, and reach for the gpt-realtime voice models when the application also has to talk back.