[Z.AI](/ai-gateway/models/labs/zai)

# GLM 5 Turbo

GLM 5 Turbo is the speed-optimized variant of Z.AI's GLM-5, released March 15, 2026. It trades some reasoning depth for faster throughput and lower latency while retaining GLM-5's multiple thinking modes and agentic capabilities. Your use is subject to Z.AI's [Terms](https://docs.z.ai/legal-agreement/terms-of-use) & [Privacy](https://docs.z.ai/legal-agreement/privacy-policy) Policies.

ReasoningTool UseImplicit Caching

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'zai/glm-5-turbo',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/glm-5-turbo) [API](/ai-gateway/models/glm-5-turbo/api) [About](/ai-gateway/models/glm-5-turbo/about) [Providers](/ai-gateway/models/glm-5-turbo/providers) [Throughput](/ai-gateway/models/glm-5-turbo/throughput) [Latency](/ai-gateway/models/glm-5-turbo/latency) [Uptime](/ai-gateway/models/glm-5-turbo/uptime) [Status](/ai-gateway/models/glm-5-turbo/status) [Similar](/ai-gateway/models/glm-5-turbo/similar) [FAQ](/ai-gateway/models/glm-5-turbo/faq)

## [Copy link to heading](#playground)Playground

Try out GLM 5 Turbo by Z.AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![zai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fzai-256.png%3Fv%3D1786428782581&w=48&q=75)GLM 5 Turbo

![zai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fzai-256.png%3Fv%3D1786428782581&w=96&q=75)

GLM 5 Turbo

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![zai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fzai-256.png%3Fv%3D1786428782581&w=48&q=75) [Z.AI](/ai-gateway/models/providers/zai) Legal:[Terms](https://docs.z.ai/legal-agreement/terms-of-use)•[Privacy](https://docs.z.ai/legal-agreement/privacy-policy) | 203K | 131K | 0.7s | 100tps | $1.20/M | $4/M | Read:$0.24/M Write:— | — |  |  |  |  | 03/15/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-zai)More models by Z.AI

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![zai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fzai-256.png%3Fv%3D1786428782581&w=48&q=75) [zai/glm-5.3-flash](/ai-gateway/models/glm-5.3-flash) | 1M | 0.5s | 126tps | $0.15/M$0.08/M | $0.50/M$0.25/M | Read:$0.03/M$0.01/M Write:— | — | +1 | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![digitalocean logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdigitalocean.png%3Fv%3D1784246703962&w=48&q=75) +11 |  |  |  | 08/26/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![zai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fzai-256.png%3Fv%3D1786428782581&w=48&q=75) [zai/glm-5.3](/ai-gateway/models/glm-5.3) | 1M | 0.3s | 192tps | $1.40/M | $4.40/M | Read:$0.14/M Write:— | — |  | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![digitalocean logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdigitalocean.png%3Fv%3D1784246703962&w=48&q=75) +7 |  |  |  | 08/18/2026 |  |
| ![zai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fzai-256.png%3Fv%3D1786428782581&w=48&q=75) [zai/glm-5.2-fast](/ai-gateway/models/glm-5.2-fast) | 1M | 0.5s | 221tps | $2.10/M | $6.60/M | Read:$0.21/M Write:— | — |  | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) |  |  |  | 06/23/2026 |  |
| ![zai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fzai-256.png%3Fv%3D1786428782581&w=48&q=75) [zai/glm-5.2](/ai-gateway/models/glm-5.2) | 1M | 0.2s | 345tps | $0.70/MFast $2.10/M | $2.20/MFast $6.60/M | Read:$0.11/M Write:— | — |  | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![crusoe logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fcrusoe.png&w=48&q=75) +15 |  |  |  | 06/16/2026 |  |
| ![zai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fzai-256.png%3Fv%3D1786428782581&w=48&q=75) [zai/glm-5](/ai-gateway/models/glm-5) | 203K | 0.7s | 130tps | $0.80/M | $2.56/M | Read:$0.16/M Write:— | — |  | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![novita logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnovita.png&w=48&q=75) +1 |  |  |  | 02/12/2026 |  |
| ![zai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fzai-256.png%3Fv%3D1786428782581&w=48&q=75) [zai/glm-4.7-flashx](/ai-gateway/models/glm-4.7-flashx) | 200K | 0.7s | 121tps | $0.06/M | $0.40/M | Read:$0.01/M Write:— | — |  | ![zai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fzai-256.png%3Fv%3D1786428782581&w=48&q=75) |  |  |  | 01/19/2026 |  |

## [Copy link to heading](#about-glm-5-turbo)About GLM 5 Turbo

GLM 5 Turbo was released March 15, 2026 as the speed-optimized variant in Z.AI's GLM-5 generation. The GLM-5 generation introduced selectable thinking modes so you can dial reasoning depth per request, and GLM 5 Turbo makes that capability affordable at production scale.

Agentic pipelines benefit the most. Many pipeline steps don't require the full GLM-5's deliberation depth, but they do benefit from the structured thinking modes when problems get harder. GLM 5 Turbo lets you route routine steps to a lightweight thinking mode for fast responses, then escalate harder steps to a deeper mode, all within the same model and API call format.

The turbo variant also inherits GLM-5's improved long-range planning and agentic coding capabilities. Combined with the lower per-token cost and faster throughput, this makes it practical to run multi-step agent workflows that would be prohibitively expensive at full GLM-5 pricing. Through AI Gateway, GLM 5 Turbo shares the same API surface as GLM-5.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: GLM 5 Turbo supports multiple thinking modes like GLM-5. Match the mode to the task: lightweight modes for extraction and classification, deeper modes for multi-step reasoning. Even deeper modes run faster than the equivalent on the full GLM-5.
- Configuration: For the most complex reasoning chains, the full GLM-5 will produce higher-quality results. Benchmark both on your hardest tasks to quantify the difference.
- Configuration: Switching between GLM-5 and GLM 5 Turbo requires only changing the model identifier. No integration changes needed.
- Zero Data Retention: Zero Data Retention is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-glm-5-turbo)When to Use GLM 5 Turbo

### Best for

- High-volume agentic pipelines: Most steps need GLM-5-class capability at lower latency and cost
- Structured data extraction: Documents where speed matters as much as accuracy
- Real-time coding assistance: Fast responses improve developer productivity without sacrificing agentic capabilities
- Production deployments at scale: Per-token cost directly impacts margins
- Multi-step workflows: Fast execution steps on GLM 5 Turbo pair with complex reasoning steps on the full GLM-5

### Consider alternatives when

- Maximum reasoning depth: The full GLM-5 provides the deepest deliberation in the generation on every request
- Vision or multimodal input: GLM-5V-Turbo adds image understanding to the turbo tier
- Frontend code focus: GLM-4.7 offers targeted frontend improvements at lower cost
- Absolute fastest inference: GLM-4.7-FlashX provides the lowest latency option when minimal capability is acceptable

## [Copy link to heading](#conclusion)Conclusion

Selectable thinking modes at production-friendly pricing make GLM 5 Turbo the practical entry point for teams adopting GLM-5 generation capabilities. Route agentic workflows through AI Gateway and scale between thinking depth levels per request.