[Moonshot AI](/ai-gateway/models/labs/moonshotai)

# Kimi K3 Fast

Kimi K3 Fast is the faster serving path for Moonshot AI's Kimi K3. It trades a higher per-token rate for lower latency, keeps the context window of 1M tokens, and routes through AI Gateway via Fireworks, Morph, Wafer. Your use is subject to Moonshot AI's [Terms](https://platform.moonshot.ai/docs/agreement/modeluse.en-US) & [Privacy](https://platform.moonshot.ai/docs/agreement/userprivacy.en-US) Policies.

ReasoningTool UseImplicit CachingFile InputVision (Image)

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'moonshotai/kimi-k3-fast',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/kimi-k3-fast) [API](/ai-gateway/models/kimi-k3-fast/api) [About](/ai-gateway/models/kimi-k3-fast/about) [Providers](/ai-gateway/models/kimi-k3-fast/providers) [Throughput](/ai-gateway/models/kimi-k3-fast/throughput) [Latency](/ai-gateway/models/kimi-k3-fast/latency) [Uptime](/ai-gateway/models/kimi-k3-fast/uptime) [Status](/ai-gateway/models/kimi-k3-fast/status) [Similar](/ai-gateway/models/kimi-k3-fast/similar) [FAQ](/ai-gateway/models/kimi-k3-fast/faq)

## [Copy link to heading](#playground)Playground

Try out Kimi K3 Fast by Moonshot AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75)Kimi K3 Fast

![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=96&q=75)

Kimi K3 Fast

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) [Fireworks](/ai-gateway/models/providers/fireworks) Legal:[Terms](https://fireworks.ai/terms-of-service)•[Privacy](https://fireworks.ai/privacy-policy) | 1M | 131K |  |  | $4.50/M | $22.50/M | Read:$0.45/M Write:— | — | +2 |  |  | 07/27/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![morph logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmorph.png&w=48&q=75) [Morph](/ai-gateway/models/providers/morph) Legal:[Terms](https://www.morphllm.com/privacy/tos)•[Privacy](https://www.morphllm.com/privacy) | 1M | 1M |  |  | $4.50/M | $22.50/M | Read:$0.45/M Write:— | — | +2 |  |  | 07/27/2026 |  |
| ![wafer logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fwafer.png&w=48&q=75) [Wafer](/ai-gateway/models/providers/wafer) Legal:[Terms](https://www.wafer.ai/terms)•[Privacy](https://www.wafer.ai/privacy-policy) | 1M | 131K |  |  | $4.50/M | $22.50/M | Read:$0.45/M Write:— | — | +1 |  |  | 07/27/2026 |  |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-moonshot-ai)More models by Moonshot AI

All

Text

Code

Video Input

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k3](/ai-gateway/models/kimi-k3) | 1M | 1.1s | 99tps | $2.90/MFast $4.50/M | $14/MFast $22.50/M | Read:$0.29/M Write:— | — | +2 | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![digitalocean logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdigitalocean.png%3Fv%3D1784246703962&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) +5 |  |  | 07/16/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.7-code-highspeed](/ai-gateway/models/kimi-k2.7-code-highspeed) | 262K | 0.1s | 131tps | $1.90/M | $8/M | Read:$0.38/M Write:— | — | +2 | ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) |  |  | 06/15/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.7-code](/ai-gateway/models/kimi-k2.7-code) | 262K | 0.5s | 190tps | $0.74/MFast $1.90/M | $3.50/MFast $8/M | Read:$0.15/M Write:— | — | +2 | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) +1 |  |  | 06/12/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.6](/ai-gateway/models/kimi-k2.6) | 262K | 0.3s | 493tps | $0.95/M | $4/M | Read:$0.16/M Write:— | — | +1 | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) +2 |  |  | 04/20/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.5](/ai-gateway/models/kimi-k2.5) | 262K | 0.6s | 44tps | $0.60/M | $3/M | Read:$0.1/M Write:— | — | +1 | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) ![novita logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnovita.png&w=48&q=75) |  |  | 01/26/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2-thinking](/ai-gateway/models/kimi-k2-thinking) | 216K | 0.8s | 39tps | $0.47/M | $2/M | Read:$0.14/M Write:— | — |  | ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) |  |  | 11/06/2025 |  |

## [Copy link to heading](#about-kimi-k3-fast)About Kimi K3 Fast

Kimi K3 Fast is the fast serving tier for Kimi K3, Moonshot AI's open-source model for long-horizon engineering work. Released July 27, 2026, Kimi K3 Fast trades a higher per-token cost for lower latency. The model's capabilities don't change.

Everything documented for the base model carries over. Kimi K3 Fast accepts text, image, and video input, supports a context window of 1M tokens, and keeps thinking mode always on. The strengths match too: long-horizon software engineering, knowledge work, deep reasoning, and tasks where code meets visual and spatial reasoning like frontend development, game development, and computer-aided design (CAD).

Two paths reach the fast tier. Set the `speed` option to `fast` while keeping the model ID on `moonshotai/kimi-k3`, and AI Gateway routes to the fast tier and falls back to standard speed when it isn't available. Or call `moonshotai/kimi-k3-fast` directly, which pins every request to the fast tier without that fallback.

Serving speed counts when someone is waiting. Interactive coding assistants stream output while a developer watches, and edit-run-fix loops repeat many times in a session, so per-turn latency compounds. Batch and offline jobs gain little. Throughput varies with load and context length, so read the live metrics on this page instead of planning around a fixed number.

AI Gateway routes Kimi K3 Fast across Fireworks, Morph, Wafer, including US-based providers for teams with data residency and compliance requirements. To keep inference in US data centers, set `inferenceRegion` with scope `zone` and geo region `us`. Zero Data Retention is configured per request with `zeroDataRetention` or turned on for every request from AI Gateway dashboard settings, and support varies by provider.

Use the identifier `moonshotai/kimi-k3-fast` with the AI SDK or any supported interface like Chat Completions, Responses, or Messages. Kimi K3 Fast supports a context window of 1M tokens and completions up to 1M tokens per request, at $4.5 per million input tokens and $22.5 per million output tokens.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: The fast tier lists above the base model, so compare $4.5 and $22.5 against Kimi K3 before you move high-volume traffic. Pinning `moonshotai/kimi-k3-fast` skips the fallback to standard speed that the `speed` option provides on the base model ID, so decide which behavior you want during a fast-tier shortage. Thinking mode stays always on, and those reasoning tokens still count toward the completion cap of 1M tokens. Measured throughput varies with load and context length, so use the live metrics on this page rather than a published figure.
- Zero Data Retention: AI Gateway supports Zero Data Retention for this model via direct gateway requests (BYOK is not included). To configure this, check the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr).
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-kimi-k3-fast)When to Use Kimi K3 Fast

### Best for

- Interactive Coding Assistants: Streaming output to a developer who watches the session
- Tight Iteration Loops: Edit-run-fix cycles where per-turn latency compounds across a session
- Latency-Budgeted Agents: Long-horizon agents that need Kimi K3 quality under a response-time cap
- Drop-In Speed Upgrade: Teams on Kimi K3 that want faster serving with the same behavior

### Consider alternatives when

- Cost-Sensitive Traffic: Kimi K3 serves the same model at the standard per-token rate
- Batch and Offline Jobs: Faster serving adds little when nobody waits on the output
- Automatic Speed Fallback: The `speed` option on the base model ID falls back to standard speed
- Optional Thinking Mode: A non-thinking Kimi K2 variant answers without a reasoning pass

## [Copy link to heading](#conclusion)Conclusion

Kimi K3 Fast gives you Kimi K3 with less waiting. When the output already meets your bar and response time is the remaining bottleneck, switch the model ID to `moonshotai/kimi-k3-fast` or set the `speed` option on the base model, then read the live metrics on this page to see what you gain.