[Moonshot AI](/ai-gateway/models/labs/moonshotai)

# Kimi K2 Instruct

Kimi K2 Instruct is Moonshot AI's Mixture-of-Experts (MoE) language model with one trillion total parameters and 32 billion active per forward pass, a context window of 131.1K tokens, available through AI Gateway via Novita AI. Your use is subject to Moonshot AI's [Terms](https://platform.moonshot.ai/docs/agreement/modeluse.en-US) & [Privacy](https://platform.moonshot.ai/docs/agreement/userprivacy.en-US) Policies.

Tool Use

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'moonshotai/kimi-k2',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/kimi-k2) [API](/ai-gateway/models/kimi-k2/api) [About](/ai-gateway/models/kimi-k2/about) [Providers](/ai-gateway/models/kimi-k2/providers) [Throughput](/ai-gateway/models/kimi-k2/throughput) [Latency](/ai-gateway/models/kimi-k2/latency) [Uptime](/ai-gateway/models/kimi-k2/uptime) [Status](/ai-gateway/models/kimi-k2/status) [Similar](/ai-gateway/models/kimi-k2/similar) [FAQ](/ai-gateway/models/kimi-k2/faq)

## [Copy link to heading](#playground)Playground

Try out Kimi K2 Instruct by Moonshot AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75)Kimi K2 Instruct

![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=96&q=75)

Kimi K2 Instruct

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![novita logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnovita.png&w=48&q=75) [Novita AI](/ai-gateway/models/providers/novita) Legal:[Terms](https://novita.ai/legal/terms-of-service)•[Privacy](https://novita.ai/legal/privacy-policy) | 131K | 131K | 0.8s | 42tps | $0.57/M | $2.30/M |  | — |  |  |  | 07/11/2025 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-moonshot-ai)More models by Moonshot AI

All

Text

Code

Video Input

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi\-k3-fast](/ai-gateway/models/kimi-k3-fast) | 1M | 1.2s | 77tps | $4.50/M | $22.50/M | Read:$0.45/M Write:— | — | +2 | ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![morph logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmorph.png&w=48&q=75) |  |  | 07/27/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k3](/ai-gateway/models/kimi-k3) | 1M | 0.4s | 82tps | $2.90/MFast $4.50/M | $14/MFast $22.50/M | Read:$0.29/M Write:— | — | +2 | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![digitalocean logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdigitalocean.png%3Fv%3D1784246703962&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) +6 |  |  | 07/16/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.7-code-highspeed](/ai-gateway/models/kimi-k2.7-code-highspeed) | 262K | 0.3s | 218tps | $1.90/M | $8/M | Read:$0.38/M Write:— | — | +2 | ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) |  |  | 06/15/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.7-code](/ai-gateway/models/kimi-k2.7-code) | 262K | 0.5s | 185tps | $0.74/MFast $1.90/M | $3.50/MFast $8/M | Read:$0.15/M Write:— | — | +2 | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) +1 |  |  | 06/12/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.6](/ai-gateway/models/kimi-k2.6) | 262K | 0.4s | 81tps | $0.95/M | $4/M | Read:$0.16/M Write:— | — | +1 | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) +2 |  |  | 04/20/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.5](/ai-gateway/models/kimi-k2.5) | 262K | 0.5s | 92tps | $0.60/M | $3/M | Read:$0.1/M Write:— | — | +1 | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) ![novita logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnovita.png&w=48&q=75) |  |  | 01/26/2026 |  |

## [Copy link to heading](#about-kimi-k2-instruct)About Kimi K2 Instruct

Kimi K2 Instruct, released July 11, 2025, is a Mixture-of-Experts (MoE) language model from Moonshot AI.

**Sparse expert routing at 32B activation.** The full trillion parameters encode broad knowledge: programming languages, API conventions, domain facts, and tool-use patterns. At inference time, a routing mechanism selects roughly 32 billion parameters per token. Latency and compute cost stay comparable to a dense 32B model, while the knowledge base spans the entire trillion-parameter budget.

With 32B active parameters for reasoning depth and a full 1T parameter budget encoding broad tool-use and coding knowledge, K2 handles structured sequences of API calls, multi-step planning, and code synthesis.

Kimi K2 Instruct is available through AI Gateway at $0.57 per million input tokens and $2.3 per million output tokens.

AI Gateway routes K2 across Novita AI, giving you automatic failover across multiple providers.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: K2 routes across Novita AI. Choose it when uptime and provider redundancy matter most.
- Zero Data Retention: Zero Data Retention is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-kimi-k2-instruct)When to Use Kimi K2 Instruct

### Best for

- Agentic pipelines: Structured sequences of API calls, data processing, and code synthesis
- Provider redundancy: Deployments where failover across multiple providers matters most
- K2 architecture baseline: Teams evaluating the K2 architecture for the first time who want the original release
- Broad knowledge at low cost: Workloads that benefit from trillion-parameter knowledge breadth at 32B-dense inference economics

### Consider alternatives when

- Chain-of-thought traces: Kimi K2 Thinking layers extended reasoning on top of this foundation
- Minimum latency: Kimi K2 Turbo is the speed-optimized variant
- September 2025 checkpoint: Use Kimi K2-0905 for expanded context and refined agentic training
- Multimodal inputs: K2 processes text only, so reach for a vision-capable model

## [Copy link to heading](#conclusion)Conclusion

Kimi K2 Instruct established that sparse expert routing can deliver dense-model responsiveness at trillion-parameter scale. Its architecture anchors the entire K2 family of specialized variants. Routing across Novita AI gives you automatic failover for high-availability production.