Kat Coder Air V2.5
Kat Coder Air V2.5 is the fast-response tier of KwaiPilot's KAT-Coder V2.5 pair. It keeps the context window of 256K tokens, output up to 80K tokens, and full tool surface of the Pro tier at a lower price. Your use is subject to KwaiPilot's Terms & Privacy Policies.
View API reference- Input and output price
- Input $0.15, Output $0.60, Per 1M tokens
- 24h uptime
- Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({ model: 'kwaipilot/kat-coder-air-v2.5', prompt: 'Why is the sky blue?'})Copy link to headingPlayground
Try out Kat Coder Air V2.5 by KwaiPilot. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.
Kat Coder Air V2.5
Copy link to headingProviders
Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.
| Provider |
|---|
Copy link to headingUptime24 hours
Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.
Copy link to headingThroughput24 hours
P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.
Copy link to headingLatency24 hours
P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.
Getting started
Call Kat Coder Air V2.5 through AI Gateway with the AI SDK generateText and streamText functions, or through the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages APIs by changing the base URL. AI Gateway authenticates the request and routes it to an available provider.
Install the AI SDK (pnpm add ai dotenv), create an API key from the API Keys page, and set it as AI_GATEWAY_API_KEY in your environment. Full setup is covered in the text generation quickstart.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'kwaipilot/kat-coder-air-v2.5', prompt: 'Why is the sky blue?', });
console.log(result.text);}
main().catch(console.error);Top-level parameters
The same Kat Coder Air V2.5 request in each API format AI Gateway supports.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'kwaipilot/kat-coder-air-v2.5', system: 'You are a concise technical assistant.', prompt: 'Summarize the tradeoffs between static generation and SSR.', maxOutputTokens: 1024, });
console.log(result.text);}
main().catch(console.error);Standard parameters like prompt, messages, temperature, and tools work as documented in the AI SDK docs. These are the parameters with model-specific behavior.
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID in the form creator/model, e.g. kwaipilot/kat-coder-air-v2.5. AI Gateway routes the request to an available provider. |
maxOutputTokens | number | No | Hard cap on generated tokens. Kat Coder Air V2.5 supports up to 80,000 output tokens. Reasoning tokens count toward this limit. |
reasoning | 'provider-default' | 'none' | 'minimal' | 'low' | 'medium' | 'high' | 'xhigh' | No | Provider-agnostic reasoning effort, available in AI SDK 7 or later. Maps to the provider’s native reasoning configuration; reasoning settings under providerOptions take precedence when both are set. See the Reasoning section below. |
providerOptions | Record<string, JSONValue> | No | AI Gateway routing options under gateway, plus any provider-native options under the provider’s own namespace — see the table below. |
Input limits
| Input | Formats | Sources | Max count | Max size | Limits |
|---|---|---|---|---|---|
| Text | — | — | — | — | Prompt and response share the 256K-token context window |
| Image | — | URL, base64, Uint8Array | — | — | Sent as image parts in messages; counts as input tokens |
Provider options
Set AI Gateway routing options under providerOptions.gateway. For provider-specific options, pass them under the provider’s namespace as documented by the AI SDK.
Learn more in the AI SDK provider docs.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'kwaipilot/kat-coder-air-v2.5', prompt: 'Why is the sky blue?', providerOptions: { gateway: { only: ['streamlake'], }, }, });
console.log(result.text);}
main().catch(console.error);These AI Gateway routing options apply to every model. Provider-specific options pass through under the provider’s own namespace (for example providerOptions.anthropic) exactly as documented by the AI SDK.
| Parameter | Type | Required | Description |
|---|---|---|---|
providerOptions.gateway.only | string[] | No | Restrict routing to these provider slugs. Requests fail over only within the listed providers. |
providerOptions.gateway.order | string[] | No | Preferred provider order. Listed providers are tried first; unlisted providers remain available as fallbacks. |
providerOptions.gateway.sort | 'cost' | 'ttft' | 'tps' | No | Rank candidate providers by price, time to first token, or tokens per second instead of the default routing order. |
providerOptions.gateway.zeroDataRetention | boolean | No | Route only to providers with a zero-data-retention policy for this model. |
Routing across providers
AI Gateway serves the same model through multiple providers and fails over automatically. order expresses a preference while keeping every provider eligible; only is a hard allowlist — if none of the listed providers are available the request fails instead of falling back.
Options under a provider's own namespace (for example providerOptions.anthropic) are forwarded to that provider with the request. Providers ignore option namespaces that don't apply to them, so it is safe to set provider options alongside gateway routing options.
Reasoning
AI Gateway bridges reasoning across every API format. The AI SDK exposes a provider-agnostic top-level reasoning level (none, minimal, low, medium, high, or xhigh); the Chat Completions and Responses formats take the same effort under reasoning.effort; and the Anthropic Messages format uses a native thinking token budget. Whichever you send, the gateway maps it to the target model’s native configuration, converting between effort levels and token budgets as needed. Reasoning-related settings under providerOptions take full precedence over the top-level reasoning value and are never merged. Reasoning tokens typically count toward your output-token usage, though how they’re reported and billed varies by provider.
Learn more in the AI Gateway reasoning guide.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'kwaipilot/kat-coder-air-v2.5', prompt: 'Explain the Monty Hall problem step by step.', reasoning: 'high', });
console.log(result.text);}
main().catch(console.error);Image input
Send images alongside text as message parts. Images count as input tokens.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'kwaipilot/kat-coder-air-v2.5', messages: [ { role: 'user', content: [ { type: 'text', text: 'Describe this image.' }, { type: 'image', image: 'https://example.com/photo.jpg' }, ], }, ], });
console.log(result.text);}
main().catch(console.error);Tool calling
Expose tools the model can call. Define each tool’s inputs with a Zod schema.
import { generateText, tool } from 'ai';import { z } from 'zod';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'kwaipilot/kat-coder-air-v2.5', prompt: 'What is the weather in San Francisco?', tools: { getWeather: tool({ description: 'Get the current weather for a location', inputSchema: z.object({ location: z.string() }), execute: async ({ location }) => ({ location, temperatureC: 18 }), }), }, });
console.log(result.text);}
main().catch(console.error);Copy link to headingAbout Kat Coder Air V2.5
Kat Coder Air V2.5 shipped alongside kat-coder-pro-v2.5 on July 10, 2026 as the fast-response option in the KAT-Coder V2.5 pair. Both models target the same job: acting inside real, executable repositories rather than generating code in a single turn. Kat Coder Air V2.5 handles issue localization, code modification, and test execution as steps in one end-to-end loop.
The capability gap between the two tiers is published rather than implied. KwaiPilot's model comparison lists a SWE-bench figure of 42.4% for Kat Coder Air V2.5 against 65.2% for kat-coder-pro-v2.5. Everything else in that comparison matches: a context window of 256K tokens, output up to 80K tokens, streaming output, context caching, MCP (Model Context Protocol), function calling, and coverage of more than 20 mainstream programming languages. KwaiPilot recommends Kat Coder Air V2.5 for fast-response agent workloads, including native OpenClaw support, and reserves the Pro tier for complex enterprise projects and SaaS integrations.
That split maps onto how coding agents actually run. Routine work dominates request volume: triaging an issue, running a test suite, applying a small patch, summarizing a diff. Send those to Kat Coder Air V2.5 and escalate to kat-coder-pro-v2.5 when a task stalls. Cached input reads bill at $0.03 per million tokens, which keeps resent repository context cheap across long sessions. Route both tiers through AI Gateway via StreamLake and switch with a model identifier change. Product details sit at .
Copy link to headingWhat To Consider When Choosing a Provider
- Configuration: Kat Coder Air V2.5 and
kat-coder-pro-v2.5differ in capability, not in limits. They share a context window of 256K tokens, an output ceiling of 80K tokens, and the same tool surface. The decision rests on whether your tasks need the Pro tier's higher SWE-bench result. - Configuration: Run both against a sample of your own issues before you route production traffic. A cheaper model that loops on failed repair attempts can cost more per resolved task than a more capable one, so measure resolution rate rather than per-token price alone.
- Configuration: Context caching helps most in multi-turn agent loops that resend the same files every turn. Check the pricing panel on this page for current input, output, and cached read rates.
- Zero Data Retention: Zero Data Retention is offered on a per-provider and model basis. See the documentation for details.
- Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.
Copy link to headingWhen to Use Kat Coder Air V2.5
Best for
- High-Volume Coding Agents: Request count drives the bill more than any single task
- Issue Triage and Tests: Localize a problem, run the suite, and report the result
- CI Automation: Repeated tasks of the same shape across many pull requests
- Interactive Development Loops: A window of 256K tokens holds the working set across turns
- Tiered Routing: Routine work stays here and hard tasks escalate to the Pro tier
Consider alternatives when
- Complex Repository Work:
kat-coder-pro-v2.5posts a higher SWE-bench figure on hard tasks - Multimodal Input Needed: Screenshots and diagrams require a model that accepts image input
- General-Purpose Work: Writing and open-ended analysis sit outside a coding model's scope
- Existing V2 Deployments:
kat-coder-pro-v2may already cover your context and output needs
Copy link to headingConclusion
Kat Coder Air V2.5 gives you the KAT-Coder V2.5 tool surface and context window at the lower price point of the pair. Use it for the routine majority of agent traffic, keep kat-coder-pro-v2.5 for the hard tasks, and route both through AI Gateway so switching tiers is a model identifier change.
Copy link to headingFrequently Asked Questions
How does Kat Coder Air V2.5 compare to KAT-Coder Pro V2.5?
KwaiPilot's model comparison lists a SWE-bench figure of 42.4% for Kat Coder Air V2.5 and 65.2% for
kat-coder-pro-v2.5. The two tiers share a context window of 256K tokens, output up to 80K tokens, context caching, function calling, and MCP support, so the difference is capability rather than limits.What is the context window and maximum output for Kat Coder Air V2.5?
256K tokens of context and up to 80K tokens of output per request, the same limits as
kat-coder-pro-v2.5.What workloads does KwaiPilot recommend for Kat Coder Air V2.5?
Fast-response agent workloads. KwaiPilot lists native OpenClaw support for the KAT-Coder V2.5 pair and points complex enterprise projects and SaaS integrations at
kat-coder-pro-v2.5instead.Does Kat Coder Air V2.5 support function calling and MCP?
Yes. KwaiPilot lists streaming output, context caching, MCP (Model Context Protocol), and function calling for Kat Coder Air V2.5.
Which programming languages does Kat Coder Air V2.5 cover?
KwaiPilot lists more than 20 mainstream programming languages, including C, C++, Java, and Python.
How do I call Kat Coder Air V2.5 through AI Gateway?
Set the model identifier to
kwaipilot/kat-coder-air-v2.5and use your AI Gateway API key. Kat Coder Air V2.5 works through the AI SDK as well as the Chat Completions API and the other request formats AI Gateway accepts.What does Kat Coder Air V2.5 cost?
Current rates appear in the pricing panel on this page, including cached input reads at $0.03 per million tokens. AI Gateway reflects provider pricing, and StreamLake serves Kat Coder Air V2.5.
When was Kat Coder Air V2.5 released?
July 10, 2026, alongside
kat-coder-pro-v2.5.Is Zero Data Retention available for Kat Coder Air V2.5?
Zero Data Retention is not currently available for this model. Zero Data Retention is offered on a per-provider basis. See https://vercel.com/docs/ai-gateway/capabilities/zdr for details.