MiMo V2.5 Pro
MiMo V2.5 Pro is the Pro tier of Xiaomi's MiMo v2.5 family, a Mixture-of-Experts (MoE) reasoning model built for agentic workflows, software engineering, and long-horizon tasks. It supports a context window of 1.1M tokens and 1.0M tokens max output tokens.
View API reference- Input and output price30% off
- Prices from: Input $0.30, Output $0.61, Per 1M tokens
- 24h uptime
- Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({ model: 'xiaomi/mimo-v2.5-pro', prompt: 'Why is the sky blue?'})Copy link to headingPlayground
Try out MiMo V2.5 Pro by Xiaomi. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.
MiMo V2.5 Pro
Copy link to headingProviders
Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.
| Provider |
|---|
Copy link to headingUptime24 hours
Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.
Copy link to headingThroughput24 hours
P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.
Copy link to headingLatency24 hours
P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.
Getting started
Call MiMo V2.5 Pro through AI Gateway with the AI SDK generateText and streamText functions, or through the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages APIs by changing the base URL. AI Gateway authenticates the request and routes it to an available provider.
Install the AI SDK (pnpm add ai dotenv), create an API key from the API Keys page, and set it as AI_GATEWAY_API_KEY in your environment. Full setup is covered in the text generation quickstart.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'xiaomi/mimo-v2.5-pro', prompt: 'Why is the sky blue?', });
console.log(result.text);}
main().catch(console.error);Top-level parameters
The same MiMo V2.5 Pro request in each API format AI Gateway supports.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'xiaomi/mimo-v2.5-pro', system: 'You are a concise technical assistant.', prompt: 'Summarize the tradeoffs between static generation and SSR.', maxOutputTokens: 1024, });
console.log(result.text);}
main().catch(console.error);Standard parameters like prompt, messages, temperature, and tools work as documented in the AI SDK docs. These are the parameters with model-specific behavior.
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID in the form creator/model, e.g. xiaomi/mimo-v2.5-pro. AI Gateway routes the request to an available provider. |
maxOutputTokens | number | No | Hard cap on generated tokens. MiMo V2.5 Pro supports up to 1,048,576 output tokens. Reasoning tokens count toward this limit. |
reasoning | 'provider-default' | 'none' | 'minimal' | 'low' | 'medium' | 'high' | 'xhigh' | No | Provider-agnostic reasoning effort, available in AI SDK 7 or later. Maps to the provider’s native reasoning configuration; reasoning settings under providerOptions take precedence when both are set. See the Reasoning section below. |
providerOptions | Record<string, JSONValue> | No | AI Gateway routing options under gateway, plus any provider-native options under the provider’s own namespace — see the table below. |
Input limits
| Input | Formats | Sources | Max count | Max size | Limits |
|---|---|---|---|---|---|
| Text | — | — | — | — | Prompt and response share the 1.1M-token context window |
Provider options
Set AI Gateway routing options under providerOptions.gateway. For provider-specific options, pass them under the provider’s namespace as documented by the AI SDK.
Learn more in the AI SDK provider docs.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'xiaomi/mimo-v2.5-pro', prompt: 'Why is the sky blue?', providerOptions: { gateway: { only: ['xiaomi', 'deepinfra'], }, }, });
console.log(result.text);}
main().catch(console.error);These AI Gateway routing options apply to every model. Provider-specific options pass through under the provider’s own namespace (for example providerOptions.anthropic) exactly as documented by the AI SDK.
| Parameter | Type | Required | Description |
|---|---|---|---|
providerOptions.gateway.only | string[] | No | Restrict routing to these provider slugs. Requests fail over only within the listed providers. |
providerOptions.gateway.order | string[] | No | Preferred provider order. Listed providers are tried first; unlisted providers remain available as fallbacks. |
providerOptions.gateway.sort | 'cost' | 'ttft' | 'tps' | No | Rank candidate providers by price, time to first token, or tokens per second instead of the default routing order. |
providerOptions.gateway.zeroDataRetention | boolean | No | Route only to providers with a zero-data-retention policy for this model. |
Routing across providers
AI Gateway serves the same model through multiple providers and fails over automatically. order expresses a preference while keeping every provider eligible; only is a hard allowlist — if none of the listed providers are available the request fails instead of falling back.
Options under a provider's own namespace (for example providerOptions.anthropic) are forwarded to that provider with the request. Providers ignore option namespaces that don't apply to them, so it is safe to set provider options alongside gateway routing options.
Reasoning
AI Gateway bridges reasoning across every API format. The AI SDK exposes a provider-agnostic top-level reasoning level (none, minimal, low, medium, high, or xhigh); the Chat Completions and Responses formats take the same effort under reasoning.effort; and the Anthropic Messages format uses a native thinking token budget. Whichever you send, the gateway maps it to the target model’s native configuration, converting between effort levels and token budgets as needed. Reasoning-related settings under providerOptions take full precedence over the top-level reasoning value and are never merged. Reasoning tokens typically count toward your output-token usage, though how they’re reported and billed varies by provider.
Learn more in the AI Gateway reasoning guide.
import { generateText } from 'ai';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'xiaomi/mimo-v2.5-pro', prompt: 'Explain the Monty Hall problem step by step.', reasoning: 'high', });
console.log(result.text);}
main().catch(console.error);Tool calling
Expose tools the model can call. Define each tool’s inputs with a Zod schema.
import { generateText, tool } from 'ai';import { z } from 'zod';import 'dotenv/config';
async function main() { const result = await generateText({ model: 'xiaomi/mimo-v2.5-pro', prompt: 'What is the weather in San Francisco?', tools: { getWeather: tool({ description: 'Get the current weather for a location', inputSchema: z.object({ location: z.string() }), execute: async ({ location }) => ({ location, temperatureC: 18 }), }), }, });
console.log(result.text);}
main().catch(console.error);Copy link to headingAbout MiMo V2.5 Pro
MiMo V2.5 Pro is the Pro variant in Xiaomi's MiMo v2.5 family, released April 22, 2026 under the MIT license. Compared to the standard tier, Pro activates a larger share of a larger parameter pool per token, which raises reasoning depth at higher per-token cost.
Like the rest of the line, MiMo V2.5 Pro uses a Mixture-of-Experts (MoE) stack with hybrid attention. Sliding-window and full attention combine in a fixed ratio, which cuts KV-cache storage versus dense attention at the same sequence length. Three multi-token prediction (MTP) blocks raise output tokens per inference step. The full window of 1.1M tokens fits long agent traces, repos, or document sets.
MiMo V2.5 Pro supports reasoning, tool calling, file input, vision, and implicit prompt caching. Call it through Xiaomi, DeepInfra, GMICloud via AI Gateway. For lower-cost everyday work, see mimo-v2.5.
Copy link to headingWhat To Consider When Choosing a Provider
- Configuration: MiMo V2.5 Pro sits at the Pro end of MiMo v2.5. Per-token cost is higher than
mimo-v2.5, but accuracy on hard math, agentic, and engineering tasks is the reason to pick it. Use AI Gateway's routing and fallback to send easy work tomimo-v2.5and reserve MiMo V2.5 Pro for the requests where reasoning depth pays off. - Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the documentation for details.
- Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.
Copy link to headingWhen to Use MiMo V2.5 Pro
Best for
- Long-Horizon Agents: Trajectories that span thousands of tool calls in a single run
- Complex Software Engineering: Issue resolution, repo-level edits, and multi-file refactors
- Math and Proofs: Long logical chains where intermediate reasoning steps matter
- Long-Context Reasoning: Documents and codebases approaching 1.1M tokens
- Pro-Tier MoE: Higher active-parameter compute for harder reasoning workloads
Consider alternatives when
- Throughput-Sensitive Workloads:
mimo-v2.5runs cheaper per token for everyday agent or code work - Short Prompt-and-Reply Calls: A smaller model is enough when you don't need deep reasoning
- Speed-First Tasks:
mimo-v2-flashfrom the previous generation is throughput-tuned - Simple Extraction Jobs: A lightweight model handles classification at lower cost
Copy link to headingConclusion
MiMo V2.5 Pro is the Pro pick in Xiaomi's MiMo v2.5 lineup. Use it for long-horizon agents, complex software engineering, and math-heavy reasoning. Pair it with mimo-v2.5 through AI Gateway routing so you can balance cost and quality across a mixed workload.
Copy link to headingFrequently Asked Questions
How does MiMo V2.5 Pro differ from
mimo-v2.5?It's the Pro tier. MiMo V2.5 Pro activates a larger share of a larger parameter pool per token than
mimo-v2.5, with higher per-token cost in return for stronger reasoning, code, and agentic scores.What architecture does MiMo V2.5 Pro use?
A Mixture-of-Experts (MoE) stack with hybrid attention. Each token activates a subset of expert blocks, and sliding-window plus full attention combine to keep KV-cache storage manageable across the 1.1M tokens window.
What's the context window for MiMo V2.5 Pro?
1.1M tokens. Hybrid attention keeps long-context runs practical, and multi-token prediction raises output tokens per inference step.
Does MiMo V2.5 Pro support tool calling and reasoning modes?
Yes. MiMo V2.5 Pro supports reasoning and tool calling, both exposed through AI Gateway. Use them through the AI SDK, the Chat Completions API, the Responses API, or any other supported format.
How do I authenticate requests to MiMo V2.5 Pro through AI Gateway?
Add your API key in AI Gateway project settings. Use
xiaomi/mimo-v2.5-proin API calls. AI Gateway routes, retries, and fails over acrossxiaomi,deepinfra,gmicloud.What does MiMo V2.5 Pro cost?
See the pricing section on this page for today's rates. AI Gateway tracks each provider's pricing for MiMo V2.5 Pro, so the numbers shown stay current.
Does MiMo V2.5 Pro support zero data retention?
Yes, Zero Data Retention is available for this model. Zero Data Retention is offered on a per-provider basis. See https://vercel.com/docs/ai-gateway/capabilities/zdr for details.
Can I route between MiMo V2.5 Pro and
mimo-v2.5automatically?Yes. AI Gateway supports fallback and routing. Send hard reasoning, code, and agentic requests to MiMo V2.5 Pro and fall back to
mimo-v2.5for simpler tasks to keep costs in check.