MAI-Voice-2.1-Flash
MAI-Voice-2.1-Flash is a text-to-speech model built for fast, low-latency generation. It produces high- fidelity, natural, and expressive speech across 23 languages and supports gated instant voice cloning, all while being optimized for real-time responsiveness. Its human-like intonation, rhythm, and emotional nuance make it ideal for voice agents, assistants, and other interactive scenarios where latency and cost are critical. Your use is subject to Microsoft AI's Terms & Privacy Policies.
- Input price
- Input $15, Per 1M characters
import { experimental_generateSpeech as generateSpeech } from 'ai';import { gateway } from '@ai-sdk/gateway';import { writeFile } from 'node:fs/promises';
const result = await generateSpeech({ model: gateway.speechModel('microsoft/mai-voice-2.1-flash'), text: 'Hello from the Vercel AI Gateway!', voice: 'en-US-Harper',});
await writeFile('speech.mp3', result.audio.uint8Array);Copy link to headingPlayground
Try out MAI-Voice-2.1-Flash by Microsoft AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.
Your generated audio will appear here
Copy link to headingProviders
Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.
| Provider |
|---|