MAI-Voice-2.1
MAI-Voice-2.1 is our highest-fidelity, most expressive text-to-speech model, delivering rich, natural speech across 23 languages. It extends the MAI-Voice family with broad multilingual coverage, gated instant voice cloning, and strong long-form generation capabilities. With its detailed prosody, nuanced expressiveness, and studio-grade audio quality, MAI-Voice-2.1 is ideal for experiences where maximum voice quality and fidelity are required - long-form narration, brand-defining audio etc. Your use is subject to Microsoft AI's Terms & Privacy Policies.
- Input price
- Input $22, Per 1M characters
import { experimental_generateSpeech as generateSpeech } from 'ai';import { gateway } from '@ai-sdk/gateway';import { writeFile } from 'node:fs/promises';
const result = await generateSpeech({ model: gateway.speechModel('microsoft/mai-voice-2.1'), text: 'Hello from the Vercel AI Gateway!', voice: 'en-US-Harper',});
await writeFile('speech.mp3', result.audio.uint8Array);Copy link to headingPlayground
Try out MAI-Voice-2.1 by Microsoft AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.
Your generated audio will appear here
Copy link to headingProviders
Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.
| Provider |
|---|