Skip to main content
The Quantum Dispatch
Back to Home
Cover illustration for Microsoft MAI Voice 2.1: Faster Speech for Voice Agents

Microsoft MAI Voice 2.1: Faster Speech for Voice Agents

Microsoft's MAI-Transcribe-2-Streaming returns first words in about 100ms across 60 languages for $0.54 an hour, alongside two new MAI-Voice-2.1 models.

Dr. Nova Chen
Dr. Nova Chen★Oct 5, 2026★3 min read

Microsoft AI released a new set of speech models on October 1, 2026: MAI-Transcribe-2-Streaming for turning speech into text, plus MAI-Voice-2.1 and MAI-Voice-2.1-Flash for generating speech. The headline is speed. If you are building a voice agent, the delay between a person finishing a sentence and the system understanding it is the thing that makes a conversation feel natural or awkward.

  • MAI-Transcribe-2-Streaming: live transcripts in 60 languages, first hypotheses in just over 100 milliseconds, priced at $0.54 per hour of audio through the end of 2026.
  • MAI-Voice-2.1: 23 languages across 26 locales, $22 per million characters.
  • MAI-Voice-2.1-Flash: 150 milliseconds end to end, 55% faster inference, $15 per million characters.
  • Availability: Microsoft Foundry, the MAI Playground, OpenRouter and Vercel, with LiveKit and Azure Voice Live integrations coming.

What Is MAI-Transcribe-2-Streaming?

This is a streaming speech-to-text model, so it produces words while someone is still talking rather than waiting for the end of a recording. Microsoft says it ranks first on the Artificial Analysis leaderboard and supports automatic, continuous language detection, which helps when a speaker switches languages mid-conversation. Microsoft also says words appear about twice as fast as the closest competitor. That comparison is Microsoft's own claim, so treat it as a vendor benchmark until independent tests arrive.

At $0.54 per hour of audio through the end of 2026, the pricing is aimed at developers who want to try live transcription at scale, from call-center assistants to meeting tools.

What Do MAI-Voice-2.1 and MAI-Voice-2.1-Flash Add?

MAI-Voice-2.1 covers 23 languages across 26 locales and can keep a single voice speaking multiple languages with what Microsoft calls a native accent. The Flash variant is tuned for high-volume work, generating up to 45 seconds of audio with 150 milliseconds of end-to-end latency and 55% faster inference at $15 per million characters.

Both models support voice cloning from a short reference clip. Microsoft says consent guardrails are built in to guard against misuse, which is worth checking when you plan a product around cloned voices.

Can a Voice Agent Now Answer Before You Finish Your Sentence?

Not literally, but the gap is shrinking. A transcription model that delivers first words in about 100 milliseconds, paired with a voice model that starts speaking in about 150, leaves much more of the latency budget for the language model in the middle. That is the formula for assistants that feel responsive. These models follow the first generation we covered in our MAI-Transcribe-1 and MAI-Voice-1 launch report. For more model news, see our AI section.

Sources: Microsoft AI, our first streaming transcription model — October 1, 2026.

More AI Stories