Microsoft launches new voice agents: MAI-Transcribe-2-Streaming and MAI-Voice-2.1 explained

Have you ever noticed how uncomfortable that silence is when communicating with a voice bot? After you have completed your thought and there is nothing else you have to say, the bot listens to you until you stop, converts everything into text form, thinks, and after all this, speaks. Microsoft AI seeks to reduce this time. At the beginning of October, it released three new models called: MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. This is not yet a fully-fledged voice agent.

Also read: What is Tavus Griffin: The Human Interaction Model that passed the Turing Test

It listens while you talk

MAI-Transcribe-2-Streaming is a live speech-to-text engine. Rather than wait until you’re done speaking, it makes guesses, referred to as partials, in slightly over 100 milliseconds and then refines them as more context comes through.

The significance of this is that an application can do something about your words before you’ve completed your sentence. The chatbot can look up your order mid-way into your sentence. Live subtitles will pop up while the speaker is talking.

The model understands 60 languages and automatically detects them as well. It is rated by Microsoft as being number 1 on Artificial Analysis when it comes to accuracy, whether in its final or partial transcript forms, with low latency. Tests have shown that it produces words twice as fast as any other competitor. This is according to Microsoft, so take it as such. The pricing is $0.54 per hour of audio – an introductory price that lasts until the end of the year.

Also read: Best smart rings to buy in 2026: Samsung, Gabit and more

MAI-Voice-2.1: One voice, many languages

This is the text-to-speech module. MAI-Voice-2.1 supports 23 languages and 26 locales, and the cool thing about it is that the same voice works in all of them. Tell it to say something in English, then switch to Mandarin, then German – it will sound like the same voice speaking with a natural accent. Other solutions apply one accent to all languages. This solution does not. It costs $22 per 1M characters.

The voice models support cloning from 5 seconds of audio reference, which includes some built-in consent guardrails.

MAI-Voice-2.1-Flash: Built for speed

Flash is the younger and lighter version of the two, optimized for bulk and latency-critical tasks. It offers the same languages as well as cross-language voice support. Microsoft boasts an end-to-end latency of 150 ms and a 55% faster inference speed, priced 60% lower than competitive options at $15/1 million characters.

Why the pairing matters

Voice agents are a loop: listen, comprehend, make a decision, and respond. Saving time milliseconds on both ends helps the middle to think, call for help and double-check the results. Combining streaming speech-to-text with Flash allows for a conversation at a human speed, not at a robotic speed.

Microsoft’s demo features customer service agents, multilingual agents responding in the language of the incoming message, as well as tutors and role-plays requiring multiple speakers.

MAI-Voice-2.1 and Flash are distributed through OpenRouter. All three models are available via Microsoft Foundry, the MAI Playground, Vercel, Azure Voice Live, and with LiveKit on the way. Microsoft has also created a demo project called Chatter on the MAI Playground.

Also read: Best coffee machines for your home under Rs 15,000

Vyom Ramani

A journalist with a soft spot for tech, games, and things that go beep. While waiting for a delayed metro or rebooting his brain, you’ll find him solving Rubik’s Cubes, bingeing F1, or hunting for the next great snack.

Connect On :