Microsoft’s new MAI models target low-latency speech-to-text and text-to-speech applications, including real-time voice agents and conversational services.
Microsoft has expanded its in-house artificial intelligence portfolio with three new speech models designed to make real-time voice conversations faster and more natural.
The company introduced MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash for developers building low-latency conversational AI applications, including voice agents and other real-time services.
The biggest addition is MAI-Transcribe-2-Streaming, which converts speech into text while a person is still speaking. Unlike the earlier MAI-Transcribe-2, which is designed for recorded audio, the streaming model begins producing partial transcription results just over 100 milliseconds after receiving audio.
It continuously updates the transcript as more context becomes available before generating the final result. Microsoft says the model supports 60 languages and can continuously detect language changes.
The company also says MAI-Transcribe-2-Streaming currently ranks first on Artificial Analysis for the accuracy of both partial and final transcripts. In Microsoft’s internal testing for dictation and subtitles, words appeared in the transcript about twice as quickly as with the closest competing model.
The streaming transcription model is priced at $0.54 per hour of audio through the end of 2026, compared with the promotional price of $0.10 per hour for the non-streaming MAI-Transcribe-2.
Microsoft has also introduced MAI-Voice-2.1, a multilingual text-to-speech model supporting 23 languages across 26 locales. It can generate speech that switches between supported languages while maintaining the same speaker identity and can use a native accent appropriate to each language.
MAI-Voice-2.1 is priced at $22 per 1 million characters.
For applications where response speed is a priority, Microsoft is offering MAI-Voice-2.1-Flash. The company says the model can generate about 45 seconds of audio with approximately 150 milliseconds of end-to-end latency. Microsoft also claims 55% faster inference and a cost about 60% lower than comparable models.
MAI-Voice-2.1-Flash costs $15 per 1 million characters.
Both voice models support voice cloning in their supported languages using only a few seconds of reference audio.
Microsoft has made all three models available through Microsoft Foundry, MAI Playground and Vercel. The two voice models are also available through OpenRouter, while LiveKit support is expected soon.
The models can also be used through Azure Voice Live, giving developers multiple options for integrating speech capabilities into real-time conversational applications.
About The Author
Muhammad Mubbashir Rauf
Mubbashir Rauf is the WEB EDITOR of Click Pakistan. He can be reached at mmubbashirrauf@gmail.com.













