The news
On October 1, 2026, Microsoft (MSFT) released MAI-Transcribe-2-Streaming, its first real-time transcription model, with two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Microsoft pitches them as building blocks for voice agents, including customer-service agents that can act during a call.
The streaming model takes audio over a WebSocket, a connection that stays open, and returns text while a person is still speaking instead of waiting for a finished recording. Microsoft says it produces its first partial transcripts in just over 100 milliseconds and that words appear twice as fast as with its closest competitor, which it did not name. The company also says the model ranks first for accuracy on both final and partial transcripts on the Artificial Analysis leaderboard. It supports 60 languages and detects the language automatically.
Microsoft set an introductory price of $0.54 per hour of audio through the end of 2026. Its earlier MAI-Transcribe-2 model, released on September 3 for audio that is not streamed, carries a limited-time price of $0.10 per hour through year-end, according to Microsoft. SiliconANGLE noted that the streaming version costs more than five times as much.
MAI-Voice-2.1 covers 23 languages and 26 locales and lets one synthetic speaker switch languages while keeping the same voice, Microsoft said. It can clone a voice from a few seconds of reference audio, with what Microsoft calls built-in consent guardrails, and costs $22 per million characters. MAI-Voice-2.1-Flash, at about $15 per million characters, is built for high-volume work; Microsoft says it can generate 45 seconds of audio with 150 milliseconds of end-to-end latency and is about 60% cheaper than comparable models.
The models are offered through Microsoft Foundry, the company's platform for AI models, as well as MAI Playground, OpenRouter, Vercel and Azure Voice Live, with LiveKit listed as coming soon. Naomi Moneypenny, senior director of product development for Microsoft Foundry Models, said the models give developers "more choice in how to build natural voice experiences," as reported by PYMNTS.
The numbers
- MAI-Transcribe-2-Streaming introductory price
- $0.54 per audio hour (through end of 2026)
- MAI-Transcribe-2 (non-streaming) limited-time price
- $0.10 per audio hour
- Transcription languages
- 60
- MAI-Voice-2.1 price
- $22 per 1M characters
- MAI-Voice-2.1-Flash price
- About $15 per 1M characters
- First partial transcript (Microsoft claim)
- Just over 100 ms
Why CEOs should care
For customer-experience leaders, Microsoft's published prices give a reference point for the next contract talk with a speech or contact-center vendor. By Tech CEO Daily's calculation, a center streaming 1 million minutes of calls a month would pay about $9,000 a month for transcription at the introductory rate, before the cost of the reasoning model and the synthetic voice. Ask incumbent vendors for per-hour streaming prices you can compare, and ask Microsoft what the rate becomes after 2026.
Benchmarks are a starting point, not a buying decision. Microsoft's accuracy and speed claims come from its own announcement and a public leaderboard, not from your calls. Run a pilot on your own recordings, with your accents, background noise and product names, and measure errors and delays against the system you use now before moving live traffic.
CISOs and legal teams should look hard at the voice-cloning feature. A model that copies a voice from a few seconds of audio raises fraud and consent questions, so ask Microsoft how its consent guardrails work, where caller audio is processed and stored, and how long it is kept. Boards should also note the strategic signal: Microsoft is building its own models rather than relying only on partners.
The bigger picture
The launch fills out an in-house voice stack. SiliconANGLE described a three-part pipeline: MAI-Transcribe-2-Streaming to hear the caller, Microsoft's MAI-Thinking-1 reasoning model to decide what to do, and the MAI-Voice models to answer aloud. In September, Microsoft compared MAI-Transcribe-2's speed against OpenAI's GPT-Transcribe, ElevenLabs' Scribe v2 and Google's Gemini 3.5 Transcribe, which shows the vendors it sees as rivals.
Cost is part of the motive. SiliconANGLE reported that Microsoft AI chief executive Mustafa Suleyman has said the goal is to reduce, and ultimately eliminate, what Microsoft pays Anthropic. Owning more of the model stack gives Microsoft room to price aggressively in markets such as customer service, where usage is measured in millions of minutes.
What’s next
Watch what Microsoft charges for the streaming model after the introductory rate ends at the close of 2026, when LiveKit support arrives, and whether independent testers confirm the accuracy and speed claims. Also watch how speech specialists and contact-center platforms respond on price.
What “Fact-checked” means
Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.
- What we checked
- Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
- How
- A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
- Who
- The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, Hussein Mukhtar. A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
- If something is wrong
- “Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error
Companies in this story








