Microsoft launches MAI‑Transcribe‑2‑Streaming for real‑time speech in 60 languages at $0.54 per hour
Microsoft AI put its MAI‑Transcribe‑2‑Streaming service into public preview on 1 October 2026, promising sub‑100 ms hypothesis generation, continuous language detection across 60 languages and a benchmark‑leading error rate.

On 1 October 2026 Microsoft AI announced that its MAI‑Transcribe‑2‑Streaming service entered public preview. The offering is positioned as a cloud‑based, real‑time speech‑to‑text engine that can handle audio streams in sixty languages, delivering both partial and final transcripts as the audio progresses. Microsoft markets the service as a low‑cost option for developers who need instant captions, live‑translation pipelines or voice‑driven analytics.
Real‑time transcription capabilities
Streaming transcription differs from batch processing in that the model emits hypotheses while the speaker is still talking. MAI‑Transcribe‑2‑Streaming claims to produce an initial hypothesis roughly one hundred milliseconds after the audio sample arrives, followed by refined partial results and a final, punctuation‑rich transcript. The system also performs continuous language detection, automatically switching its acoustic and language models when it detects a change among the supported sixty languages, eliminating the need for a separate language‑selection step.
Developers can reach the service through Microsoft Foundry, which exposes a Realtime API that mirrors the WebSocket patterns used by OpenAI’s streaming endpoints. The same functionality is also available via the Azure Speech SDK, allowing integration into Windows, Linux, mobile and web applications with minimal code changes. Microsoft’s documentation emphasizes that the API accepts raw audio frames and returns JSON‑encoded transcript fragments, making it compatible with existing real‑time pipelines.
Benchmark performance and accuracy
Artificial Analysis, an independent benchmarking platform, placed MAI‑Transcribe‑2‑Streaming at the top of its streaming word‑error‑rate (WER) leaderboard, which compared thirty‑eight models on a common eight‑hour test mix. According to the benchmark, the model achieved a 2.5 % WER for final transcripts with an average latency of 0.13 seconds, and the same 2.5 % WER for first partial results delivered in 0.12 seconds. MarkTechPost reported these figures, noting that the test mix included a variety of speakers, accents and background noises.
Microsoft’s own measurements underpin the benchmark numbers, and the company cautions that speed claims outside the Artificial Analysis test should be treated as internal estimates. The public preview does not include a production‑grade service‑level agreement, meaning that enterprises planning mission‑critical deployments will need to evaluate reliability and latency under their own workloads before committing to a paid tier.
Pricing, related models and preview limits
For the duration of the preview, Microsoft has set the introductory price at fifty‑four cents per hour of processed audio, a rate that remains in effect through the end of 2026. The preview does not carry a formal SLA, and usage beyond the announced benchmark figures is not guaranteed. Alongside the transcription service, Microsoft also released MAI‑Voice‑2.1 and Voice‑2.1‑Flash, the latter advertised as capable of generating forty‑five seconds of synthetic speech with an end‑to‑end latency of roughly 150 milliseconds at a cost of fifteen dollars per million characters.
The Flash model is intended for low‑latency voice generation scenarios such as interactive assistants or real‑time dubbing, complementing the streaming transcription engine by providing a fast, high‑quality voice output path. Both models share the same integration surface through Microsoft Foundry, allowing developers to chain transcription and synthesis in a single WebSocket session if desired.
- Supports 60 languages with automatic detection
- Initial hypothesis generated in ~100 ms after audio receipt
- Benchmark‑reported 2.5 % WER at 0.13 s latency for finals
- Introductory price of $0.54 per audio hour (preview period)
- Access via Microsoft Foundry Realtime API and Azure Speech SDK
The combination of low latency, multilingual coverage and a transparent pricing model opens new possibilities for developers building live captioning, multilingual call‑center analytics, or real‑time translation services. Because the service streams partial results, applications can begin processing text before the speaker finishes, reducing end‑to‑end response times for downstream AI components such as sentiment analysis or intent detection. The modest cost structure also lowers the barrier for startups and research teams that previously avoided cloud transcription due to high per‑minute fees.
For organizations operating in English‑speaking markets, the preview means that real‑time speech interfaces can now be deployed at a fraction of the cost of earlier Azure Speech offerings, while still covering a broad set of languages for global teams. Companies can experiment with live captioning for webinars, integrate instant subtitles into video platforms, or augment voice‑driven customer support with on‑the‑fly transcription, all without committing to a production SLA until the service graduates from preview.
Sources
- Our first streaming transcription model debuts at no. 1 on Artificial AnalysisMicrosoft AI · October 1, 2026
- Microsoft AI Releases MAI-Transcribe-2-StreamingMarkTechPost · October 2, 2026



