Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Microsoft releases MAI-Transcribe-2-Streaming, a real-time speech-to-text model - OpenSmartRoute
Introduction - Microsoft releases MAI-Transcribe-2-Streaming as the first real-time speech-to-text model
Microsoft has introduced a new speech-to-text model called MAI-Transcribe-2-Streaming. It is the first model from Microsoft that can transcribe speech in real time. This means it can turn spoken words into written text instantly as someone talks. The launch happened on October 1, 2026. The model is part of a set of new speech and text tools Microsoft announced at the same time.
This new model is designed to work in live situations. It aims to improve how voice assistants, live captions, and dictation services perform. Unlike older models that process audio after recording, MAI-Transcribe-2-Streaming works while the speaker is still talking. This allows for faster responses and more natural interactions. The model is also ranked as the most accurate among 38 tested speech-to-text models.
Model overview - Features, target uses, and how it differs from batch models
MAI-Transcribe-2-Streaming is a real-time speech-to-text model. It can transcribe speech as it happens, unlike batch models that process recorded audio after the fact. Batch models analyze large chunks of audio at once, which can cause delays. Streaming models like this one process audio continuously, providing immediate text output.
The model supports 60 languages and can automatically detect which language is being spoken. It can handle audio streams that go on without stopping. As the speaker talks, the model sends back partial transcripts. These are early guesses that get better as more audio arrives. When the speaker finishes, the model produces a final, stable transcript. This setup helps voice agents, live captioning, and dictation services work more smoothly.
The model is designed for applications where low latency is critical. It can start giving partial results just over 100 milliseconds after hearing the first sounds. It updates these partials as more context becomes available. This means users see words appear faster than with previous models. It also allows agents to start reasoning or calling tools before the speaker finishes a sentence. The model's internal tests show it can produce words twice as fast as its closest competitors.
Mirror Particle raises capital to create an AI engine that simulates changing human motivations. The company rejects large language models in favor of a foundation model trained on longitudinal data.
Mistral AI launched Mistral Large 4 with one trillion parameters. It is not yet an open-weight model but will release weights in three weeks.
Performance metrics - Accuracy, latency, and comparison with competitors
Microsoft reports that MAI-Transcribe-2-Streaming is the most accurate speech-to-text model tested so far. It scored a 2.5% word error rate (WER) on a standard evaluation index called AA-WER Streaming. WER measures how many words are wrong in the transcript. A lower WER means higher accuracy.
The model produces its first partial transcript about 120 milliseconds after the speaker starts talking. Its final transcript appears roughly 130 milliseconds after the speaker stops. This means it delivers both early guesses and final results very quickly. The final transcript also has a 2.5% WER, matching the accuracy of the first partial. This is significant because early results are as reliable as the final ones.
Compared to other models, MAI-Transcribe-2-Streaming outperforms several competitors. For example, one rival, Grok Voice Transcribe 2.0, has a WER of 2.7% but takes about 490 milliseconds to produce the first partial. Another, Muse Voice Transcribe, has a WER of 3.1% and appears about 160 milliseconds after speech begins. A different model, Cartesia Ink-2, can produce final results in just 70 milliseconds but has a higher WER of 4.0%. This shows that MAI-Transcribe-2-Streaming offers a good balance of speed and accuracy.
Microsoft states that its internal tests show words appear twice as fast as in competing models. This speed advantage is crucial for real-time applications. It allows voice agents to respond quickly and improves user experience in live settings. The model is also placed on the accuracy versus latency Pareto frontier, meaning it offers a good trade-off between speed and correctness.
Model details - Languages, partial and final transcript timings, and evaluation setup
MAI-Transcribe-2-Streaming can transcribe speech in 60 languages. It also detects the language automatically as the conversation continues. This feature is useful in multilingual environments or when the language is unknown beforehand. The model streams audio continuously, which means it can handle ongoing conversations without stopping.
The model produces its first hypotheses, called partial transcripts, just over 100 milliseconds after hearing the first sounds. These partials are early guesses that get refined as more audio arrives. The model revises these partials as it gathers more context, ensuring the text improves over time. When the speaker stops talking, the model commits to a final transcript that is stable and accurate.
The evaluation of the model used about 8 hours of audio from different sources. These sources include voice agents, public speech datasets, and earnings calls. The mix was about half from voice agents, with the rest from general speech datasets. The timing of the partial and final transcripts was measured from the moment speech ended, as detected by a voice activity detection system called SileroVAD.
The results show that the first partial transcript appears roughly 120 milliseconds after speech begins. It has the same accuracy as the final transcript, which appears about 130 milliseconds after speech ends. This quick turnaround allows for faster responses in live applications. The model’s high accuracy and low latency make it suitable for real-time voice systems.
Pricing and availability - Cost, API options, and integration paths
Microsoft offers MAI-Transcribe-2-Streaming at an introductory price of $0.54 per hour of audio. This rate is valid through the end of 2026. When normalized to minutes, it costs about $9.00 per 1,000 minutes of audio. This pricing is higher than some other models, such as the batch version of MAI-Transcribe-2, which costs $0.10 per hour.
The higher cost reflects the model’s real-time capabilities and high accuracy. Microsoft charges more than some competitors like xAI and Meta, but roughly matches Google’s estimated rates for similar services. The model is available through two main API paths. The Realtime API is suitable for apps already using an OpenAI-compatible WebSocket connection. The Azure Speech SDK manages connection, retries, and audio streaming, and both options return partial and final results.
The model can be integrated into applications using these APIs. It is also available in the MAI Playground, which runs on Vercel and Azure Voice Live. Support for LiveKit is expected soon. These options make it easier for developers to add real-time speech transcription to their products.
Microsoft pairs MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash for full voice loops. Flash can generate 45 seconds of audio at about 150 milliseconds latency for $15 per million characters. MAI-Voice-2.1 supports 23 languages and 26 locales at $22 per million characters. These tools together enable complete voice-based interactions, from speech synthesis to transcription.
Related models and tools - Voice generation models and their roles
Microsoft offers other models that complement MAI-Transcribe-2-Streaming. One such model is MAI-Voice-2.1-Flash, which generates speech audio. It produces 45 seconds of audio with very low latency. This is useful for creating voice responses in real time.
Another related tool is MAI-Voice-2.1, which covers 23 languages and 26 locales. It is designed for high-quality speech synthesis, allowing applications to produce natural-sounding voices. These voice models are used in virtual assistants, customer service bots, and other voice-enabled systems.
Microsoft also announced two text-to-speech models alongside MAI-Transcribe-2-Streaming. These models help generate speech from text, completing the cycle of voice interaction. Combining speech-to-text and text-to-speech models enables full voice communication systems.
These tools are part of Microsoft's broader effort to improve voice technology. They aim to make voice interactions more natural, fast, and accurate. Developers can combine these models to build advanced voice agents and communication apps.
Background - What speech-to-text models are and how they are used today
Speech-to-text models convert spoken words into written text using artificial intelligence. They analyze audio signals and recognize words, phrases, and sentences. These models have become essential in many applications.
Today, speech-to-text is used in voice assistants, live captioning, transcription services, and dictation tools. They help people communicate more easily, especially in noisy environments or for those with disabilities. These models also support multilingual communication by recognizing multiple languages.
Most older models processed audio after it was recorded. This batch approach caused delays and limited real-time use. Newer models, like MAI-Transcribe-2-Streaming, process speech continuously. They provide immediate text output, making voice applications more responsive.
Speech-to-text models are trained on large datasets of spoken language. They learn to recognize patterns and words from diverse voices and accents. This training improves their accuracy and ability to handle different speaking styles.
The development of these models has advanced with improvements in AI, computing power, and data availability. They are now faster, more accurate, and capable of supporting many languages. This progress enables new applications and improves existing voice services.
Why it matters - Impact on voice-based applications and real-time communication
The introduction of MAI-Transcribe-2-Streaming impacts many voice-based applications. Its high accuracy and low latency make live interactions more natural. Users can see captions, commands, or transcriptions almost instantly.
For voice agents, this means faster understanding and response times. It allows agents to start reasoning or calling tools before the speaker finishes. This improves the flow of conversations and makes voice assistants more useful.
Live captions for meetings, videos, and broadcasts also benefit. They can now be more accurate and appear faster. This helps people who rely on captions for understanding speech in real time. It can also assist in noisy environments or for those with hearing impairments.
The model's ability to handle multiple languages automatically is important in global settings. It supports multilingual conversations without needing to switch models. This flexibility simplifies deployment in diverse environments.
Overall, the model advances real-time communication. It reduces delays and errors, making voice interactions more seamless. This can change how people use voice in work, education, and entertainment.
What to do - How engineers and managers can adopt or evaluate the model
Engineers can start testing MAI-Transcribe-2-Streaming through its APIs. They should compare its speed and accuracy with existing models in their applications. Using the Realtime API or Azure Speech SDK, they can integrate the model into voice agents or captioning tools.
Managers should consider the model’s cost and benefits. They can evaluate whether the improved speed and accuracy justify the higher price. Testing in real scenarios can help determine its impact on user experience.
Both engineers and managers can explore the model in the MAI Playground. This platform allows hands-on testing without full integration. It helps understand how well the model performs with their data.
For deployment, organizations should plan for API management and scaling. They can also prepare fallback options if needed. Monitoring performance and user feedback will guide further improvements.
Finally, staying updated on new features and support options, such as upcoming support for LiveKit, will help maximize the model’s potential. Combining it with other voice tools can create comprehensive voice solutions.