
Microsoft AI published MAI-Transcribe-2 on September 3, 2026, positioning it as a speech-to-text model for real-world audio. The official announcement names clinical note-taking, legal documentation, accessibility, closed captioning, and other workflows that need searchable text from speech. The release is notable not only for its language count, but for putting speaker identity, timing, and domain terminology inside the same transcription component.
MAI-Transcribe-2 supports 60 languages, automatic language identification, and code switching. It can continue through a conversation that naturally changes language instead of requiring each audio segment to be assigned one language in advance. That could reduce preprocessing and model switching for multilingual support, meetings, and media. Actual quality still needs to be tested against local accents, language mixtures, recording conditions, and proper nouns.
Speaker diarization separates speakers, while word-level timestamps locate each word in the recording. Keyword biasing lets a business provide recognition hints for product names, drug names, abbreviations, or internal terminology. Those capabilities map directly to contact-center QA, caption search, meeting navigation, clinical-note drafts, and editing. Keyword biasing remains a hint, however, not a guarantee that a term will be transcribed correctly.
The model also offers configurable transcription styles. Verbatim preserves filler words, false starts, and other speech details for compliance, quality review, and analysis. Clean removes some spoken fillers to produce more readable captions, notes, and publication drafts. This split is more useful than optimizing for one minimum Word Error Rate because compliance records and public-facing content have different requirements.
Microsoft says MAI-Transcribe-2 reached an average 5.2% Word Error Rate on FLEURS across 60 languages, sat on the Artificial Analysis accuracy-latency Pareto frontier, ranked second on its WER comparison, and was up to 10 times faster than leading competitors. These figures come from Microsoft's product announcement and the evaluations it cites. They should not be treated as outcomes that every audio environment will reproduce.
The launch price is listed as $0.10 per hour of audio through the end of 2026 as a limited-time offer. Microsoft says the model can be demonstrated through Microsoft Foundry, MAI Playground, and OpenRouter. Before a production integration, teams should confirm preview status, region, quotas, data handling, SLA, audio-size limits, and the real cost calculation. For long-form audio, speed affects more than spend: it changes whether search, QA, and downstream agent work can happen in near real time.
The model's value ultimately depends on what happens after transcription. Text with speaker identity, timing, terminology, and a verifiable audio provenance can feed search, summarization, CRM, captions, or voice agents. A plain text block without provenance leaves downstream automation vulnerable to silent errors. Evaluation should measure word errors, missed critical content, speaker mismatches, timestamp drift, human edits, and cost per audio hour together.
The release therefore shows Microsoft moving transcription from "audio in, text out" toward a workflow component that can be embedded in enterprise systems. Its benchmark, speed, and price deserve comparison, but while the service is in preview the safer path is to build a golden set from the organization's own recordings, validate data boundaries and human review, and only then decide which follow-up actions an agent may run.



