Microsoft releases three new MAI speech AI models for transcription and voice synthesis.
Microsoft introduced three documented MAI speech models: MAI-Transcribe-2 for speech-to-text, and MAI-Voice-2 and MAI-Voice-2-Flash for text-to-speech. MAI-Transcribe-2 offers multilingual transcription, speaker identification, and word-level timestamps but lacks documented real-time streaming. The voice models provide high-fidelity and low-latency speech generation for different application needs. These releases give developers distinct tools for building voice-based applications and workflows.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in