Google Gemini 3.5 Transcribe Models Differ Sharply on Speaker Diarization Support

A developer building a macOS meeting translation app discovered that Google released two similarly named models — gemini-3.5-transcribe-live and gemini-3.5-transcribe — with significantly different capabilities. The live streaming version uses WebSocket via the Live API and supports real-time interim transcription, but does not support speaker diarization or word-level timestamps. Speaker diarization, which identifies who said what, is only available in the non-streaming version, which requires uploading a complete audio file after recording and supports up to eight speakers. This architectural limitation meant the developer could not achieve real-time speaker identification and had to design the app around a post-meeting upload workflow for diarized transcripts. The live model does offer features like automatic language detection, a SMART cleanup mode that removes filler words, and a custom vocabulary option supporting up to 1,000 technical terms.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in