Google Launches Gemini 3.5 Transcribe, a Dedicated Speech-to-Text Model
Google has released Gemini 3.5 Transcribe, a speech-to-text model built on the Gemini audio understanding core and optimized for fast, accurate, and cost-effective transcription. Unlike general-purpose Gemini models, it offers native features such as speaker diarization, word-level millisecond timestamps, and a custom vocabulary dictionary supporting up to 1,000 terms. The model supports over 85 language locales and can handle both verbatim transcription and cleaned, reading-optimized output without complex prompting. It runs on the Google GenAI SDK version 2.0 and above, with audio input handled through the Files API. Developers can test the model via an interactive Colab notebook or directly through Google AI Studio without writing any code.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in