Google launches Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps

Google has released two new AI models via the Gemini API and Google AI Studio to help developers build real-time, voice-first applications. Gemini 3.8 Live and its Extended Thinking variant enable conversational voice agents capable of background reasoning, asynchronous tool calls, and multilingual support across 97+ languages. The Extended Thinking model has ranked first on Artificial Analysis' Speech-to-Speech leaderboard for complex, multi-step tasks. Separately, Gemini 3.5 Transcribe — launched last month — offers dedicated speech-to-text conversion across 85+ languages with a word error rate as low as 2.6% in non-streaming mode. Both models are accessible through the Gemini API, with integration support from partners including LangChain, LiveKit, and Vercel.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in