Google launches Gemini 3.5 Transcribe for smarter voice applications
Google’s new speech-to-text model targets lower-latency, more accurate transcription for voice agents, captions and call analytics.

Google introduced Gemini 3.5 Transcribe on Aug. 26, 2026, positioning it as a speech-to-text model built for precise, real-time voice interactions. In its Google DeepMind Blog announcement, the company said the model turns raw audio into polished, formatted text rather than simply recognizing words.
The release gives developers two integration paths. The Live API uses gemini-3.5-transcribe-live for continuous, bidirectional streaming with sub-second latency, aimed at interactive voice agents and live captioning. The Interactions API uses gemini-3.5-transcribe for recorded meetings, calls and other audio, adding speaker attribution and word-level timestamps.
Built for messy, real-world speech
Gemini 3.5 Transcribe can remove filler words, clean up self-corrections and adapt to custom vocabulary, including specialized terms and unusual spellings. It supports automatic detection and transcription across more than 85 languages, and can identify up to three speakers in pre-recorded audio; support for more than three remains experimental.
Google cites Artificial Analysis results showing a 4.0% average word error rate for streaming and 2.6% for non-streaming use cases. The company also reports a 70% improvement in time to final transcription compared with its earlier Chirp 3 model. On the FLEURS benchmark, Gemini 3.5 Transcribe recorded word error rates of 5.50% in streaming mode and 5.04% in non-streaming tests.
For teams building AI workflows, the model’s importance is less about transcription alone than what cleaner text enables: more reliable voice commands, searchable call records and downstream function calls to other Gemini models. Google says the technology is already used in voice features across the Gemini app and Android, and is now accessible through Google AI Studio and the Gemini Enterprise Agent Platform.
Source: Google DeepMind Blog
Comments
Log in to join the discussion