Google launches Gemini 3.5 Transcribe, filtering out filler words
Gemini 3.5 Transcribe is Google's new speech-to-text model: it detects over 85 languages, cleans up self-corrections and filler words in real time, and reaches a 2.6% word error rate on recorded audio, according to DeepMind.

Google DeepMind announced Gemini 3.5 Transcribe on August 26, 2026, describing it as its "most precise speech-to-text model yet," built to convert raw audio directly into accurate, polished, formatted text, according to the company's official blog. The system is designed for what Google calls "intelligent voice interactions," moving beyond conventional speech recognition that struggles with background noise, jargon and disfluencies.
According to figures Google attributes to independent benchmarking firm Artificial Analysis, Gemini 3.5 Transcribe reaches an average Word Error Rate (WER) of 4.0% in streaming mode and 2.6% for pre-recorded audio. On the multilingual FLEURS benchmark, the model posts a 5.50% WER in streaming and 5.04% in non-streaming use, and Google says time to final transcription improves by 70% compared with its previous model, Chirp 3.
What makes it different from standard speech-to-text
The model's core pitch is "smart transcription": it seamlessly handles self-corrections — Google's own example is a speaker saying "let's meet Tuesday — no, Wednesday" — strips out filler words like "um" and "ah," and auto-formats the resulting text. Developers and end users can also feed it custom vocabulary lists so it correctly renders specialized jargon or unusual spellings instead of guessing.
Gemini 3.5 Transcribe automatically detects and transcribes more than 85 languages, handling regional accents and dialects without the user specifying a language in advance, per DeepMind's announcement. In pre-recorded audio, it can also attribute speech to up to three separate speakers with word-level timestamps; Google flags support for more than three speakers as still experimental.
Beyond transcription, the model can trigger "function calling" to hand off tasks — such as image generation or file analysis — to other Gemini models, a capability currently live in the Gemini app on macOS. German outlet The Decoder additionally reported that the model can delegate web searches this way, a use case Google's own blog post does not explicitly list.
Two APIs, and a growing list of surfaces
- Real-time streaming, via the Live API (model gemini-3.5-transcribe-live), for voice agents and live captioning with sub-second latency
- Pre-recorded processing, via the Interactions API (model gemini-3.5-transcribe), for meetings, call logs and post-call analytics
- Rambler, a new dictation feature in Gboard on Android, in select countries and languages
- The Gemini app on macOS, currently in English, pairing transcription with voice commands and screen context
- Google Antigravity, where it combines screen context and chat history to transcribe file names and technical terms accurately
- Google AI Studio's Build mode, for describing and building apps by voice
Third-party platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents already build on the Live API to ship voice-driven interfaces, and Google cites early feedback from Vivo, Intellitek Health and Lingopal. Chrome support, letting users dictate into any web field, is described as "coming soon."
A quiet release inside a bigger wait
The Verge noted that Gemini 3.5 Transcribe arrives while Google's long-promised Gemini 3.5 Pro model, expected in June, still has not shipped. The outlet also reported — then corrected after Google reached back out — that two further audio models, Gemini 3.5 Live and 3.5 Live Experimental, were not in fact launching alongside Transcribe despite information Google had initially provided; no new date was given for them. Availability itself remains staged: the transcription model is in public preview for developers via Google AI Studio and Antigravity, in public preview for enterprises via the Gemini Enterprise Agent Platform, and only partially rolled out to consumers.
For companies that process large volumes of multilingual audio — customer support calls, sales meetings, podcasts, subtitling — a transcription layer that removes filler words and attributes speakers automatically could cut the manual editing time that currently follows most transcripts. The lower WER on pre-recorded audio (2.6%) versus live streaming (4.0%) suggests batch use cases like call-center analytics are currently the safer bet, while real-time captioning and voice agents remain a slightly noisier proposition. Businesses evaluating it should also note that speaker attribution beyond three participants and full non-English support on macOS are not yet ready, meaning larger meetings and non-English desktop users will need to wait before relying on it in production.
Sources
- Introducing Gemini 3.5 TranscribeGoogle DeepMind · August 26, 2026
- Google's new AI transcription edits out your 'ums' and 'ahs'The Verge · August 26, 2026
- Googles Gemini 3.5 Transcribe erkennt über 85 Sprachen und filtert Füllwörter in Echtzeit herausThe Decoder · August 26, 2026



