Google rolls out Gemini Audio upgrades that erase filler words and boost multilingual transcription
Google has introduced new Gemini 3.5 Audio models that automatically remove filler sounds like “um” and improve transcription across more than 85 languages.
Google announced a set of Gemini 3.5 Audio enhancements, adding Live, Live Experimental, and the first-ever Transcribe model to its portfolio. The Transcribe variant claims a substantial leap over the prior Chirp 3 system, especially in handling multilingual input and reducing wording errors. It automatically excises filler utterances such as “um” and “uh,” formats spoken text, and supports user-defined vocabularies to preserve specialized terminology.
The model can also attribute speech to up to three speakers in pre-recorded audio and supplies precise timestamps for each word. Gemini 3.5 Live improves handling of mid-sentence interruptions and visual processing, while the Experimental version narrates its reasoning steps in real time. The rollout begins today for English users on macOS and for Android dictation in selected regions, with API access offered in public preview and Chrome support slated for the near future.
Why it matters
The upgrade makes voice dictation more natural and accurate, helping users and developers create cleaner, multilingual text without manual editing.
In this story
