EnterpriseModelsTools

Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

Source: The Verge · Jess Weatherbed

Intel Summary

Google has introduced Gemini 3.5 Transcribe to its Gemini Audio suite, adding automatic speech recognition capabilities across more than 85 languages. The model is designed to recognize domain-specific jargon and automatically filter out speech disfluencies, such as filler words. The addition follows the release of Google's 3.5 Live Translate feature as the company continues expanding its Gemini 3.5 model portfolio ahead of anticipated larger foundation model releases.

Why It Matters

Automated filtering of filler words combined with multilingual jargon detection reduces post-processing overhead for meeting summaries, transcription workflows, and automated voice pipelines. The release enhances Google's native speech-to-text capabilities within the Gemini ecosystem, increasing competitive pressure on standalone transcription services and alternative speech models like Whisper.

Part of an ongoing development

Source

Google announces Gemini 3.5 Transcribe

Google has launched Gemini 3.5 Transcribe, an automated speech-to-text model supporting more than 85 languages. Google reports a 4.0 percent word error rate in streaming mode alongside a 70 percent reduction in latency compared to its previous Chirp 3 system. Claims are as reported; this summary makes no determination about accuracy or significance.

Confidence
High confidence
Corroboration
Corroborated

More coverage of this development

Organizations & Entities