EnterpriseModelsTools

Google's Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles

Source: The Decoder · Matthias Bastian

Intel Summary

Google has launched Gemini 3.5 Transcribe, an automated speech-to-text model supporting more than 85 languages. According to the company, the model performs real-time transcription with automatic filler word removal and speech stumble correction. Google reports a 4.0 percent word error rate in streaming mode alongside a 70 percent reduction in latency compared to its previous Chirp 3 system. The model also supports function calling to hand off transcription workflows directly to other Gemini models.

Why It Matters

Real-time speech transcription with integrated correction and native function calling reduces post-processing overhead for voice-driven enterprise workflows. Developers can build low-latency voice agents and automated documentation pipelines without intermediary parsing tools, intensifying competition across multilingual speech recognition and conversational AI infrastructure.

Part of an ongoing development

Independent reporting

Google announces Gemini 3.5 Transcribe

Google has launched Gemini 3.5 Transcribe, an automated speech-to-text model supporting more than 85 languages. Google reports a 4.0 percent word error rate in streaming mode alongside a 70 percent reduction in latency compared to its previous Chirp 3 system. Claims are as reported; this summary makes no determination about accuracy or significance.

Confidence
High confidence
Corroboration
Corroborated

More coverage of this development

Organizations & Entities

Topics