Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text
Source: Ars Technica · Ryan Whitwam
Intel Summary
Google has announced Gemini 3.5 Transcribe, a dedicated speech-to-text model designed to expand its voice transcription capabilities across consumer and enterprise products. According to reporting from Ars Technica, the underlying AI technology previously used in Gboard's Rambler feature will now be integrated into additional Google platforms, including the Chrome web browser. The release reflects Google's continued strategy of modularizing its Gemini model family for specific multimodal tasks like real-time audio transcription.
Why It Matters
Expanding automated speech-to-text natively into widespread software like Chrome lowers the barrier for voice-first user interfaces and accessibility tools across the web. For developers and enterprise IT teams, dedicated audio models from hyperscalers increase competition against specialized transcription providers, potentially lowering transcription costs while embedding baseline automated transcription directly into everyday browser-based workflows.
Part of an ongoing development
Independent reportingGoogle announces Gemini 3.5 Transcribe
Google has launched Gemini 3.5 Transcribe, an automated speech-to-text model supporting more than 85 languages. Google reports a 4.0 percent word error rate in streaming mode alongside a 70 percent reduction in latency compared to its previous Chirp 3 system. Claims are as reported; this summary makes no determination about accuracy or significance.
- Confidence
- High confidence
- Corroboration
- Corroborated
More coverage of this development
Organizations & Entities
Topics
Related Intelligence
- ReportSame development
Google's Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles
Google has launched Gemini 3.5 Transcribe, an automated speech-to-text model supporting more than 85 languages. According to the company, the model performs real-time transcription with automatic filler word removal and speech stumble correction. Google reports a 4.0 percent word error rate in streaming mode alongside a 70 percent reduction in latency compared to its previous Chirp 3 system. The model also supports function calling to hand off transcription workflows directly to other Gemini models.
The Decoder - ReportSame development
Google’s new AI transcription edits out your ‘ums’ and ‘ahs’
Google has introduced Gemini 3.5 Transcribe to its Gemini Audio suite, adding automatic speech recognition capabilities across more than 85 languages. The model is designed to recognize domain-specific jargon and automatically filter out speech disfluencies, such as filler words. The addition follows the release of Google's 3.5 Live Translate feature as the company continues expanding its Gemini 3.5 model portfolio ahead of anticipated larger foundation model releases.
The Verge - Report
OpenAI’s next big AI model has ‘entered the AGI era’
The Verge reports that OpenAI has introduced GPT-6 Astra, which the vendor describes as a generational leap in capabilities across software engineering, science, professional work, computer use, and cybersecurity. OpenAI also designated the release as its first model to meet its internal critical cybersecurity capability threshold.
The Verge - Report
Nvidia confirms $12.9B acquisition of AI hosting platform Hugging Face
Nvidia has agreed to acquire AI hosting and development platform Hugging Face for just over $12.93 billion, according to reporting by SiliconANGLE confirming recent deal discussions.
SiliconANGLE