Gemini 3.1 Flash TTS: the next generation of expressive AI speech
Google DeepMind announced Gemini 3.1 Flash TTS, an updated text-to-speech model designed for expressive audio generation. According to the company, the model introduces granular audio tags that provide users with precise directional control over synthetic speech. The system aims to enhance controllability and emotional nuance in generated audio, allowing developers to steer vocal performance more predictably across downstream interactive applications.