Skip to main content
ModelsTools

ElevenLabs' new v4 speech model makes AI voices more expressive and consistent

Source: The Decoder (opens in a new tab) · Jonathan Kemper

Intel Summary

ElevenLabs has released Eleven v4, an updated speech model that improves voice consistency over long productions and responds more accurately to emotional cues like whispering and laughter. The release includes a Turbo variant featuring a 150-millisecond latency for real-time voice agents. According to Artificial Analysis' Voice Arena leaderboard, Eleven v4 ranks ahead of Cartesia and Google's Gemini.

Why It Matters

Sub-200 millisecond response times and emotional nuance are essential benchmarks for deploying interactive voice agents and scalable audio production. The updated rankings indicate accelerating competition among specialized audio AI developers and frontier model providers in speech synthesis performance.

Part of an ongoing development

Source

ElevenLabs released Eleven v4 speech model

ElevenLabs has released Eleven v4, an updated speech model that improves voice consistency over long productions and responds more accurately to emotional cues like whispering and laughter. The release includes a Turbo variant featuring a 150-millisecond latency for real-time voice agents. Claims are as reported; this summary makes no determination about accuracy or significance.

Organizations & Entities