ElevenLabs releases Eleven v4 speech model with expressive voice features
1 source · The Decoder- Boom: Eleven v4 ranks above Cartesia and Google Gemini on Artificial Analysis Voice Arena leaderboard
- Boom: Turbo variant achieves 150-millisecond latency, designed for real-time voice agent use
- Boom: Model improves accuracy on expressive vocal cues including laughter and whispering
- Boom: Voice consistency across long productions like audiobooks is a stated improvement over prior versions
The story in full
ElevenLabs released its Eleven v4 speech model on September 29, 2025, adding improved accuracy for vocal cues such as laughter and whispering, along with better voice consistency across long-form content like audiobooks. A Turbo variant targets real-time voice agent applications with a 150-millisecond time-to-first-speech latency.
On Artificial Analysis' Voice Arena leaderboard, Eleven v4 ranks ahead of Cartesia and Google's Gemini, placing it among the leading publicly benchmarked speech models at launch.
Analysis
341 wordsOn September 29, 2025, ElevenLabs released Eleven v4, an updated speech synthesis model that the company says improves accuracy on expressive vocal cues, specifically laughter and whispering, and maintains voice consistency across long-form content such as audiobooks. A companion release called the Turbo variant is built for real-time voice agent applications and achieves a time-to-first-speech latency of 150 milliseconds. At launch, Eleven v4 ranked ahead of both Cartesia and Google's Gemini on Artificial Analysis' Voice Arena leaderboard, a publicly accessible benchmark for comparing speech models.
The release matters because voice synthesis is becoming infrastructure, not just a novelty. Latency at the 150-millisecond range is considered close enough to natural human response times to make automated voice agents feel conversational rather than robotic, which has direct implications for customer service, accessibility tools, and companion applications. The benchmark placement above Google's Gemini is notable given Google's resources, though leaderboard positions shift frequently and independent evaluations often surface trade-offs that aggregate scores obscure. What remains genuinely in dispute is how much benchmark rankings reflect real-world production performance, and whether improvements in expressiveness are sufficient to close the gap between synthetic and human voice talent for professional audio work like audiobooks.
No reactions from the Pro-AI, Anti-AI, or Middle Ground camps have been published yet on this story. Typically, the Pro-AI camp would highlight the latency and benchmark figures as evidence that AI voice technology is reaching commercial maturity, while the Anti-AI camp would raise concerns about the displacement of voice actors and narrators and the potential for misuse in deepfake audio. A middle-ground perspective would likely acknowledge the genuine utility for accessibility and automation while calling for clearer disclosure standards when synthetic voices are used in consumer-facing products.
The argument about real-world quality versus benchmark performance will likely be tested as developers and audio producers begin integrating v4 into production pipelines. Reviews from audiobook publishers and voice agent platforms in the weeks following launch would provide the most meaningful signal about whether the stated improvements in consistency and expressiveness hold up outside controlled evaluations.
Where do you stand?
Add your take
0 reader votesSign in with Google to pick a side and post. Your vote moves the story's Doom / Boom score.

