The Controlled Voice Arena has released its new leaderboard comparing various Text to Speech models using a consistent set of cloned voices. This assessment aims to standardize performance evaluations by cloning identical voices across models, which distinguishes voice preference from the overall quality of the models. Leading the overall rankings is Cartesia Sonic 3.5, with an Elo score of 1,122, followed closely by ElevenLabs Eleven v3 and Inworld Realtime TTS-2 – Research Preview. The leaderboard also highlights performance differences between US and UK accents, with Cartesia Sonic 3.5 continuing to excel in both categories.

Inworld: Inworld AI focuses on real-time AI for character interactions and has expanded into text-to-speech with voice cloning support. The company’s Realtime TTS-2 – Research Preview model places highly in the Controlled Voice Arena for both accent groups.
Cartesia: Cartesia develops advanced text-to-speech AI models optimized for natural voice synthesis. In the news, the company’s Sonic 3.5 model leads the Controlled Voice Arena leaderboard across US and UK accent categories through standardized voice cloning evaluations.
ElevenLabs: ElevenLabs specializes in high-quality voice AI and text-to-speech technologies with voice cloning features. Its Eleven v3 model ranks near the top of the Controlled Voice Arena results in the provided leaderboard.
Fish Audio: Fish Audio builds open-weight and open-source text-to-speech models emphasizing accessibility and performance. Its S2 Pro model leads the open weights category in the Controlled Voice Arena benchmark.
Mistral AI: Mistral AI is an AI company that develops large language models and has entered the text-to-speech space with models supporting voice cloning. Its Voxtral TTS ranks among the top open weights entries in the Controlled Voice Arena.
Resemble AI: Resemble AI creates voice cloning and text-to-speech solutions for customizable audio generation. Its Chatterbox model is included in the open weights section of the Controlled Voice Arena leaderboard.

Evaluation Methodology: The Controlled Voice Arena standardizes evaluations by cloning identical voices across models to separate voice preference from overall model quality in text-to-speech comparisons.
Open Weights Benchmarking: Open weights text-to-speech models are directly compared against proprietary systems using the same cloned voice samples and categories.
Accent-Specific Performance: Models demonstrate different relative strengths when evaluated separately on US versus UK accent voices within the leaderboard framework.