Soniox has launched its new Soniox v5 Real-Time, a low latency streaming Speech to Text model that balances accuracy and speed at a competitive price of $2 per 1,000 minutes, the lowest among proprietary models tested. This model joins the recently released Soniox v5 Async as part of an updated suite optimized for streaming and batch processing. In benchmarking against competitors like Cartesia, ElevenLabs, and Deepgram, Soniox v5 Real-Time delivers a 4.5% Word Error Rate (WER) with a 0.05-second response time for final transcription, outperforming faster models like Deepgram Flux while remaining cost-effective. Additionally, it supports over 60 languages, enhancing its utility for multilingual conversations.

Soniox: Soniox is a speech AI company focused on developing multilingual real-time speech-to-text, text-to-speech, and translation APIs for live applications. It recently launched Soniox v5 Real-Time as its latest streaming model, building directly on the non-streaming Soniox v5 Async released the prior week. The company emphasizes structured, accurate transcripts suitable for real-time voice products and analytics.

Model Release: Soniox v5 Real-Time joins the recently released Soniox v5 Async as part of an updated suite of speech-to-text models optimized for both streaming and batch processing.
Competitive Landscape: It is benchmarked against other proprietary streaming models including those from Cartesia, ElevenLabs, and Deepgram on metrics of accuracy and latency.
Language Capabilities: The model supports identification and real-time translation across more than 60 languages for multilingual conversations.