Agora Media Lab reported that GPT-Live, launched by OpenAI in July 2026 to enhance natural human-AI conversations, shows improvements in voice AI performance primarily due to increased consistency in response times rather than significant reductions in raw speed. The median response time only improved by 205 milliseconds compared to the previous voice mode, but the standard deviation of latency dropped from 489 milliseconds to 104 milliseconds. This decline in variability means that users can expect responses that are more predictably timed, contributing to a smoother interaction experience.
OpenAI: OpenAI develops foundational AI models including the GPT series for language and multimodal tasks. In July 2026 the company introduced GPT-Live, a new generation of voice models designed for more natural, bidirectional conversations in ChatGPT. The latency benchmark discussed in the news directly evaluates this recent GPT-Live release against prior voice capabilities.
AgoraIO: AgoraIO supplies real-time communication infrastructure and SDKs used across voice, video, and interactive applications. Its Agora Media Lab performed the independent iPhone-based measurements of GPT-Live end-to-end latency described in the news. The work underscores AgoraIO’s ongoing role in providing tools and testing methodologies for emerging real-time AI voice systems.
Agora Media Lab: Agora Media Lab conducts applied research on media processing, networking, and AI interactions. In the current news it executed controlled benchmark tests using repeatable speech, dual-track recordings, and network impairments to assess GPT-Live performance. These measurements highlight consistency improvements that complement OpenAI’s latest voice model rollout.
Consistency Focus: Recent independent testing emphasizes that reduced variability in response timing, rather than raw minimum latency alone, drives smoother perceived performance in real-time voice AI.
Voice AI Advancement: OpenAI launched GPT-Live in July 2026 specifically to enable simultaneous listening and speaking for more natural human-AI conversations in ChatGPT Voice.
