Discover the capabilities of OpenAI’s new GPT-Live voice model—it speaks just like a human
OpenAI has launched GPT-Live, a next-generation voice model that enables more natural conversations, real-time interaction, and enhanced AI capabilities across ChatGPT.
OpenAI has introduced GPT-Live, a next-generation voice model designed to deliver more natural, human-like conversations within ChatGPT. The company describes it as its most advanced speech model to date, offering significant improvements in responsiveness, conversational flow, and overall user experience.
The new model powers ChatGPT Voice and enables users to engage in real-time spoken conversations with AI that can listen and speak simultaneously.
Full-Duplex Architecture Enables Real-Time Conversations
At the core of GPT-Live is a full-duplex architecture, allowing the model to process incoming speech while responding at the same time rather than waiting for the user to finish speaking.
According to OpenAI, this approach creates a more fluid and natural conversation. The model can acknowledge the speaker with brief verbal cues such as "mm-hmm" or "I understand," pause naturally when users need time to think, and participate in fast-paced dialogue without interrupting the flow of conversation.
The company said these improvements represent a major step toward making ChatGPT Voice feel more intelligent and conversational while laying the foundation for future voice-powered AI agents capable of handling longer and more complex tasks.
Two Models Available for ChatGPT Users
OpenAI has released two versions of the new technology for ChatGPT users worldwide:
GPT-Live-1
GPT-Live-1 mini
The company also plans to make both models available to developers through its API in the near future.
Moving Beyond Earlier Voice Systems
OpenAI explained that earlier versions of ChatGPT Voice relied on three separate models that sequentially converted speech into text, generated a response, and then converted the response back into speech.
While that approach introduced conversational AI to millions of users, it also increased latency and occasionally resulted in information loss between processing stages.
Later, Advanced Voice Mode unified speech processing into a single model, reducing response times. However, conversations still followed a turn-based structure, requiring users to finish speaking before the AI responded, making interactions feel less natural and more susceptible to interruptions caused by brief pauses or background noise.
GPT-Live addresses these limitations by enabling simultaneous listening and speaking, allowing for smoother conversations and opening the door to capabilities such as real-time spoken translation.
Intelligent Task Delegation
Another key innovation is GPT-Live's ability to separate conversation management from computationally intensive tasks.
When users request web searches, advanced reasoning, or AI agent-based actions, GPT-Live can delegate those tasks to more specialized models such as GPT-5.5 while continuing the voice conversation without disrupting the interaction.
This architecture allows users to maintain a continuous dialogue while more complex processing happens in the background.
Enhanced Voice Experience in ChatGPT
OpenAI said more than 150 million people use ChatGPT's voice conversation and voice dictation features every week.
Users rely on these capabilities for everyday assistance, language learning, bedtime storytelling, hands-free productivity, and casual conversations while on the move.
With GPT-Live, pressing the Voice button in ChatGPT now provides:
More natural voice conversations
Smarter AI responses
Improved listening capabilities
Visual responses during conversations
Adjustable reasoning modes, including Instant for faster replies and Medium for more thoughtful responses
Multimodal Features Expand Voice Capabilities
GPT-Live supports ChatGPT's broader multimodal ecosystem, including:
Web search
Memory
Image understanding
File uploads
This allows users to discuss documents during voice conversations, ask questions about uploaded images, receive photography advice, and access rich visual cards for topics such as weather, financial markets, and sports.
New Safety Measures
OpenAI also introduced enhanced safety systems designed specifically for voice interactions.
If the model detects that a response could be unsafe, it can redirect the conversation toward safer guidance, provide additional safety resources, or terminate the voice session entirely in high-risk situations.
For conversations involving self-harm, GPT-Live incorporates specialized support mechanisms adapted for voice interactions, including information about crisis helplines developed in consultation with safety experts.
The company also strengthened protections for teenagers by training the model to provide age-appropriate responses and reduce the likelihood of unsuitable content.
Parents can manage access to ChatGPT Voice through parental control tools, while linked guardians may receive notifications in high-risk situations involving potential signs of self-harm or suicidal intent.
Global Rollout Begins
GPT-Live is rolling out globally through the ChatGPT apps for iOS, Android, and ChatGPT.com.
For subscribers to Go, Plus, and Pro, GPT-Live-1 will become the default model powering ChatGPT Voice, while GPT-Live-1 mini will serve as the default model for free users.
OpenAI said it has improved the model's performance across many of ChatGPT's most widely used languages. However, it acknowledged that some languages may currently exhibit non-native accents or reduced fluency, adding that further improvements are ongoing.

