shiichan

Talk and listen at the same time! ChatGPT's voice evolves into full-duplex "GPT-Live"!

Hey everyone, it's Shii! I found some news about a big evolution in ChatGPT's voice experience, and it got me really excited! Let me share it with you.

OpenAI News openai.com

What was announced?

OpenAI News announced a new voice model architecture called GPT-Live. Its headline feature is full-duplex processing, meaning it can listen and speak at the same time, which is a big step up from previous cascaded or turn-based voice models. It's rolling out globally as ChatGPT's voice feature on iOS, Android, and ChatGPT.com. There are two versions: GPT-Live-1 for paid tiers, and a lighter GPT-Live-1 mini for free users.

The story so far

Previous voice interactions were either cascaded, chaining speech recognition, text processing, and speech synthesis one after another, or turn-based, where one side had to finish speaking before the other could respond. With that setup, it was hard to reproduce things that come naturally in human conversation, like nodding along while someone talks or jumping in mid-sentence.

What changes

GPT-Live's full-duplex processing lets it handle input and output continuously and simultaneously. That means it can make decisions like speaking, listening, pausing, interrupting, or calling a tool multiple times a second, in real time. This enables more human-like exchanges: natural verbal cues like "mhmm" or "yeah," strategic silence while the user thinks, and visual response cards for things like weather, stocks, or sports. You can also pick from three reasoning levels, Instant, Medium, and High, and it supports live translation.

Dive Deep

Another interesting architectural piece is how it delegates complex work. When a task needs web search, deep reasoning, or agentic work, GPT-Live hands that off to a backend model (GPT-5.5) while keeping the conversation's pace uninterrupted. It's a balance of running heavy processing in the background while preserving the natural rhythm of the conversation.

On the safety side, it includes dedicated safety training covering topics like self-harm, psychosis, emotional reliance, violence, and sexual content, with real-time safeguards that can redirect unsafe outputs mid-conversation. Given how real-time a modality voice is, it makes sense that this kind of safety work gets special attention.

Note that API access for developers and enterprises currently requires registering through a dedicated form, so it's not yet something everyone can use right away. Also, voice combined with video or screen sharing isn't supported yet.

Wrap-up

  • GPT-Live is a new full-duplex voice model that can listen and speak at the same time
  • Complex work is delegated to a backend GPT-5.5 model while conversational pace is preserved
  • Rolling out globally as ChatGPT's voice feature on iOS, Android, and ChatGPT.com (GPT-Live-1 / GPT-Live-1 mini)
  • API access requires registration; voice with video/screen sharing isn't supported yet
  • Includes dedicated safety training around self-harm, violence, and sexual content

For anyone who's wanted a voice assistant that feels like a genuinely natural conversation, this news feels like a pretty big step forward!