Reviewed September 29, 2026. For years, a voice assistant could be drawn as three boxes: speech recognition turns audio into text, a language model decides what to say, and text-to-speech turns the answer back into audio. The sequence is often summarized as “listen, think, speak.” It remains useful, especially when a workflow needs a transcript, strict checks or replaceable components. But it is no longer the only serious design for voice AI.
Newer systems can process live audio directly, begin work before a turn is neatly finished, and keep listening while they speak. The change is not simply a faster synthetic voice. It is a move from exchanging audio files or completed turns to managing a continuous conversation: pauses, interruptions, emphasis, background sounds and all.
What is a “listen, think, speak” voice pipeline?
A conventional voice AI pipeline usually has three core stages:
- Listen: automatic speech recognition (ASR), also called speech-to-text, converts the user’s audio into words.
- Think: a language model or dialogue system interprets those words, may call tools, and produces a text response.
- Speak: text-to-speech (TTS) synthesizes the response as audio.
In a basic implementation, each stage waits for the previous one to finish. That makes the architecture understandable and controllable, but it also makes every boundary a possible delay or information bottleneck. Modern cascaded systems can stream partial results between stages, so “pipeline” should not be confused with “slow by definition.” The real distinction is whether text remains the mandatory handoff between listening, reasoning and speaking.
Why the three-stage model feels less like conversation
Latency accumulates at boundaries
A voice system must decide whether a pause means “I am finished” or merely “I am thinking.” If it waits too long, the conversation drags. If it commits too early, it talks over the user. After endpoint detection, recognition, model inference, speech synthesis, networking and any tool calls each consume part of the response-time budget.
Human timing sets a demanding reference point. A 2009 study of question-and-answer sequences across ten languages found a shared tendency to minimize silence and overlap, with response timing varying within a relatively narrow range. This does not create a universal 200-millisecond product requirement: human-to-human conversation is not the same as a support call, a translation session or a voice agent using a slow database. It does explain why an extra second of unexplained silence is easy to notice.
OpenAI offered a useful historical comparison when it introduced GPT-4o in 2024. Its earlier ChatGPT Voice Mode used three separate models and averaged 2.8 seconds with GPT-3.5 and 5.4 seconds with GPT-4, while the end-to-end GPT-4o model was reported to respond to audio in as little as 232 milliseconds and 320 milliseconds on average. Those figures describe that launch and its test conditions, not a guarantee for every network, device or tool-using application. They show why architecture became part of the user experience.
A transcript is a compressed version of speech
Text captures words well, but a conventional transcript may discard cues that change how those words should be understood: hesitation, stress, tempo, laughter, accent, overlapping speakers and non-speech sounds. “That’s fine” can signal agreement, resignation or irritation. A text-only reasoning stage may receive the same sentence in all three cases unless another component preserves those acoustic cues.
This does not mean every system should infer emotion, or that vocal style reveals a person’s intent reliably. Emotion recognition can be culturally brittle and inappropriate in high-stakes settings. The narrower point is architectural: audio contains information that plain text does not, so a design that throws away the audio cannot use it later.
Conversation is not a clean sequence of turns
People interrupt, overlap, restart sentences and use short backchannels such as “right,” “mm-hm” or “I see.” A half-duplex assistant treats one side as active at a time: it listens, stops listening, then speaks. A full-duplex system can continue monitoring the user while producing audio, allowing it to stop, yield or adapt when the user speaks.
This is harder than adding an interrupt button. The system has to distinguish a real barge-in from background speech or a harmless backchannel, stop unplayed audio, and keep its conversation state consistent with what the user actually heard. Full-Duplex-Bench, introduced in 2025, reflects this shift in evaluation by testing turn-taking behavior rather than only transcript or answer quality.
What is replacing the old pipeline?
There is no single replacement. Voice AI is branching into three broad patterns.
1. Streaming cascades
ASR, language reasoning and TTS remain separate, but they no longer operate as three sealed boxes. Recognition emits partial text, the model can prepare work incrementally, and speech synthesis can start from the first stable part of a response. A turn model uses both silence and linguistic context to estimate whether the user is done.
This approach preserves explicit transcripts, modular vendors and familiar policy checkpoints. It is often the practical choice for structured support, regulated workflows and existing text agents. Its quality depends on orchestration: partial transcripts can change, speculative work can be wrong, and aggressive endpointing can save time while increasing interruptions.
2. Native speech-to-speech models
A speech-native model accepts audio representations and generates audio without requiring a transcript as the central reasoning interface. Text may still be generated for captions, logging or internal assistance, but the model can work from acoustic information directly.
Kyutai’s Moshi research is a clear public example. It represents user and assistant audio as parallel streams and was designed to listen and generate continuously. The paper reports a theoretical latency of 160 milliseconds and about 200 milliseconds in practice for its research system. Moshi also uses time-aligned text in an “Inner Monologue,” a reminder that speech-native does not necessarily mean text-free.
3. Hybrid voice front ends with agent back ends
The emerging production pattern is often hybrid. A live audio model manages timing, turn-taking and spoken delivery, while a separate agent or service handles deeper reasoning, retrieval, permissions and tool use. The voice layer can acknowledge the user or ask a clarifying question while the back end checks a calendar, looks up an order or plans several steps.
This division matters for devices as well as call centers. In an AI agent phone, voice may be the quickest way to state a goal, but usefulness depends on what happens after the request: context, tools, permissions, execution and recovery. Meydo’s overview of Meydo OS and DroiClaw describes a related system-level goal—connecting multimodal input to permission-based action—rather than treating voice as a standalone chat feature.
The new unit of design is the interaction loop
Once audio is continuous, teams have to design more than recognition accuracy and voice quality. The interaction loop includes:
- Turn detection: Is the user pausing, finished, or expecting a response?
- Barge-in: Should user speech cancel, pause or redirect the current answer?
- Backchannels: Can the system acknowledge without taking over the turn?
- Grounding: Which words, audio cues, screen context and tool results support the response?
- Action boundaries: Which tasks can proceed, and which require confirmation?
- Recovery: Can the user correct a name, undo an action or resume after an interruption?
- State: Does the model remember what was spoken, what was played and what a tool actually completed?
A voice that sounds natural can make errors feel more trustworthy, not less. For consequential actions, conversational speed should not remove confirmation, authorization or a visible record. A fast “done” is useful only after the underlying action has been verified.
Why cascaded systems are not disappearing
Native audio has real advantages, but the architectural decision is a tradeoff, not a maturity ranking. OpenAI’s current voice-agent guidance presents speech-to-speech sessions as a fit for natural, low-latency interaction and chained pipelines as a fit for predictable workflows, existing text agents and explicit control over intermediate stages.
A cascade may be the better choice when a team needs:
- a durable, inspectable transcript;
- deterministic text checks before anything is spoken;
- independent replacement of ASR, model or voice vendors;
- language-specific components or custom vocabulary;
- clear audit points around tools and regulated decisions; or
- an economical voice layer around a proven text workflow.
A native or hybrid live-audio design may be preferable when interruption, expressive delivery, rapid turn-taking or acoustic context is central to the experience. Many products will support both: a fluid voice path for ordinary conversation and a controlled path for sensitive actions.
How to evaluate a modern voice AI system
Do not judge it from a polished one-turn demo. Test the entire loop under realistic conditions.
- Measure more than one latency. Record end-of-turn to first audio, time to a useful answer and time to completed action. Report typical and slow cases, not only the best run.
- Interrupt it. Try early corrections, late barge-ins and short backchannels. Confirm that playback stops and that the next answer does not rely on words the user never heard.
- Vary the audio. Use accents, code-switching, names, numbers, quiet speech, noise and multiple speakers that reflect the intended users.
- Inspect tools and confirmations. Verify arguments before an action and read the target system back afterward. Spoken confidence is not evidence of completion.
- Check transcript behavior. If transcripts are stored, ask whether they are authoritative or only a rough representation of what the audio model understood.
- Test failures. Disconnect the network, delay a tool, deny a permission and correct a mistaken entity. Recovery quality often matters more than the ideal path.
- Review privacy and retention. Determine where raw audio, transcripts, voiceprints and tool results are processed, retained and accessible.
What comes after “listen, think, speak”?
The next model is closer to listen continuously, interpret incrementally, act carefully and speak adaptively. Those activities can overlap. The assistant may detect that a turn is ending while preparing a response, continue listening while it talks, and delegate a task while keeping the user informed.
That does not eliminate the old pipeline. It removes its status as the default mental model for every voice experience. The winning architecture will be the one that fits the task: native audio where timing and expression matter, modular stages where inspection and control matter, and hybrids where conversation must connect to reliable action.
Frequently asked questions
What is speech-to-speech AI?
Speech-to-speech AI accepts spoken input and produces spoken output while processing audio directly. A transcript may still be available, but it is not necessarily the mandatory bridge between an ASR system, a text model and TTS.
What does full duplex mean in voice AI?
Full duplex means the system can listen and produce audio at the same time. In practice, a useful full-duplex agent also needs to recognize interruptions and backchannels, stop or adapt its output, and maintain accurate state after overlapping speech.
Are native audio models always faster than ASR–LLM–TTS pipelines?
No. Native models remove some serial boundaries, but real latency also depends on model size, hardware, network transport, endpointing, tool calls and playback. A well-streamed cascade can outperform a poorly deployed native model. Measure the complete application.
Do speech-native models still use text?
They can. Text may support reasoning, captions, search, safety checks or logs. “Speech-native” means the model can consume or generate audio representations directly; it does not require every internal or external representation to be audio-only.
Is a more human-sounding voice AI more accurate?
No. Natural timing and expressive speech improve interaction, but they do not verify facts or actions. Accuracy, tool results, permissions and post-action checks must be evaluated separately.
Disclosure and sources
This article is an analysis of public technical documentation and research, not a benchmark of a specific commercial system. Meydo has a commercial interest in personal AI devices. Performance depends on models, hardware, network conditions, language and application design.
- OpenAI: Hello GPT-4o (May 13, 2024)
- OpenAI API: Voice agents—speech-to-speech and chained architectures (accessed September 29, 2026)
- OpenAI API: Realtime conversations (accessed September 29, 2026)
- Kyutai: Moshi—A Speech-Text Foundation Model for Real-Time Dialogue (2024)
- Full-Duplex-Bench: A Benchmark to Evaluate Full-Duplex Spoken Dialogue Models on Turn-taking Capabilities (2025)
- Stivers et al.: Universals and cultural variation in turn-taking in conversation (PNAS, 2009)
