
Is it my turn to speak?
Why voice agents talk over people or go quiet: how humans manage turn-taking and how machines try to detect it.
Published on · updated on · Voice agents
Cascaded, speech-to-speech or hybrid voice agents: how each architecture works, what you gain and lose in latency, control and voice, and when to choose each.


A voice agent listens, understands, decides and answers out loud. There are three ways to organize those pieces: in a chain of separate components, with a single model that goes from speech to speech, or with a mix of the two. Each one solves a problem and creates another. The choice depends mostly on what matters to you most: speed, control, voice or cost.
This is the classic architecture. The audio goes through a speech recognizer that turns it into text. A language model reads that text and decides what to answer, calling whatever tools it needs. And a synthesizer turns the answer into speech.
It has clear advantages:
And two serious drawbacks. The first is latency: each piece adds its own time, and the total is noticeable. In a conversation between people, the silences between turns are very short, on the order of tenths of a second. A study of ten languages from around the world found that same pattern in all of them: we avoid overlapping and we avoid long silences (Stivers et al., 2009). A cascade that takes a second and a half to answer breaks that rhythm.
The second is that converting to text loses how things were said. An angry tone, hesitation or irony don’t reach the language model, and the voice that answers doesn’t know whether it should sound reassuring or cheerful.
In May 2024, OpenAI introduced GPT-4o, a model that works directly with audio and responded in 320 milliseconds on average, close to the pace of human conversation (OpenAI, 2024). In September, the French lab Kyutai released Moshi, an open model that can listen and speak at the same time, with a practical latency of about 200 milliseconds (Défossez et al., 2024). Since August 2025, OpenAI has offered its speech-to-speech model for production use, with tool calling and phone connectivity (OpenAI, 2025).
Google followed a similar path. In August 2024 it launched Gemini Live, the voice conversation mode of the Gemini app, where you can interrupt the assistant while it’s talking (TechCrunch, 2024). For developers it offers the Live API, with native audio models that listen and respond by voice without going through text. The latest, Gemini 3.8 Live, introduced in September 2026, can call tools without interrupting the conversation, accept images as context and speak more than 97 languages (Google, 2026).
What you gain:
What you lose:
Quality in Spanish, and especially in underrepresented accents, varies a lot from one model to another. Test it with your users before deciding.
The most common hybrid architecture merges part of the chain into a single model. Instead of a separate speech recognizer and language model, a single model listens to the audio directly and answers in text. That text goes to a synthesizer that generates the voice as the words arrive, in streaming, without waiting for the sentence to end. It’s usually called a half-cascade.
Ultravox is a good example. Its open-weights model combines an audio encoder with a language model: it turns audio directly into the representation the language model uses, without transcribing it, and streams text that’s then turned into speech (Ultravox). Its creators justify it like this: by skipping the transcript, the cues about how the person speaks aren’t lost (Ultravox).
You can also build it with some of the big providers’ real-time models, configured to respond in text instead of audio, plus your own synthesizer. Not all of them allow it: Gemini’s native audio models, for example, only respond by voice. LiveKit documents this as a half-cascade for Gemini Live and other real-time models: you take advantage of how a speech-to-speech model understands audio and keep full control of the output voice (LiveKit).
What you gain:
What you lose compared with a full speech-to-speech model: the output voice doesn’t hear the person, so it adapts its tone less well to how they spoke, and the synthesizer still adds some latency, although streaming reduces it a lot.
You don’t need to build the chain from scratch. There are open frameworks that handle the common parts: sending and receiving audio in real time, detecting when the person has finished speaking, handling interruptions and connecting to the phone network.
With either one, switching architecture or provider is a matter of configuration, which makes it easier to try several options with your users before deciding.
| Cascaded | Speech-to-speech | Hybrid | |
|---|---|---|---|
| Pieces | Recognizer, language model and synthesizer | One model | Audio-to-text model and synthesizer |
| Latency | Highest | Lowest | In between |
| Own or brand voice | Yes | Usually not | Yes |
| Tone and emotion | Lost when converted to text | Preserved | Preserved when listening, not when speaking |
| Reviewing what happened | Easy | Hard | Easy: the answer is text |
| Tasks with many rules | Most reliable | Less predictable | Depends on the model |
If your agent represents a brand with its own voice, works in a regulated sector or follows processes with many rules, start with a well-optimized cascade or a hybrid architecture. If what matters most is that the conversation flows, for example for companionship, language practice or simple customer service, a speech-to-speech model can give you an experience that’s hard to match.
In every case, measure. The latency that matters is the one the person on the other end of the line perceives, with their connection and their noise, not the one on the provider’s spec sheet.
At Monoceros we design and build voice agents with any of these three architectures, and we choose with each client the one that fits their case, their rules and their budget. When the agent needs to sound like no one else, we create custom synthetic voices, cloned with permission or designed from scratch, that can be used in a cascade or a hybrid architecture. And before launch, we evaluate it with real conversations: latency, interruptions, recognition errors and compliance with business rules.
Share this article: