Published on · updated on · Voice agents

Three ways to build a voice agent, and none is perfect

Cascaded, speech-to-speech or hybrid voice agents: how each architecture works, what you gain and lose in latency, control and voice, and when to choose each.

3 ways to build a voice agent

A voice agent listens, understands, decides and answers out loud. There are three ways to organize those pieces: in a chain of separate components, with a single model that goes from speech to speech, or with a mix of the two. Each one solves a problem and creates another. The choice depends mostly on what matters to you most: speed, control, voice or cost.

Cascaded: one piece for each job

This is the classic architecture. The audio goes through a speech recognizer that turns it into text. A language model reads that text and decides what to answer, calling whatever tools it needs. And a synthesizer turns the answer into speech.

It has clear advantages:

  • You can choose each piece. The recognizer that best understands your users, the language model you prefer and the voice you want, including a custom brand voice.
  • Everything goes through text, so you can read what the system understood and what it decided. Reviewing errors and passing audits is much easier.
  • Text language models are currently the most reliable at following instructions and using tools.

And two serious drawbacks. The first is latency: each piece adds its own time, and the total is noticeable. In a conversation between people, the silences between turns are very short, on the order of tenths of a second. A study of ten languages from around the world found that same pattern in all of them: we avoid overlapping and we avoid long silences (Stivers et al., 2009). A cascade that takes a second and a half to answer breaks that rhythm.

The second is that converting to text loses how things were said. An angry tone, hesitation or irony don’t reach the language model, and the voice that answers doesn’t know whether it should sound reassuring or cheerful.

Speech-to-speech: a single model

In May 2024, OpenAI introduced GPT-4o, a model that works directly with audio and responded in 320 milliseconds on average, close to the pace of human conversation (OpenAI, 2024). In September, the French lab Kyutai released Moshi, an open model that can listen and speak at the same time, with a practical latency of about 200 milliseconds (Défossez et al., 2024). Since August 2025, OpenAI has offered its speech-to-speech model for production use, with tool calling and phone connectivity (OpenAI, 2025).

Google followed a similar path. In August 2024 it launched Gemini Live, the voice conversation mode of the Gemini app, where you can interrupt the assistant while it’s talking (TechCrunch, 2024). For developers it offers the Live API, with native audio models that listen and respond by voice without going through text. The latest, Gemini 3.8 Live, introduced in September 2026, can call tools without interrupting the conversation, accept images as context and speak more than 97 languages (Google, 2026).

What you gain:

  • Speed. There aren’t three pieces waiting on each other.
  • The way of speaking goes in and out of the model. It can notice someone hesitating and respond in a fitting tone.
  • Interruptions are handled more naturally, especially in models that listen while they speak.

What you lose:

  • The voice. You usually choose among the voices the provider offers. Your own brand voice, or a voice cloned with permission, usually isn’t available.
  • Control and visibility. It’s harder to know why it said what it said, and harder to fix a specific pronunciation.
  • Reliability on complex tasks. They’ve improved a lot at following instructions and using tools, but in long processes with many rules, a text model is still more predictable.
  • Cost predictability. Audio is billed differently from text, so you have to work it out per minute of conversation and compare it with the cascade.

Quality in Spanish, and especially in underrepresented accents, varies a lot from one model to another. Test it with your users before deciding.

Hybrid: fewer pieces, the voice you want

The most common hybrid architecture merges part of the chain into a single model. Instead of a separate speech recognizer and language model, a single model listens to the audio directly and answers in text. That text goes to a synthesizer that generates the voice as the words arrive, in streaming, without waiting for the sentence to end. It’s usually called a half-cascade.

Ultravox is a good example. Its open-weights model combines an audio encoder with a language model: it turns audio directly into the representation the language model uses, without transcribing it, and streams text that’s then turned into speech (Ultravox). Its creators justify it like this: by skipping the transcript, the cues about how the person speaks aren’t lost (Ultravox).

You can also build it with some of the big providers’ real-time models, configured to respond in text instead of audio, plus your own synthesizer. Not all of them allow it: Gemini’s native audio models, for example, only respond by voice. LiveKit documents this as a half-cascade for Gemini Live and other real-time models: you take advantage of how a speech-to-speech model understands audio and keep full control of the output voice (LiveKit).

What you gain:

  • One piece fewer than the cascade, so less waiting.
  • The model receives the original audio, with the person’s tone and hesitations, not just a transcript.
  • You choose the voice, including your own brand voice.
  • The answer exists as text, so it can be reviewed, filtered and stored.

What you lose compared with a full speech-to-speech model: the output voice doesn’t hear the person, so it adapts its tone less well to how they spoke, and the synthesizer still adds some latency, although streaming reduces it a lot.

What to build them with

You don’t need to build the chain from scratch. There are open frameworks that handle the common parts: sending and receiving audio in real time, detecting when the person has finished speaking, handling interruptions and connecting to the phone network.

  • Pipecat, maintained by Daily and its community, supports both component chains and speech-to-speech models from OpenAI, Gemini, Ultravox and other providers (Pipecat).
  • LiveKit Agents, open source, works for cascades, for speech-to-speech models such as OpenAI Realtime, Gemini Live or Ultravox, and for the half-cascade. It includes support for phone calls (LiveKit).

With either one, switching architecture or provider is a matter of configuration, which makes it easier to try several options with your users before deciding.

Quick comparison

CascadedSpeech-to-speechHybrid
PiecesRecognizer, language model and synthesizerOne modelAudio-to-text model and synthesizer
LatencyHighestLowestIn between
Own or brand voiceYesUsually notYes
Tone and emotionLost when converted to textPreservedPreserved when listening, not when speaking
Reviewing what happenedEasyHardEasy: the answer is text
Tasks with many rulesMost reliableLess predictableDepends on the model

How to choose

If your agent represents a brand with its own voice, works in a regulated sector or follows processes with many rules, start with a well-optimized cascade or a hybrid architecture. If what matters most is that the conversation flows, for example for companionship, language practice or simple customer service, a speech-to-speech model can give you an experience that’s hard to match.

In every case, measure. The latency that matters is the one the person on the other end of the line perceives, with their connection and their noise, not the one on the provider’s spec sheet.

How we help

At Monoceros we design and build voice agents with any of these three architectures, and we choose with each client the one that fits their case, their rules and their budget. When the agent needs to sound like no one else, we create custom synthetic voices, cloned with permission or designed from scratch, that can be used in a cascade or a hybrid architecture. And before launch, we evaluate it with real conversations: latency, interruptions, recognition errors and compliance with business rules.

References

Written by

Carlos Muñoz-Romero

Co-founder at Monoceros Labs

Co-founder of Monoceros Labs. A computer engineer with a master's in data science, he leads Fonos's voice technology. Previously Chief Innovation Officer at BEEVA (now BBVA Technology).

Share this article:

Related posts

Is it my turn to speak?

Is it my turn to speak?

Why voice agents talk over people or go quiet: how humans manage turn-taking and how machines try to detect it.