Published on · updated on · Voice agents Conversation design

Is it my turn to speak?

Why voice agents talk over people or go quiet: how humans manage turn-taking and how machines try to detect it.

Is it my turn to speak?

You’re reading out your phone number to a voice agent. You say the first three digits, pause to remember the next ones, and the agent replies: “Sorry, that number isn’t valid.” Or the other way around: you finish speaking and a second goes by, then two, and you don’t know if it heard you. Both failures have the same root. Knowing when it’s each person’s turn to speak is one of the hardest parts of a conversation, and humans do it so well we don’t even notice.

How people do it

In 1974, sociologists Harvey Sacks, Emanuel Schegloff and Gail Jefferson described something that seems obvious today: in a conversation, almost always only one person speaks at a time, and turns change with very little silence and very little overlap, without anyone organizing it (Sacks, Schegloff and Jefferson, 1974). Decades later, a study of ten languages from around the world found the same pattern in all of them (Stivers et al., 2009).

The surprising part is the speed. Stephen Levinson and Francisco Torreira calculated that the typical gap between two turns is 100 to 300 milliseconds, while preparing a sentence to say it takes at least 600 (Levinson and Torreira, 2015). The math doesn’t add up unless we start preparing our answer while the other person is still talking. And that’s exactly what we do: we anticipate when the other person will finish from what they say, how they say it and the context.

We also use signals to give up or claim the turn. A direct question hands over the floor. An “uh-huh” or “right” while the other person talks shows we’re still listening, without taking their turn. Lowering our intonation at the end of a sentence signals that we’re done. And when two people start talking at once, one of them stops almost immediately.

How machines try to do it

The simplest way to decide someone has finished speaking is to wait for silence. Many systems detect when there’s speech and when there isn’t, and if the silence lasts longer than a threshold, they treat the turn as over. In OpenAI’s Realtime API, for example, that threshold is 500 milliseconds by default, and the documentation itself warns that shorter values make the model respond sooner but may interrupt during short pauses (OpenAI).

The problem is that silence isn’t a good signal. We pause to think, to remember a detail or to find the right word. And sometimes we finish a sentence with no pause at all, tacking on a “right?”

That’s why current systems combine silence with what’s being said. If the sentence sounds incomplete (“I’d like to move my appointment on the…”), they wait longer. If it sounds finished, they respond sooner. OpenAI calls this semantic turn detection (OpenAI). In 2024 LiveKit released an open model that decides when a turn ends from the text, and its current version listens to the audio directly so it can also take intonation and rhythm into account (LiveKit).

The next step is models that listen while they speak, like Moshi, which can hear an “uh-huh” or an interruption without stopping to wait for their turn (Défossez et al., 2024). We cover this in three ways to build a voice agent.

What design can do

The technology has improved a lot, but conversation design still prevents many of the problems.

End the turn clearly

If the agent ends with a direct question, the person knows it’s their turn. If it ends with a long statement, they may wait for it to continue. It’s one of the first tips we gave back in 2018 for Alexa apps (eight tips for building good Alexa skills in Spanish), and it still holds.

Allow more time when the answer needs it

A phone number, an ID number, an address or a date are said with pauses. At those moments the agent should wait longer before treating the turn as over, or ask for the details in parts.

Let people interrupt

If the agent is explaining something and the person starts talking, the natural thing is to stop and listen. An agent that keeps talking over people is very annoying. But you have to tell an interruption apart from an “uh-huh” or background noise, or the agent will cut itself off every time someone coughs.

Don’t leave silences unexplained

When the agent needs to look something up and it’ll take a moment, it should say so: “One moment, let me check.” A two-second silence on a call sounds like a failure.

Short answers

The longer the agent’s turn, the more likely the person will want to interrupt and the harder it is to remember what was said. We explain this in what reads well doesn’t always sound good.

Test it with people

Reading scripts is a poor way to test turn-taking. You need to listen to real conversations, with people who hesitate, correct themselves, talk over background noise or read out numbers. That’s where the cut-offs and awkward silences you don’t see in internal testing show up. It’s part of our work in conversation and agent design.

References

Written by

Nieves Ábalos

Co-founder at Monoceros Labs

Co-founder of Monoceros Labs. A computer engineer researching dialogue systems since 2009, now doing a PhD in AI at the University of Granada.

Share this article:

Related posts

3 ways to build a voice agent

Three ways to build a voice agent, and none is perfect

Cascaded, speech-to-speech or hybrid voice agents: how each architecture works, what you gain and lose in latency, control and voice, and when to choose each.

Before you build

What to ask yourself before building an assistant

Before choosing technology for a chatbot or voice assistant: which conversations you want to have, where, through which channel, how you'll measure success and what to prototype.

6 adjustments for writing for the ear

What reads well doesn't always sound good

Text written for reading usually sounds wrong in a voice assistant. Six adjustments for writing for the ear: sentences, options, punctuation, numbers and pauses.