You’re reading out your phone number to a voice agent. You say the first three digits, pause to remember the next ones, and the agent replies: “Sorry, that number isn’t valid.” Or the other way around: you finish speaking and a second goes by, then two, and you don’t know if it heard you. Both failures have the same root. Knowing when it’s each person’s turn to speak is one of the hardest parts of a conversation, and humans do it so well we don’t even notice.
How people do it
In 1974, sociologists Harvey Sacks, Emanuel Schegloff and Gail Jefferson described something that seems obvious today: in a conversation, almost always only one person speaks at a time, and turns change with very little silence and very little overlap, without anyone organizing it (Sacks, Schegloff and Jefferson, 1974). Decades later, a study of ten languages from around the world found the same pattern in all of them (Stivers et al., 2009).
The surprising part is the speed. Stephen Levinson and Francisco Torreira calculated that the typical gap between two turns is 100 to 300 milliseconds, while preparing a sentence to say it takes at least 600 (Levinson and Torreira, 2015). The math doesn’t add up unless we start preparing our answer while the other person is still talking. And that’s exactly what we do: we anticipate when the other person will finish from what they say, how they say it and the context.
We also use signals to give up or claim the turn. A direct question hands over the floor. An “uh-huh” or “right” while the other person talks shows we’re still listening, without taking their turn. Lowering our intonation at the end of a sentence signals that we’re done. And when two people start talking at once, one of them stops almost immediately.
How machines try to do it
The simplest way to decide someone has finished speaking is to wait for silence. Many systems detect when there’s speech and when there isn’t, and if the silence lasts longer than a threshold, they treat the turn as over. In OpenAI’s Realtime API, for example, that threshold is 500 milliseconds by default, and the documentation itself warns that shorter values make the model respond sooner but may interrupt during short pauses (OpenAI).
The problem is that silence isn’t a good signal. We pause to think, to remember a detail or to find the right word. And sometimes we finish a sentence with no pause at all, tacking on a “right?”
That’s why current systems combine silence with what’s being said. If the sentence sounds incomplete (“I’d like to move my appointment on the…”), they wait longer. If it sounds finished, they respond sooner. OpenAI calls this semantic turn detection (OpenAI). In 2024 LiveKit released an open model that decides when a turn ends from the text, and its current version listens to the audio directly so it can also take intonation and rhythm into account (LiveKit).
The next step is models that listen while they speak, like Moshi, which can hear an “uh-huh” or an interruption without stopping to wait for their turn (Défossez et al., 2024). We cover this in three ways to build a voice agent.
What design can do
The technology has improved a lot, but conversation design still prevents many of the problems.
End the turn clearly
If the agent ends with a direct question, the person knows it’s their turn. If it ends with a long statement, they may wait for it to continue. It’s one of the first tips we gave back in 2018 for Alexa apps (eight tips for building good Alexa skills in Spanish), and it still holds.
Allow more time when the answer needs it
A phone number, an ID number, an address or a date are said with pauses. At those moments the agent should wait longer before treating the turn as over, or ask for the details in parts.
Let people interrupt
If the agent is explaining something and the person starts talking, the natural thing is to stop and listen. An agent that keeps talking over people is very annoying. But you have to tell an interruption apart from an “uh-huh” or background noise, or the agent will cut itself off every time someone coughs.
Don’t leave silences unexplained
When the agent needs to look something up and it’ll take a moment, it should say so: “One moment, let me check.” A two-second silence on a call sounds like a failure.
Short answers
The longer the agent’s turn, the more likely the person will want to interrupt and the harder it is to remember what was said. We explain this in what reads well doesn’t always sound good.
Test it with people
Reading scripts is a poor way to test turn-taking. You need to listen to real conversations, with people who hesitate, correct themselves, talk over background noise or read out numbers. That’s where the cut-offs and awkward silences you don’t see in internal testing show up. It’s part of our work in conversation and agent design.
References
- Sacks, H., Schegloff, E. A. and Jefferson, G. (1974). A simplest systematics for the organization of turn-taking for conversation. Language, 50(4), 696-735.
- Stivers, T., Enfield, N. J., Brown, P. et al. (2009). Universals and cultural variation in turn-taking in conversation. PNAS, 106(26), 10587-10592.
- Levinson, S. C. and Torreira, F. (2015). Timing in turn-taking and its implications for processing models of language. Frontiers in Psychology, 6, 731.
- OpenAI. Voice activity detection (VAD). Realtime API documentation, accessed September 2026.
- LiveKit. Solving end-of-turn detection: LiveKit Turn Detector v1.0.
- Défossez, A., Mazaré, L., Orsini, M. et al. (2024). Moshi: a speech-text foundation model for real-time dialogue. arXiv.