Published on · updated on · Speech recognition
Why doesn't my voice assistant understand me?
Noise, accent, unusual words or a cut-off turn: where the chain breaks between what you say and what a voice assistant understands, and what you can do about it.


You ask the assistant on your phone, in your car or in your kitchen for something, and it responds with something else, or with “sorry, I didn’t catch that.” It’s tempting to think we’re speaking badly. It’s almost never that. There are several steps between what you say and what the assistant does, and the failure can be in any of them.
What happens between speaking and getting an answer
In a very short time, a voice assistant does roughly this:
- It detects that you’re talking to it, through a wake word or because you press a button.
- It decides when you’ve finished speaking.
- It turns your voice into text. That’s automatic speech recognition.
- It interprets that text: what you want and with what details.
- It decides what to do and responds.
Errors in step three are measured with the word error rate: how many words are added, missing or changed compared with what was actually said (Jurafsky and Martin, ch. 16). A small error in a key word can be enough for everything else to go wrong.
The microphone and the noise
The first problem is physical. Speaking into a phone held to your mouth is not the same as speaking to a speaker across the living room, with the TV on and the dishwasher running. With distance, the voice arrives weaker and mixed with the room’s echo. In the car you add the engine, the air conditioning and the road.
Today’s systems handle noise much better than those of a few years ago. Whisper, OpenAI’s model, was trained on 680,000 hours of all kinds of audio, and that volume gave it a robustness that didn’t exist before (Radford et al., 2022). But noise is still the first cause of failure we look for when an assistant doesn’t understand in its real environment.
Your accent, your age and the way you speak
A recognition system learns from the voices it’s trained on. If it has heard little of an accent or a way of speaking, it makes more mistakes with it.
In 2020, a Stanford University study analyzed the systems from Amazon, Apple, Google, IBM and Microsoft using interviews with white and African American speakers in the United States. The average word error rate was 0.35 for African American speakers and 0.19 for white speakers: almost double (Koenecke et al., 2020). The cause lay in the acoustic models, trained with little data from those voices.
The same happens with Spanish varieties. A 2026 study of YouTube’s automatic captions in seven Latin American countries found error rates ranging from 16% for Puerto Rican women to 24% for speakers from Argentina (Jimenez and Kern, 2026). Same language, same system, and a notable difference depending on where you’re from.
Age and health matter too. Children’s and older people’s voices tend to be underrepresented in the data. And a study of Whisper found that about 1% of transcripts included made-up phrases nobody had said, more often for people with aphasia, who take long pauses when speaking (Koenecke et al., 2024).
Words the system doesn’t know
Even if it hears you perfectly, a system has to know which words exist. Proper names, brands, a company’s products or technical terms are tricky territory. If you ask for “the Tarifa Plus 30 plan” and the system has never seen it written, it will most likely write something that sounds similar and that it does know.
Most services let you add custom vocabulary or provide context so they recognize those words better. It’s one of the first things we set up in an assistant for a specific company.
It cut you off, or it’s still waiting
Step two on the list, deciding when you’ve finished, fails more often than you’d think. If you pause to think, the assistant may decide your sentence is over and respond to half of it. If it waits too long, the conversation feels slow. Long numbers, like a phone number or an ID number, are a typical case: people pause between groups of digits and the system cuts them off halfway.
It heard you, but didn’t understand you
Sometimes the transcript is perfect and the error is in the interpretation. “Set the alarm for seven” is clear, but “wake me up early tomorrow” requires working out what early means. Rule- and intent-based assistants failed a lot here, because they only recognized the sentences someone had anticipated. Today’s language models understand many more ways of saying the same thing, although they can also confidently interpret something you didn’t say.
What you can do
If you use an assistant, speak naturally, without shouting or over-enunciating, and move closer to the microphone when it’s noisy. If you build one, there’s much more room for improvement:
- Measure recognition with recordings of your real users, in their environment, broken down by group: accents, ages, channel. The average hides the people who struggle most.
- Test with real noise, not just in the office.
- Add your own vocabulary: products, brands, place names.
- Confirm important details before acting, such as an amount or a date.
- When something fails, help the person move forward instead of repeating “I didn’t catch that.”
- Offer another route when voice isn’t a good idea, such as a screen or a keyboard.
We run this kind of testing when we evaluate assistants and agents before they launch. There’s almost always some group of users for whom the assistant works quite a bit worse than in internal testing.
References
- Jurafsky, D. and Martin, J. H. Speech and Language Processing, chapter 16: Automatic Speech Recognition (3rd edition, draft).
- Radford, A., Kim, J. W., Xu, T. et al. (2022). Robust Speech Recognition via Large-Scale Weak Supervision. arXiv.
- Koenecke, A., Nam, A., Lake, E. et al. (2020). Racial disparities in automated speech recognition. PNAS, 117(14), 7684-7689.
- Jimenez, I. D. and Kern, C. (2026). Dialect and Gender Bias in YouTube’s Spanish Captioning System. arXiv.
- Koenecke, A., Choi, A. S. G., Mei, K. X., Schellmann, H. and Sloane, M. (2024). Careless Whisper: Speech-to-Text Hallucination Harms. FAccT 2024.
Share this article: