An AI assistant has no opinions of its own, but it does learn the ones in its data. If the texts it was trained on mostly show women at home and men at the office, the model will repeat that. If the voices used to train a speech recognizer were almost all from one accent, the model will understand people with other accents less well. Bias isn’t a rare glitch: it’s what you should expect if nobody looks for it.
Where it comes from
There are three usual ways in:
- The data. A model learns from what it sees, imbalances included. What’s underrepresented is learned less well, and stereotypes that appear a lot are learned very well.
- Design decisions. Which voice is the default, what the assistant is called, how it responds to an insult. Nobody decides out of malice, but every choice sends a message.
- How it’s evaluated. If you only measure average accuracy, a system can work very well for most people and very badly for one particular group without anyone noticing.
How it shows up in a conversation
In what it says
In March 2024, UNESCO published a study of language models such as GPT-3.5, GPT-2 and Llama 2. It found they described women in domestic roles far more often than men, up to four times as often in one model, and associated female names with “home,” “family” and “children,” and male names with “business,” “executive” and “salary” (UNESCO, 2024).
The hardest biases to spot are the ones that don’t appear when you ask directly. A study published in Nature in 2024 showed that models that expressed positive views when asked directly about African Americans assigned less prestigious jobs to a person just because they wrote in African American English (Hofmann et al., 2024). In 2025, part of the same team found similar biases against speakers of German dialects in every model they evaluated (Bui et al., 2025). There’s no reason to think Spanish, with all its variety, is spared.
In how it understands
Bias also affects who a voice assistant understands well. As we explain in why doesn’t my voice assistant understand me, a Stanford study found that five commercial speech recognition systems made almost twice as many errors with African American speakers as with other American speakers (Koenecke et al., 2020). For people who aren’t understood, the assistant simply doesn’t work.
In how it sounds
Voices send messages too. In 2019, UNESCO devoted a report to voice assistants, which at the time almost all had a female voice by default. Its title, I’d blush if I could, was Siri’s answer to a sexist insult. The report warned that these voices reinforce the idea of women as obliging helpers and normalize speaking to them rudely (UNESCO, 2019).
Something similar happens with accents. If every synthetic voice speaks standard Spanish, the implicit message is that this is the correct Spanish. With LLYC we created Free the Voices, the first bank of diverse synthetic voices in Spanish, precisely to broaden what people hear.
The project sought gender diversity, not just accent diversity. The voice is one of the traits that triggers the most discrimination against LGBTQ+ people, and many end up changing or hiding the way they speak. To create the bank, voices from the community were collected in twelve countries, and from them we generated synthetic voices that don’t copy any specific person. We designed two new identities, Libertas and Vega, which together with Cástor, Fulu and Hatysa make up the first five diverse synthetic voices in Spanish.
What you can measure
A bias you don’t measure doesn’t get fixed. Some ways to do it:
- Break down results by group. Instead of average speech recognition accuracy, accuracy by accent, age or gender. Instead of average satisfaction, satisfaction for each user profile.
- Use tests designed to detect stereotypes. BBQ, for example, asks questions where the stereotyped answer and the correct one differ, and measures how often the model picks the stereotype (Parrish et al., 2022). Almost all of these tests are in English, so for Spanish you usually have to adapt them.
- Test with sentence pairs. The same request, changing only a name, a gender or a way of speaking. If the answer changes without justification, there’s a problem.
- Listen to the people using the assistant. Complaints from a particular group are often the first clue.
What you can do
Measuring is the first step. Then, depending on where the problem is: fill in the data with the missing voices and texts; review the assistant’s instructions and its default answers; choose the voice, name and personality deliberately; and repeat the tests every time the model changes, because each version has its own biases.
None of this makes a system bias-free, but it reduces the most harmful biases and lets you know what remains. It’s part of what we review when we evaluate assistants and agents and when we design custom synthetic voices.
References
- UNESCO (2024). Generative AI: UNESCO study reveals alarming evidence of regressive gender stereotypes.
- Hofmann, V., Kalluri, P. R., Jurafsky, D. and King, S. (2024). AI generates covertly racist decisions about people based on their dialect. Nature, 633, 147-154.
- Bui, M. D., Holtermann, C., Hofmann, V., Lauscher, A. and von der Wense, K. (2025). Large Language Models Discriminate Against Speakers of German Dialects. arXiv.
- Koenecke, A., Nam, A., Lake, E. et al. (2020). Racial disparities in automated speech recognition. PNAS, 117(14), 7684-7689.
- UNESCO and EQUALS Skills Coalition (2019). I’d blush if I could: closing gender divides in digital skills through education.
- Parrish, A., Chen, A., Nangia, N. et al. (2022). BBQ: A Hand-Built Bias Benchmark for Question Answering. Findings of ACL 2022.
- LLYC and Monoceros Labs. Free the Voices.