Published on · updated on · Synthetic voice

What exactly does a synthetic voice copy?

Timbre, intonation, pauses, pronunciation and accent: what a speech synthesis model learns from recordings, what it imitates well and where it still falls short.

What a synthetic voice copies

When we say a synthetic voice “sounds like” someone, we’re mixing several things. You recognize a voice by its timbre, but also by how it rises when asking a question, where it breathes, how it pronounces its s’s or how fast it talks. A speech synthesis model learns almost all of that from the recordings it’s trained on, though not everything with equal success. Let’s go trait by trait through what it copies well and where the seams show.

Timbre: what you recognize first

Timbre is the quality that lets you tell two people apart even when they say the same sentence at the same pitch. It depends on each person’s body: the vocal folds, the throat, the mouth, the nose.

An important part of it is the fundamental frequency, or F0: the number of times per second the vocal folds vibrate. In speech it usually ranges between 80 and 450 hertz, lower in men than in women and children (Aalto University). What we perceive as a higher or lower voice is mostly that frequency. On top of it, the shape of the mouth and throat boosts some frequencies and dampens others, and that gives each voice its own color.

Timbre is what current models copy best. It’s stable throughout a recording and present in every second of audio, so there’s plenty to learn from. That’s why even a very small sample produces a voice reminiscent of the original, as we explained in our post on how many minutes it takes to clone a voice.

Prosody: how it’s said

Prosody is everything that sits above individual sounds: intonation, word stress, rhythm, pauses and speed. The Instituto Cervantes defines it as the set of phenomena that span more than one phoneme, and notes that in Spanish the two most relevant are stress and intonation (Centro Virtual Cervantes). Prosody also conveys emotion, origin and attitude.

With prosody, the same sentence changes meaning. “Ya has llegado” (“you’re here already”) can be a statement, a question or a reproach. In writing, punctuation and context signal it. A synthetic voice has to work it out.

Models learn prosody from data. If the recordings include questions, exclamations and lists, the model learns their intonation curves. If they’re almost all read-out statements, its questions will come out flat. There are three situations where this shows most:

  • Long sentences, where the model doesn’t always know where to breathe or how to group the words.
  • Emphasis: which word matters in a sentence depends on meaning, and that can’t always be inferred from the text.
  • Questions that don’t start with a question word, like “¿Vienes mañana?” (“You’re coming tomorrow?”), where only intonation signals that it’s a question.

Research has spent years trying to control prosody separately from timbre. In 2018, two Google papers showed that a model could learn prosody from a reference recording and apply it to another text and another voice (Skerry-Ryan et al., 2018), and that it could discover speaking styles without anyone labeling them (Wang et al., 2018).

Pronunciation and accent

To speak, the model has to turn letters into sounds. In Spanish this is easier than in English, because spelling closely matches pronunciation, but there are pitfalls: foreign words, proper names, acronyms and brands.

When a name isn’t pronounced the way it’s written, the model gets it wrong with confidence. For Victoria, the voice we created with PRISA for sports news, the phonetic dictionary grew past 3,000 terms, mostly names of players and stadiums, and had to be expanded every month. It’s manual work, and it’s still necessary.

Regional accent is also learned from the recordings: seseo, the aspirated s, Canarian or River Plate intonation. A voice cloned from someone from Seville will speak with a Seville accent if the data has it. The problem appears when the base model has mostly heard standard Spanish: it then tends to pull the voice toward that accent, and the mix sounds odd.

What the text doesn’t say

Before synthesizing, you have to decide how to read what’s written. In Spanish, “1/2” can be “un medio” (one half) or “uno de febrero” (February 1). “Km” is read “kilómetros.” “Felipe II” is read “Felipe segundo” (Philip the Second), but “tomo II” is “tomo dos” (volume two). And “3-1” in a soccer report is “tres a uno” (three to one). In English the problem is even bigger: the same number, like 2026, is read differently if it’s a year or a quantity. This step is called text normalization, and errors here change the message completely. Richard Sproat and Navdeep Jaitly, at Google, showed this in 2016: neural networks with very good overall accuracy still made unacceptable errors, like misreading a figure, and had to be combined with rules (Sproat and Jaitly, 2016). Today’s models do better, but numbers, dates and acronyms should always be checked.

Style and emotion

The same person doesn’t speak the same way reading the news as telling a story or answering a phone call. If a synthetic voice has to switch from one style to another, it needs examples of each.

Many systems today let you ask for a style in words. In March 2025, OpenAI introduced a speech synthesis model you can instruct on how to speak, for example “like a sympathetic customer service agent” (OpenAI, 2025). It works well for general changes in tone. Believable emotion sustained over several minutes, or irony that comes across, is still hard.

What gets copied by accident

The model doesn’t distinguish between the voice and what surrounds it. If the recordings have echo, background noise, mouth clicks or bumps on the table, it learns them and reproduces them where you least expect. The same goes for breaths: used well, they add naturalness, but if they’re badly trimmed they turn up in the middle of a word. That’s why we review the audio by hand before training.

What this means if you want a synthetic voice

If timbre is what matters most, for example so a brand voice is recognizable, you have it relatively easy. If what matters is how it speaks, how it asks questions, how it emphasizes or what accent it has, the work lies in designing the recordings well: sentences that cover the sounds of the language, the types of questions and the styles that will be used. And in any case you need a pronunciation dictionary and someone who listens to the result carefully.

It’s the least glamorous part of creating a voice and the one that shows most in the end. We do it in every custom synthetic voice project.

References

Written by

Carlos Muñoz-Romero

Co-founder at Monoceros Labs

Co-founder of Monoceros Labs. A computer engineer with a master's in data science, he leads Fonos's voice technology. Previously Chief Innovation Officer at BEEVA (now BBVA Technology).

Share this article:

Related posts

Cover: What would this face sound like? Synthetic voices from a single photo. IberSpeech 2026.

What would this face sound like? Synthetic voices from a single photo

At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.

Whose voice is it?

Before you clone a voice, ask whose it is

A voice identifies a person and the law protects it. What Spanish and European rules say, which cases show it, and what a voice licensing agreement should include.

3 s or hours of audio to clone a voice

How many minutes of audio does it take to clone a voice?

Cloning a voice can take three seconds of audio or several hours. We explain zero-shot cloning, professional cloning and voice conversion, and when to use each.