Today, a few seconds of recording are enough to clone a voice. The latest speech synthesis systems listen to a short sample and say whatever you ask in that voice. But what happens when there’s no recording at all?
Think of a historical figure from before microphones, the protagonist of a novel or the supporting characters in a video game. There’s an image, a portrait or a design, but no voice. Until now, the solution was to eyeball a stock voice that “fit.”
It also happens with some people who can’t speak and communicate through a device that speaks for them, known as augmentative and alternative communication. People who lose their voice to an illness sometimes manage to record it beforehand and can get it back in synthetic form. People who have never spoken have no recordings, so they use a stock voice, the same one many others use. A voice generated from their face could be another option: synthetic, but plausible and different for each person.
In a paper we presented at IberSpeech 2026, together with José A. González López (University of Granada), we explore this path: generating a plausible voice from a single photo of a face. The preprint is on arXiv and the code is on GitHub.
Does a voice really resemble a face?
A little, and it’s best not to overstate it.
There’s a biological basis: the same hormones that make the jaw grow during development also thicken the vocal folds. That’s why a face gives clues about a voice, such as approximate age, sex or build. But many other things affect how someone sounds that you can’t see in a photo: breathing, the shape of the throat, accent, speaking habits.
So the link between face and voice exists, but it’s weak and doesn’t determine a specific voice. Two people who look very alike can sound very different.
This creates a well-known problem in the field. If you train a system to get each face’s exact voice right, it learns that the safest bet is not to take risks: it ends up generating an average male voice and an average female voice, both equally generic. This is called collapse, and it was the first thing we wanted to avoid.
How we did it
The central idea is not to touch the part that already works well.
We start from an open speech synthesis model that already sounds very natural. That model stores a summary of how each voice sounds, its “style.” Normally that style is extracted from a recording. We extract it from a face.
To do that, we leave the voice model as it was and add a small component that translates what a face recognizer sees into the language of voices. That’s why we call the method Freeze-Align: freeze the voice and align the face with it.
The key is what we ask that component to learn. We don’t require it to get each face’s exact voice right, because that leads to generic voices. We ask it to place each person’s face closer to their own voice than to the others, to keep the voices varied and to respect basic cues such as apparent age and sex.
Timbre from the face, intonation from the model
A photo says nothing about how someone intones, or about their rhythm or pauses. If everything came from the face, the voice would sound monotonous.
So we separate the two. The face provides the timbre, and the voice model itself provides the intonation. There’s also a control for how much weight the face carries. With a lot of weight, the voice resembles what the photo suggests more closely, but loses naturalness. With little, it sounds very natural but more generic. Each use can choose its own balance.
What we found
We trained on videos of talks in English and tested the system on faces of people it had never seen.
The voices it generates sound natural and varied: they don’t collapse into a couple of generic voices. And they resemble each person’s real voice more than you’d expect by chance, although the resemblance is still modest.
We also tested it in Spanish. We adapted the voice model to our language and used the same component that translates faces, without retraining it. It worked, which suggests the face-voice relationship it learns doesn’t depend much on the language.
What we don’t know yet
The face contributes a small part of the voice. Most of it still comes from what the model already knows about how people sound. We’ve tested it with a small number of people and with automatic quality metrics; a listening test with people is still missing, and that’s what really tells you whether a voice is convincing.
We also haven’t tested it with people who use augmentative communication, and since we trained on adults, children’s voices are out of scope for now.
This is research, not a product.
Next steps
We see three paths toward voices that resemble the person more and more.
The first is to start from voice models that know more voices. Ours learned from a limited number of speakers, and can only propose voices similar to those it has heard. A model trained on many more voices, of more ages, accents and languages, would have far more to choose from.
The second is to learn from many more faces. The more face-voice pairs the system sees, and the more diverse they are, the better it will understand which facial features say something about the voice and which don’t.
The third is to refine how face and voice are linked. Today that connection is still approximate. With more precise alignment methods, the voice that comes out of a photo should get closer to that person’s voice.
None of these steps is immediate, but together they’d bring us closer to more faithful voices for the cases we mentioned at the start: characters who never had a voice, and people who need one.
A technology to use with care
One thing should be clear: this isn’t voice cloning. The system doesn’t listen to any recording, it only sees a face, so it can’t reproduce anyone’s real voice. What it generates is a voice that sounds plausible for that face, not that person’s voice. Even so, it carries risks. Someone could present that voice as a real person’s authentic voice. And since it learns from real data, it can reinforce stereotypes about how someone “should” sound based on how they look.
We believe responsible use requires people’s consent, making clear where each voice comes from and evaluating its biases before bringing it into any product. Fictional characters, video games, historical outreach or giving one more option to people who can’t speak are good areas to explore it. Using it with real people without their permission is not.
If you’re interested in the technical details, see the paper on arXiv and the code on GitHub. And if you’re thinking about voices for characters, get in touch.