
What would this face sound like? Synthetic voices from a single photo
At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.
Published on · updated on · Synthetic voice
Cloning a voice can take three seconds of audio or several hours. We explain zero-shot cloning, professional cloning and voice conversion, and when to use each.


It’s the question we’re asked most when someone wants their own synthetic voice. The short answer is that it depends on what you mean by cloning. Some systems imitate a voice from three seconds of audio, and others need several hours of studio recording. Both say they clone, but they don’t give the same result or serve the same purpose.
Cloning a voice means creating a speech synthesis model that speaks in a specific person’s voice, from recordings of them. The model receives new text that person has never read and generates the audio as if they were saying it.
What gets copied can be just the timbre, the color of the voice that lets you recognize someone on the phone, or also the way they speak: rhythm, pauses, intonation, accent. How much audio you need depends mostly on how much of this you want it to match.
There are two main paths, already described in 2018 by a Baidu team in one of the first papers on the subject (Arik et al., 2018):
In that paper, adaptation gave more naturalness and similarity, while encoding was much faster and cheaper to deploy. Models have changed a lot since then, but that difference still explains much of what you can expect from each option.
Zero-shot cloning is the second path: the system listens to a short sample of a voice it has never heard and imitates it on the spot, without training.
The trick is that the model has already learned, from thousands of hours of other people, how voices vary. Google showed this in 2018 by taking a model trained to verify who’s speaking and using it to condition the synthesis (Jia et al., 2018). In January 2023, Microsoft introduced VALL-E, which imitates a voice from a 3-second recording. To get there, they trained it on 60,000 hours of English speech (Wang et al., 2023). Today this is within anyone’s reach with open models: XTTS-v2 clones from a 6-second clip and speaks 17 languages, Spanish among them (Coqui), and F5-TTS, trained on 100,000 hours of multilingual speech, released its code and weights (Chen et al., 2024). Commercial services ask for a bit more: ElevenLabs recommends 1 to 2 minutes of audio for its instant cloning (ElevenLabs).
A few seconds of your voice are enough because other voices already did the heavy lifting. That has consequences:
Professional cloning is the first path: you start from a model and train it on recordings of the person. Now we’re talking minutes or hours.
You can start with very little. The open YourTTS model showed it’s possible to fine-tune a model with less than a minute of speech and get reasonable similarity (Casanova et al., 2022). In our experience, from 10 minutes of clean audio you get a recognizable voice, and with more than 20 it clearly improves. Besides the audio you need the transcripts, reviewed so that every sound matches what’s said. If the text doesn’t match the audio, the model learns the pronunciation wrong.
Commercial services are along the same lines. For its professional cloning, ElevenLabs asks for at least 30 minutes of audio and recommends getting close to 2 or 3 hours for the best result (ElevenLabs).
When the voice is going to represent a brand, it’s recorded in a studio with a script designed to cover all the sounds of the language and the styles that will be used. For Victoria, the synthetic voice we created with PRISA, it took more than 12 hours in the studio and more than 4,200 sentences. After phonetic review, just over 4 hours of clean audio remained for training. For reference, one of the classic datasets for training a single-speaker voice from scratch, LJ Speech, has about 24 hours of audio (Ito and Johnson, 2017).
More audio mostly improves the way of speaking. The cloned voice asks questions, lists things and pauses like the original person.
There’s another technique that’s also called cloning and works differently: voice conversion, or speech-to-speech. Here you don’t start from text. Someone speaks, and the system transforms their audio so it sounds like another voice, like a filter.
RVC, an open voice conversion tool with tens of thousands of stars on GitHub, recommends about 10 minutes of clean audio of the target voice (RVC-Project). The result keeps the cloned voice’s timbre, but the intonation, rhythm and pauses belong to the person speaking. That’s why it’s useful for dubbing a performance, and less so for generating audio from text without a person behind it.
| Method | Audio of the person | What it imitates | What it’s for |
|---|---|---|---|
| Zero-shot | A few seconds to 2 minutes | Timbre | Prototypes and quick tests |
| Voice conversion | About 10 minutes | Timbre, over someone else’s performance | Dubbing, changing the voice in a recording |
| Professional cloning | 10 minutes to several hours | Timbre and way of speaking | Brand voices, narration, assistants |
Audio quality matters as much as duration. An hour recorded in a room with echo, background noise or different microphones can give a worse result than twenty well-recorded minutes. The model learns the echo and the bumps on the table too.
What gets recorded matters as well. If every sentence is a statement read in a neutral tone, the cloned voice won’t know how to ask a question naturally or sound enthusiastic. That’s why we design scripts with the final use in mind: reading sports news is not the same as answering as an assistant.
The fewer seconds it takes, the easier it is to clone a voice without the person knowing. In 2023, Meta decided not to release its Voicebox model, which imitated a voice from a two-second sample, because of the risk of misuse (Meta AI, 2023). Serious providers ask for proof of consent: Microsoft, for example, requires a recording of the person reading a statement authorizing the use of their voice (Microsoft Learn). And Article 50 of the EU AI Act, applicable since August 2, 2026, requires synthetic audio to be marked as AI-generated and requires disclosure when a deepfake is published (Regulation (EU) 2024/1689).
We only clone a voice with the person’s explicit permission. We also explain this in our post on text-to-speech.
If you’re thinking about a voice of your own, start by deciding where and what it will be used for. That determines the method and how much needs to be recorded. It’s part of what we do in custom synthetic voices.
Share this article:


At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.


A voice identifies a person and the law protects it. What Spanish and European rules say, which cases show it, and what a voice licensing agreement should include.


Timbre, intonation, pauses, pronunciation and accent: what a speech synthesis model learns from recordings, what it imitates well and where it still falls short.