Published on · updated on · Synthetic voice Responsible AI

Before you clone a voice, ask whose it is

A voice identifies a person and the law protects it. What Spanish and European rules say, which cases show it, and what a voice licensing agreement should include.

Whose voice is it?

Cloning a voice keeps getting easier. With a few seconds of audio, there are open models that can imitate someone recognizably, as we explained in how many minutes of audio it takes to clone a voice. That’s why, when someone asks us for a synthetic voice, the first question isn’t technical: whose voice is it, and what have they authorized?

This post isn’t legal advice. It sets out what we take into account in every project and the rules we rely on.

The voice is protected by law

In Spain, Organic Law 1/1982 on the right to honor, personal and family privacy and one’s own image considers it an unlawful intrusion to “use a person’s name, voice or image for advertising, commercial or similar purposes” (Article 7.6, our translation). The same law states that there’s no intrusion if the person gives express consent, and that this consent can be revoked at any time, with compensation for damages where applicable (BOE, Organic Law 1/1982, in Spanish).

What’s more, a voice recording of an identifiable person is personal data, and the General Data Protection Regulation applies to processing it. If the voice is technically processed to identify or verify that person, it becomes biometric data: the GDPR defines it as data resulting from specific technical processing of physical, physiological or behavioral characteristics that allow the unique identification of a person (GDPR, Article 4). Training a model on someone’s voice means processing their personal data, and it needs a clear legal basis.

The fact that audio is published online changes none of this. An interview on YouTube or a podcast doesn’t authorize anyone to build a model with that voice.

What happens when it’s forgotten

In February 2023, audiobook narrators discovered that the contract with Findaway Voices, Spotify’s audiobook distributor, included a clause allowing their recordings to be used to train machine learning models. Those files had reached Apple, which had just launched audiobooks narrated with synthetic voices. Many didn’t know they’d signed that clause. After protests and intervention by the SAG-AFTRA union, both companies stopped using the files for that purpose (AppleInsider, 2023; Writer Beware, 2023).

In January 2024, ahead of the New Hampshire primary, voters received robocalls with a voice imitating President Joe Biden telling them not to vote (NPR, 2024). The person responsible was eventually charged with voter suppression (New Hampshire Department of Justice). A few weeks later, the US Federal Communications Commission (FCC) ruled that AI-generated voices count as artificial voices under the law regulating robocalls, which makes them subject to the same prohibitions and penalties (FCC, 2024).

These are two different problems. In the first, voices were used under a consent nobody understood. In the second, with no consent at all and to deceive. Both show that protection has to be clear before recording.

What a voice licensing agreement should include

When we work with a voice actor to create a voice, we sign a specific document, separate from the usual recording contract. These are the points that shouldn’t be missing:

  • Purpose: What the voice will be used for: an assistant, news narration, a campaign. The more specific, the better.
  • Excluded uses: What will never be done with it. For example, political messages, advertising for other brands or content the person wouldn’t approve.
  • Duration and territory: For how long and where the synthetic voice can be used.
  • The model: Who keeps it, who can access it and what happens to it when the agreement ends: whether it’s deleted and how that’s proven.
  • Revocation: How the person can withdraw consent and what the consequences are for both parties, in line with what the law provides.
  • Compensation: Whether there’s a fixed fee, a per-use or a time-based payment, and whether the synthetic voice may replace voice-over work that used to be recorded live.
  • Transparency: How the audience will be told the voice is synthetic.

A generic consent along the lines of “I license my voice for any use” is exactly what went wrong in the audiobook case.

What if the voice belongs to no one?

Some synthetic voices don’t imitate any person: they’re created by blending traits learned from several voices, and the result is a new voice. Even so, they’ve been trained on recordings of real people, and those people also need to know what their voice is used for and to have authorized it. The fact that the result doesn’t sound like any of them doesn’t remove that obligation.

Telling people the voice is synthetic

The consent of the person lending their voice is one part. The other is not deceiving the listener. Article 50 of the EU AI Act, applicable since August 2, 2026, requires AI-generated audio to be marked so it can be detected, and requires disclosure when manipulated content imitating a real person is published (Regulation (EU) 2024/1689, Article 50).

Some providers already require it on their own. Microsoft, for example, asks for a recording in which the person reads a statement agreeing to the use of their voice before training a custom voice (Microsoft Learn).

In our custom synthetic voice projects, the agreement with the person lending their voice is signed before the first recording. It takes time at the start, and it prevents problems that can’t be fixed later.

References

Written by

Carlos Muñoz-Romero

Co-founder at Monoceros Labs

Co-founder of Monoceros Labs. A computer engineer with a master's in data science, he leads Fonos's voice technology. Previously Chief Innovation Officer at BEEVA (now BBVA Technology).

Share this article:

Related posts

Cover: What would this face sound like? Synthetic voices from a single photo. IberSpeech 2026.

What would this face sound like? Synthetic voices from a single photo

At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.

Was it written by AI?

Can you tell if a text was written by AI?

Watermarks and AI-generated text detectors: how they work, why they fail on human writing, and why the problem is further along in audio and images.

The prejudices in the data

An assistant repeats the prejudices in its data

Where the biases in AI assistants and agents come from, how they show up in a conversation or a voice, and what you can measure to reduce them.