Project

Synthetic voices for audio verification research

Working with the RTVE-UGR Chair to develop synthetic voices that support research into detecting fake audio, helping fight disinformation.

← All projects
RTVE-UGR Chair collaboration on audio verification

Working with the RTVE-UGR Chair, a research chair run by RTVE, Spain’s public broadcaster, and the University of Granada, Monoceros Labs took part in a multidisciplinary project focused on developing tools to detect fake audio.

Our contribution focused on creating high-quality cloned synthetic voices to train and evaluate audio deepfake detection models, that is, fake audio that impersonates a person.

The project developed an audio verification tool available to VerificaRTVE and other agencies through the IVERES project.

Interface of the deepfake detection tool that was built.

The challenge

The spread of fake audio, or deepfakes, is a growing threat to the integrity of information and of our society. The project aims to:

  • Develop effective tools to detect fake audio.
  • Create a corpus of synthetic voices to train models that detect impersonations of public figures.

More information in the scientific paper published in 2024:

Deep Speech Synthesis and Its Implications for News Verification: Lessons Learned in the RTVE-UGR Chair

Our contribution

Developing synthetic voices

Our work focused on creating cloned synthetic voices of public figures: King Felipe VI, Prime Minister Pedro Sánchez and Deputy Prime Minister Yolanda Díaz. We used two different techniques:

  1. Speech synthesis (TTS):

    • Models that imitate a person’s full prosody.
    • A focus on naturalness and expressiveness.
    • Adaptation to different speaking styles.
  2. Voice conversion (STS):

    • Models that specifically imitate the person’s timbre.
    • Preserving identifying vocal characteristics.
    • Real voices as a base, with a focus on expressiveness.

Methodology

Developing the voice clones followed a rigorous approach:

  1. Selecting voice data:

    • Speeches and parliamentary appearances.
    • Prioritizing high-quality audio.
    • Selecting samples without background noise.
  2. Audio processing:

    • Cleaning and normalizing samples.
    • Accurate phonetic transcription.
    • Thorough quality control.
  3. Controlled, private and secure training:

    • Our own on-premises infrastructure.
    • A fully private and secure process.
    • No access to public clouds.

Once the voice models were built, we generated fake audio and, together with the real audio used for training, created a dataset to train a classifier that detects fake audio of the target people.

A classifier to detect deepfakes

As part of the project, a deepfake detection classifier was developed specifically for the cloned voices. It was trained on thousands of audio clips per target voice, both real and fake, that is, generated with the cloned voices.

The classifier uses an architecture based on FastAudio, which can adapt dynamically to the characteristics of spoofing threats. Unlike traditional systems, our classifier uses filter layers that are tuned during training, allowing more accurate detection of audio manipulation.

What’s new about it

The project stands out for several reasons:

  1. A dual approach: creating both high- and low-quality voices to train the classifier.
  2. Security: a fully controlled process on our own infrastructure.
  3. Ethics: responsible, supervised use of the technology with a goal that benefits society.
  4. Collaboration: joint work between academia, media and industry:
  • RTVE: contributing journalistic knowledge and real use cases.
  • University of Granada: leading the research and model development.
  • Monoceros Labs: providing expertise in speech synthesis.

Impact and results

Our contribution was key to:

  • Creating a corpus of synthetic voices for research.
  • Developing more robust detection tools.
  • Advancing the fight against disinformation.
  • Establishing good practices in using AI for verification.

Recognition

This collaboration, part of the IVERES project mentioned above, received the following recognition:

More information

This project is an example of how speech synthesis technology can be applied responsibly to fight disinformation. Collaboration between different players allowed us to build more effective tools for audio verification, contributing to the integrity of information online.

Related services

Shall we build the next one?

Tell us about your voice or conversation challenge.