Published on · Launches AI Synthetic voice PRISA

Victoria, the synthetic voice that reads sports news

A pioneering project to create an AI-generated voice for PRISA Radio to read sports news, the first of its kind launched in Spain.

Victoria la voz del fútbol (PRISA)

What is this project about?

“Victoria, la voz del fútbol” (“Victoria, the voice of soccer”) is the virtual contributor for Cadena SER (Carrusel Deportivo) and AS that you can talk to on Alexa. She gives you your team’s information and reads you its latest news from AS.com.

On our side, we created Victoria’s voice, an AI-generated synthetic voice that blends features of the voices it learns from without being identifiable as any of them. The project was carried out together with the PRISA Radio team and in collaboration with Amazon Alexa.

You can talk to Victoria on Alexa by saying “Alexa, abre Victoria” or from this link.

What does Victoria’s voice sound like?

Victoria’s voice is a custom voice with vocal characteristics that make it ideal for its use case: energy, speed and a medium-low pitch. It also has two different prosody styles, used for reading sports news and for the conversational experience on Alexa. The news-reading prosody has the intonation typical of sports news and a pace that makes it more pleasant to listen to than other, more neutral synthetic voices designed for general use. The prosody for conversation with users is different: slower, more expressive and more natural, so the conversation is more effective. The synthetic voice’s audio quality is tuned for use on Alexa.

How was this AI-generated synthetic voice created?

To create Victoria’s synthetic voice, with expressiveness and different prosody styles, we followed these steps in an iterative process until we got the result we wanted:

The process of creating an AI-generated synthetic brand voice

  1. Defining Victoria’s personality, as a conversational interface (an assistant available on Alexa) and as a voice, in a workshop with PRISA and the AS and Carrusel Deportivo teams.

  2. Translating the personality and voice traits into sound attributes that would help us choose the voice in terms of prosody.

  3. Studio recording of a selection of sentences designed specifically for the use case. Four female voices were used, with different prosody styles: conversational, news reading and sports news reading. Sentences with more expressiveness were also recorded so the voice would have enough energy.

  4. Data preparation: although it took more than 12 hours in the studio to record the more than 4,200 sentences, the resulting training dataset contained just over 4 hours of clean audio after phonetic review and preparation, a process that took weeks.

  5. Training several of our own multispeaker speech synthesis models, experimenting with the voices included in each model, the expressiveness and the prosody style. Our technology uses generative methods such as generative adversarial networks, or GANs, among others, to obtain unique voices. We ran many tests on prosody styles, comparing the news reading with Alexa’s voice. Training each model took weeks, including the final model trained from scratch, which needed more than six full days and over 350,000 training iterations.

  6. Selecting Victoria’s voice: blending characteristics and variables the different models had learned, we selected the voice and ran several evaluations of it. We measured the MOS (Mean Opinion Score), which rates voice quality based on listeners’ subjective ratings and perceptions. We also tested different voice configurations until we found the one that best matched the defined personality traits.

  7. Using the voice: from any text received, the system generates audio narrated in Victoria’s voice. Texts can be received programmatically through an API, for example from AS.com, or from Fonos. To pronounce foreign words, such as soccer players and stadiums, there’s a phonetic dictionary with more than 3,000 terms, growing every month.

How was it received?

It launched to the public in November 2022 and was the first AI-generated synthetic voice in Spanish created this way, available on Alexa and taking part in the radio show “Carrusel Deportivo.” It drew significant media attention, was sponsored by a leading car brand, and added up to about 100,000 user interactions on Alexa.

This project puts Cadena SER at the forefront of synthetic media and has opened the door to developing more AI products. Victoria has also become a tool for sharing sports content across the group’s different verticals, opening new ways to bring its content to new audiences.

The project won the “Best New Audio and Voice Product 2023” category at the International News Media Association’s (INMA) Global Media Awards 2023.

More information

Acknowledgments

NEOTEC - CDTI

Our speech synthesis technology is part of the project “Speech synthesis in Spanish for building natural conversational systems,” funded by CDTI (2021-2022).

Written by Monoceros Labs

Share this article:

Related posts

Cover: What would this face sound like? Synthetic voices from a single photo. IberSpeech 2026.

What would this face sound like? Synthetic voices from a single photo

At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.

Whose voice is it?

Before you clone a voice, ask whose it is

A voice identifies a person and the law protects it. What Spanish and European rules say, which cases show it, and what a voice licensing agreement should include.

What a synthetic voice copies

What exactly does a synthetic voice copy?

Timbre, intonation, pauses, pronunciation and accent: what a speech synthesis model learns from recordings, what it imitates well and where it still falls short.