“We want a voice of our own that sounds like us, not like everybody else.”

Services

Custom synthetic voices

We design and train synthetic voices with identity: the voice of a brand, a media outlet or a character. We take care of pronunciation, reading styles and the consent of the people who lend their voice.

Spanish text-to-speech: how we build a custom brand voice

Speech synthesis, or text-to-speech, turns any text into audio in a specific voice. A custom brand voice is one only you use, trained for your content in Spanish and its accents. For Victoria, PRISA's voice for sports news, we recorded more than 12 hours in the studio and over 4,200 sentences, which became more than 4 hours of clean audio to train the model.

Voice personality

Together with your team, we define how it should sound: personality, tone and reading styles.

Casting or designed voice

We choose the voice actor, with their consent, or design a voice that doesn't identify any real person. You approve it before we go on.

Recording and data

We write the recording scripts, direct the sessions and prepare the data with phonetic review. For Victoria, that meant over 12 hours in the studio and 4,200 sentences.

Training and tuning

We train the model and tune intonation and pacing for each context: news, conversation or narration.

Pronunciation

Dictionaries so names, places, acronyms and industry terms are always read correctly.

Voice delivery

Use it from your own systems through an API, or from Fonos, our voice and localization studio.

Before building a voice, hear the ones that exist

Spanish voices, in all their accents, in Fonos, our voice and localization studio.

Fonos

How we work

01

Explore

We understand the challenge and review what today's technology can do for your case, your language and your data.

02

Prototype

We build just enough to validate the idea with real people.

03

Pilot

We try it out in your organization, measure and adjust.

04

Hand over

We document and train your team so they can keep evolving it.

What you get

  • Exclusive synthetic voice
  • Reading styles for each context
  • Pronunciation dictionary
  • Integration into your systems or Fonos

Who it's for

  • Media outlets and radio stations
  • Brands that want a voice of their own
  • Organizations that need accessible audio information

Related projects

All projects
Victoria, la voz del fútbol (PRISA)
Project Awarded 2023

Victoria: the first synthetic voice for sports news in Spanish

A pioneering project to create an AI-generated voice for PRISA Radio to read sports news, the first of its kind launched in Spain.

RTVE Elecciones with artificial intelligence
Project Awarded 2023

Accessible election information thanks to AI: RTVE Elecciones

Synthetic voices to make real-time election coverage accessible for towns with fewer than 1,000 inhabitants, an AI project by RTVE.

RTVE weather with artificial intelligence
Project

Catalan synthetic voices for local weather information

Making local weather information accessible in Spanish and Catalan with AI, an innovative project by RTVE.

Free the Voices, by LLYC and Monoceros Labs
Project Awarded 2025

Free the Voices: the first bank of diverse synthetic voices

An initiative that uses AI to create synthetic voices representing the diversity of the LGBTQ+ community, promoting inclusion and reducing bias in voice technology.

Synthetic voice and Spanish text-to-speech FAQ

Does the voice have to belong to a real person?

No. Victoria, PRISA's voice for sports news, was designed not to identify any real person. We can also start from a voice actor's voice, always with their consent.

Can you create voices in other languages or accents?

We've built Catalan voices for RTVE and Spanish accents such as Andalusian in Fonos. For each language or accent we assess the available data and the quality we can reach.

How is the voice used afterwards?

Integrated into your systems or through Fonos, our voice and localization studio.

How much audio does a custom voice need?

It depends on the quality and the reading styles you need. For Victoria, PRISA's voice for sports news, we recorded more than 12 hours in the studio and over 4,200 sentences written for that purpose, which became just over 4 hours of clean audio after phonetic review. Other uses need far less: from a few seconds of audio for instant (zero-shot) cloning to a minimum of 15 to 30 minutes for professional cloning.

What drives the timeline and the price?

Three things: whether we start from a real person or a designed voice, how many reading styles you need (news, conversation, narration), and how many languages or accents.

What can't a synthetic voice do well yet?

Proper names, acronyms and industry terms come out wrong unless it's taught them, which is why we deliver a pronunciation dictionary. A voice trained to read the news won't narrate a story or hold a conversation equally well: each style has to be prepared. Depending on the technology, it can also hallucinate: skip or repeat words, or say something that isn't in the text. And a voice based on a real person is only created with their consent and clear terms of use.

Services

Got a voice or conversation challenge?

Tell us what you want to achieve. If we're not the right team, we'll say so.