Published on · Text-to-speech AI Synthetic voice TTS Technology

Technology for turning text into speech

Want to know which tools to use to bring your written content to life? Here are a few keys.

Tecnología de texto a voz (TTS)

What is speech synthesis technology?

The technology for turning text into audio, also called speech synthesis or text-to-speech (TTS), brings a text to life by generating audio with the characteristics of a human voice. If you use voice assistants like Google Assistant, Siri or Alexa, you’ve already heard a synthetic voice or digital voice answering all kinds of questions, from the weather to directions home.

This technology isn’t new, but recent advances are making it an increasingly powerful tool for listening to long texts effectively, such as a blog post or a news story. Besides improving accessibility, it lets anyone distribute audio content quickly and easily.

Where can I find synthetic voices in Spanish?

You’ll find Spanish synthetic voices on several online platforms with TTS APIs (application programming interfaces). On their websites you can try the available voices (male and female) by typing in some text and, in some cases, download the resulting audio.

The main online platforms with synthetic voice APIs in Spanish:

1. Amazon AWS Polly

To listen to AWS Polly’s voices you need an AWS account. You can access it from this link. You’ll find two quality levels: neural and standard.

1.1 Neural:

  • Spanish (Spain): Lucía and Sergio.
  • Spanish (Mexico): Mia and Andrés.
  • Spanish (United States): Pedro and Lupe.

1.2 Standard:

  • Spanish (Spain): Lucía, Conchita and Enrique.
  • Spanish (Mexico): Mia.
  • Spanish (United States): Penélope, Miguel and Lupe.

2. Microsoft Azure Text to Speech

With a Microsoft Azure account you can start using the available voices, all of them neural quality:

  • Spanish (Spain): Elvira and Álvaro (and more variations).
  • … and more than 20 voices in other Spanish varieties, such as Mexico, Argentina, Colombia or Chile.

3. Google Cloud TTS

Google Cloud’s synthetic voices come in different quality levels, depending on the technology used. There are basic or standard voices, neural voices built with WaveNet and second-generation neural voices with more advanced technology.

  • Basic: Standard A, B, C and D.
  • WaveNet (neural, WaveNet technology): WaveNet B, C and D.
  • Neural2 (neural, latest technology): Neural2 A through F.

4. IBM Watson

The few Spanish voices you can use from IBM Watson are all neural quality:

  • Spanish (Spain): Laura and Enrique.
  • Spanish (Americas, North / South): Sofía.

Synthetic voices in Spanish (2023)Synthetic voices in Spanish (2023)

There are also other synthetic voice websites or marketplaces that act as intermediaries or resellers of the platforms above, making it easier for non-developers to access these voices and download audio. Some even include audio editing features, and others let you use custom voices (voice cloning or AI-generated voices) or voice filters (STS, speech-to-speech). The market has moved forward in English, but it’s hard to find tools that work in Spanish. Here are a few for Spanish voices:

  • Fonos, for creating and editing audio content with synthetic voices, with your own voice (voice cloning) or with one from its catalog of custom AI-generated voices. Disclaimer: Fonos is a Monoceros Labs product and is currently in closed beta; you can request access on its website.
  • Reseller marketplaces with a catalog of API voices in Spanish: Murf.ai.
  • For voice filters, Voicemod is a very good option.

What kinds of synthetic voices are there?

The quality of these voices varies with the technology used to generate them, from standard voices with more monotonous prosody, which use parametric or concatenative technology, to more expressive voices that use neural technology (based on neural networks, like WaveNet since 2016). However, most API platforms don’t offer a wide variety of Spanish voices, which makes it hard to create custom, unique audio experiences. It’s easy to come across Microsoft Azure’s Álvaro voice in any recent app (we used it ourselves in the Cervecistas Alexa skill, for example).

On the other hand, with AI-based neural speech synthesis technology, like the one we build at Monoceros Labs, which also incorporates the latest advances in generative methods, we can get voices with custom characteristics in prosody and timbre, gaining expressiveness and diversity.

We can create two kinds of custom voices:

  • Cloned voices: We talk about cloning when the technology learns the phonetic characteristics of a specific voice from recordings of the original voice, imitating them in detail.
  • AI-generated blended voices: Here the technology learns from several voices and can use the characteristics learned from all of them to create a different synthetic voice that doesn’t identify any of the original voices. You could call it a voice that doesn’t exist. These voices are ideal for brands that want a 100% original voice.

What else should I keep in mind?

If you’re going to clone a voice, you need permission from the person the voice belongs to to create a model of their voice. This is very important; otherwise the voice could be used maliciously, as seen in many audio and video deepfakes.

At Monoceros Labs, we believe that building any product or service based on artificial intelligence must rest on firm principles of ethics and responsibility. This Manifesto sets out the principles we’re committed to.

Questions answered? Time to get your own synthetic voice

If you’d like your own synthetic voice, request access to our tool on the Fonos website, or get in touch with us directly through our contact page.

We can’t wait to hear what you’ll use it for!

Acknowledgments

NEOTEC - CDTI

This technology is part of the project “Speech synthesis in Spanish for building natural conversational systems,” funded by CDTI.

Written by

Nieves Ábalos

Co-founder at Monoceros Labs

Co-founder of Monoceros Labs. A computer engineer researching dialogue systems since 2009, now doing a PhD in AI at the University of Granada.

Share this article:

Related posts

Cover: What would this face sound like? Synthetic voices from a single photo. IberSpeech 2026.

What would this face sound like? Synthetic voices from a single photo

At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.

Whose voice is it?

Before you clone a voice, ask whose it is

A voice identifies a person and the law protects it. What Spanish and European rules say, which cases show it, and what a voice licensing agreement should include.

What a synthetic voice copies

What exactly does a synthetic voice copy?

Timbre, intonation, pauses, pronunciation and accent: what a speech synthesis model learns from recordings, what it imitates well and where it still falls short.