Want to know which tools to use to bring your written content to life? Here are a few keys.
What is speech synthesis technology?
The technology for turning text into audio, also called speech synthesis or text-to-speech (TTS), brings a text to life by generating audio with the characteristics of a human voice. If you use voice assistants like Google Assistant, Siri or Alexa, you’ve already heard a synthetic voice or digital voice answering all kinds of questions, from the weather to directions home.
This technology isn’t new, but recent advances are making it an increasingly powerful tool for listening to long texts effectively, such as a blog post or a news story. Besides improving accessibility, it lets anyone distribute audio content quickly and easily.
Where can I find synthetic voices in Spanish?
You’ll find Spanish synthetic voices on several online platforms with TTS APIs (application programming interfaces). On their websites you can try the available voices (male and female) by typing in some text and, in some cases, download the resulting audio.
The main online platforms with synthetic voice APIs in Spanish:
Google Cloud’s synthetic voices come in different quality levels, depending on the technology used. There are basic or standard voices, neural voices built with WaveNet and second-generation neural voices with more advanced technology.
Basic: Standard A, B, C and D.
WaveNet (neural, WaveNet technology): WaveNet B, C and D.
Neural2 (neural, latest technology): Neural2 A through F.
The few Spanish voices you can use from IBM Watson are all neural quality:
Spanish (Spain): Laura and Enrique.
Spanish (Americas, North / South): Sofía.
Synthetic voices in Spanish (2023)
There are also other synthetic voice websites or marketplaces that act as intermediaries or resellers of the platforms above, making it easier for non-developers to access these voices and download audio. Some even include audio editing features, and others let you use custom voices (voice cloning or AI-generated voices) or voice filters (STS, speech-to-speech). The market has moved forward in English, but it’s hard to find tools that work in Spanish. Here are a few for Spanish voices:
Fonos, for creating and editing audio content with synthetic voices, with your own voice (voice cloning) or with one from its catalog of custom AI-generated voices. Disclaimer: Fonos is a Monoceros Labs product and is currently in closed beta; you can request access on its website.
Reseller marketplaces with a catalog of API voices in Spanish: Murf.ai.
For voice filters, Voicemod is a very good option.
What kinds of synthetic voices are there?
The quality of these voices varies with the technology used to generate them, from standard voices with more monotonous prosody, which use parametric or concatenative technology, to more expressive voices that use neural technology (based on neural networks, like WaveNet since 2016). However, most API platforms don’t offer a wide variety of Spanish voices, which makes it hard to create custom, unique audio experiences. It’s easy to come across Microsoft Azure’s Álvaro voice in any recent app (we used it ourselves in the Cervecistas Alexa skill, for example).
On the other hand, with AI-based neural speech synthesis technology, like the one we build at Monoceros Labs, which also incorporates the latest advances in generative methods, we can get voices with custom characteristics in prosody and timbre, gaining expressiveness and diversity.
We can create two kinds of custom voices:
Cloned voices: We talk about cloning when the technology learns the phonetic characteristics of a specific voice from recordings of the original voice, imitating them in detail.
AI-generated blended voices: Here the technology learns from several voices and can use the characteristics learned from all of them to create a different synthetic voice that doesn’t identify any of the original voices. You could call it a voice that doesn’t exist. These voices are ideal for brands that want a 100% original voice.
What else should I keep in mind?
If you’re going to clone a voice, you need permission from the person the voice belongs to to create a model of their voice. This is very important; otherwise the voice could be used maliciously, as seen in many audio and video deepfakes.
At Monoceros Labs, we believe that building any product or service based on artificial intelligence must rest on firm principles of ethics and responsibility. This Manifesto sets out the principles we’re committed to.
Questions answered? Time to get your own synthetic voice
If you’d like your own synthetic voice, request access to our tool on the Fonos website, or get in touch with us directly through our contact page.
We can’t wait to hear what you’ll use it for!
Acknowledgments
This technology is part of the project “Speech synthesis in Spanish for building natural conversational systems,” funded by CDTI.
What would this face sound like? Synthetic voices from a single photo
At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.
Before you clone a voice, ask whose it is
A voice identifies a person and the law protects it. What Spanish and European rules say, which cases show it, and what a voice licensing agreement should include.
What exactly does a synthetic voice copy?
Timbre, intonation, pauses, pronunciation and accent: what a speech synthesis model learns from recordings, what it imitates well and where it still falls short.