
What would this face sound like? Synthetic voices from a single photo
At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.
Published on · Launches AI Synthetic voice RTVE
Real-time election coverage with artificial intelligence for towns with fewer than 1,000 inhabitants.


RTVE’s Technology Strategy department (RTVE is Spain’s public broadcaster) has launched and is leading this proof of concept of new AI-based technologies to generate content, automatically and in real time, for the municipal elections of May 28, 2023, covering the nearly 5,000 towns in Spain with fewer than 1,000 inhabitants. To ensure the information is accurate and reliable, the news stories are generated from election results obtained from official sources, such as the Ministry of the Interior.
The information is available at rtveia.es and on the Alexa virtual assistant by saying “Alexa, abre elecciones inteligencia artificial” or through this link if you have an Alexa account.
At Monoceros Labs, we contributed our experience in conversational AI and speech synthesis so the written news can be listened to, thanks to two synthetic voices with a news-reading style. We also made the most important election information accessible in real time through conversation with the voice assistant in the Alexa skill “RTVE Elecciones Inteligencia Artificial”, where you can listen to the summary story, ask about a town and find out how the count, turnout or a given party is doing there, among other things. Besides all the design and development of the voice app, the synthetic voice used in the conversation with Alexa was also created by our company.
Try the skill live by saying “Alexa, abre elecciones inteligencia artificial.”
The proof of concept involves the University of Castilla-La Mancha, the University of Granada through the RTVE-UGR Chair, ONCE and AWS (Amazon Web Services). The technology company Narrativa is in charge of generating the news stories and the written and visual information available on the website.
Artificial intelligence is involved in reading data, writing text, structuring the information, generating images, conversing with the voice assistant, and speech synthesis to create and use synthetic voices to narrate the texts. The entire technology-driven process, as well as decisions on content and quality, is constantly overseen by professionals from different areas of RTVE, including engineers, technicians and journalists, and by the institutions and companies involved in this proof of concept.
Besides images of the towns from RTVE’s archives or recordings from its regional centers, images created by some of the best-known generative AI tools will also be published, also as a test. Another custom detail is the specific theme music that accompanies the narration, composed by RTVE’s musicians and sound designers.
Narrativa was in charge of generating the news stories. From the data supplied by official sources, its artificial intelligence technology, “Gabriele,” can interpret it and turn it into a news text.
The stories include a headline with the winning political party (by number of votes) and the town, varying its structure so the headlines aren’t all too uniform in style. A subheading covers the highlights of parties and votes or the result in the town. In the body of the story, once the count is complete, it explains which party is considered the winner (for getting the most votes in the town), how the other parties did, how many votes each one got (in number and percentage) and turnout on the day compared with the previous election, among other data or the usual graphics such as images, maps, parliament charts or coalition calculators.
This process, created automatically with natural language text, has always been trained and overseen by professionals from different areas of RTVE to certify and verify its news, editorial and visual quality. So that nobody has any doubts about authorship, and in the interest of transparency, each publication states that it was generated automatically by AI from official data.
At Monoceros Labs, we took care of how the information is consumed by voice. We worked on three synthetic voices: two for reading the news, one male and one female, and a third female voice with a conversational style for the dialog in the Alexa skill. All three are unique voices, custom in style and timbre, because, thanks to generative AI, they’re a blend of people’s voices and their characteristics, something unusual that’s barely being done in Spanish, where the usual and most available option is synthetic voices cloned from people (both links in Spanish).
The attention to detail went into every area, for example making sure both towns and political parties are pronounced as correctly as possible. These voices “read” the texts and make the news content more accessible to people with disabilities, to people without strong digital skills, and to anyone listening on smart devices.
Beyond the synthetic voice, the voice app, or Alexa skill, uses conversational AI technologies to recognize speech, understand people’s requests and the context of the conversation, and respond by text and voice, this time with the custom synthetic voice. We handled the conversation design, writing the responses, the implementation and integration with Narrativa’s system, building the language understanding model, and running various tests with users. The experience is multimodal: you don’t just hear the information, you can also see it on a screen.
To explore this information, anyone with Alexa or an Echo device just has to say:
Alexa, abre elecciones inteligencia artificial.
It’s also available by clicking this link.
The University of Castilla-La Mancha takes part in this test, contributing its academic knowledge and experience in journalism and the application of new technologies.
For the conversational experience on Alexa, we also worked with Zoraida Callejas’s team at the University of Granada, through the RTVE-UGR Chair, and as always, we had the support of the Amazon Alexa team in Spain.
The project also builds in important accessibility features, both in the voices and in the structure of the website. For this, it had the participation and advice of ONCE, the Spanish National Organization of the Blind, whose specialists carefully reviewed the content and how it’s presented so it’s accessible.
It’s a combination of collaborations in the service of society, aiming to help maintain, or even improve, the quality of democracy by making all this information accessible as accurately and reliably as possible.
The system, designed to create content for rural or less populated Spain that meets accessibility requirements, provides a service that isn’t possible with traditional means, and it could be used to cover regional, national or European elections as long as the necessary official data is available. For any of them, this resource makes it possible to complement traditional coverage with information for analysis, such as the distribution of votes, whether political trends shift, support for national or regional parties or local initiatives, among many other possibilities currently being explored.
This test, focused on the least populated towns in Spain, could be extended to provide full coverage, since the system can generate results for all of them, from the smallest to big cities like Madrid, Barcelona, Valencia or Zaragoza.
This innovative project aims to position RTVE as a pioneering public broadcaster serving citizens, to support strategies and policies for a humanist, inclusive digital transition, and to back Spanish business and technology initiatives that take the Spanish language as their reference. It also seeks to strengthen its public service commitment by closing territorial gaps, offering content that didn’t exist until now and guaranteeing all citizens the same right of access to information, whether it’s about the main cities or the smallest towns, in this case on a subject as sensitive as voting and the election results that shape institutions.
The project also won the Best AI implementation project award in the third edition of the TM Broadcast awards.

Our speech synthesis technology is part of the project “Speech synthesis in Spanish for building natural conversational systems,” funded by CDTI.
Written by Monoceros Labs
Share this article:


At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.


A voice identifies a person and the law protects it. What Spanish and European rules say, which cases show it, and what a voice licensing agreement should include.


Timbre, intonation, pauses, pronunciation and accent: what a speech synthesis model learns from recordings, what it imitates well and where it still falls short.