
What would this face sound like? Synthetic voices from a single photo
At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.
Co-founder at Monoceros Labs
Carlos Muñoz-Romero is co-founder of Monoceros Labs and leads the technology behind Fonos, our voice and localization studio. He coordinates the development of our own Spanish speech synthesis models, which grew out of a NEOTEC project funded by Spain's CDTI (2020) in collaboration with Rey Juan Carlos University and RTVE, and which power Victoria, PRISA's voice for sports news, and RTVE's spoken election-night news in 2023.
He holds a degree in Computer Engineering from Carlos III University of Madrid, a master's in Data Science from the Open University of Catalonia (2026) and an MBA in media business from Carlos III. His master's thesis, on synthesizing speech from an image of a face, led to a paper accepted at IberSPEECH 2026. He co-authored the RTVE-UGR Chair paper on speech synthesis and news verification (Applied Sciences, 2024).
Before Monoceros he built the innovation department at BEEVA, now BBVA Technology, from scratch, and served as its global head of innovation and a member of the executive committee, with innovation lab and user experience teams in Spain and Mexico. He has taught generative AI on UNIR's master's program since 2024.

At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.

Watermarks and AI-generated text detectors: how they work, why they fail on human writing, and why the problem is further along in audio and images.

Multi-agent systems: when splitting a task across several AI agents pays off, when it creates more problems than it solves, and what changes on a phone call.

Cascaded, speech-to-speech or hybrid voice agents: how each architecture works, what you gain and lose in latency, control and voice, and when to choose each.

Noise, accent, unusual words or a cut-off turn: where the chain breaks between what you say and what a voice assistant understands, and what you can do about it.

A voice identifies a person and the law protects it. What Spanish and European rules say, which cases show it, and what a voice licensing agreement should include.

Timbre, intonation, pauses, pronunciation and accent: what a speech synthesis model learns from recordings, what it imitates well and where it still falls short.

Cloning a voice can take three seconds of audio or several hours. We explain zero-shot cloning, professional cloning and voice conversion, and when to use each.

After our third year, we look back on how we've evolved inside and out, along with our new brand.