Carlos Muñoz-Romero

Co-founder at Monoceros Labs

Carlos Muñoz-Romero is co-founder of Monoceros Labs and leads the technology behind Fonos, our voice and localization studio. He coordinates the development of our own Spanish speech synthesis models, which grew out of a NEOTEC project funded by Spain's CDTI (2020) in collaboration with Rey Juan Carlos University and RTVE, and which power Victoria, PRISA's voice for sports news, and RTVE's spoken election-night news in 2023.

He holds a degree in Computer Engineering from Carlos III University of Madrid, a master's in Data Science from the Open University of Catalonia (2026) and an MBA in media business from Carlos III. His master's thesis, on synthesizing speech from an image of a face, led to a paper accepted at IberSPEECH 2026. He co-authored the RTVE-UGR Chair paper on speech synthesis and news verification (Applied Sciences, 2024).

Before Monoceros he built the innovation department at BEEVA, now BBVA Technology, from scratch, and served as its global head of innovation and a member of the executive committee, with innovation lab and user experience teams in Spain and Mexico. He has taught generative AI on UNIR's master's program since 2024.

Posts by Carlos Muñoz-Romero

Cover: What would this face sound like? Synthetic voices from a single photo. IberSpeech 2026.

What would this face sound like? Synthetic voices from a single photo

At IberSpeech 2026 we presented a method that generates a plausible voice from a photo. Here's how it works, what we measured and how far it goes.

Was it written by AI?

Can you tell if a text was written by AI?

Watermarks and AI-generated text detectors: how they work, why they fail on human writing, and why the problem is further along in audio and images.

Several agents for one call

How many agents does it take to answer a phone call?

Multi-agent systems: when splitting a task across several AI agents pays off, when it creates more problems than it solves, and what changes on a phone call.

3 ways to build a voice agent

Three ways to build a voice agent, and none is perfect

Cascaded, speech-to-speech or hybrid voice agents: how each architecture works, what you gain and lose in latency, control and voice, and when to choose each.

Noise, accents and unusual words

Why doesn't my voice assistant understand me?

Noise, accent, unusual words or a cut-off turn: where the chain breaks between what you say and what a voice assistant understands, and what you can do about it.

Whose voice is it?

Before you clone a voice, ask whose it is

A voice identifies a person and the law protects it. What Spanish and European rules say, which cases show it, and what a voice licensing agreement should include.

What a synthetic voice copies

What exactly does a synthetic voice copy?

Timbre, intonation, pauses, pronunciation and accent: what a speech synthesis model learns from recordings, what it imitates well and where it still falls short.

3 s or hours of audio to clone a voice

How many minutes of audio does it take to clone a voice?

Cloning a voice can take three seconds of audio or several hours. We explain zero-shot cloning, professional cloning and voice conversion, and when to use each.

En evolución constante.

Always evolving

After our third year, we look back on how we've evolved inside and out, along with our new brand.