An emergency medical service in Spain attends people who don’t speak Spanish every day. In an emergency, medical staff need to know what’s wrong with the patient, what they’ve taken or whether they have allergies, and there’s no time to find an interpreter. Together with a major telecom operator, which would integrate the solution, we prepared a pilot, starting with a proof of concept, for a face-to-face translation app on the ambulances’ devices.
The problem
The conversation happens on the street, in someone’s home or inside a moving ambulance, with background noise. Medical staff have their hands full and can’t keep pressing buttons. Nobody knows in advance what language the patient will speak. And what’s said is medical information, so privacy was the service’s first concern.
How speech technology solves it
The app works like an interpreter that takes turns on its own. It detects when someone starts speaking and opens the microphone without anyone touching it. A voice isolation system separates the voice from the noise of the street or the engine before transcribing it. Speech recognition identifies the language and turns it into text, and machine translation puts it into the other person’s language. Finally, speech synthesis reads it aloud, while the screen shows the written conversation in both languages.
The voice is never stored. Saving the transcript is optional and up to the service, according to its data protection criteria.
The proof-of-concept app
It’s an Android app designed to be used with one hand or hands-free. It has a conversation mode in which the turn passes automatically from one person to the other, and it recognizes dozens of languages.
Before the pilot we made a demo with real conversations between Spanish, English and Chinese, recorded without cuts. The pilot was defined from that demo, to be carried out in phases, including support.
What the pilot would measure
The pilot was designed to check three things in real conditions. First, how long translation takes and whether that time allows a fluid conversation. Second, how many transcription errors occur with the devices’ microphones and the usual noise. Third, whether the experience fits each place where medical staff work: the ambulance, the street or a home.