A very common mistake when building a voice assistant is to reuse existing text: the website FAQ, the chatbot’s messages, a brochure. It’s well written, but it’s written to be read. When a synthetic voice reads it aloud, sentences drag on forever, options are forgotten before the list ends, and numbers sound strange. The problem lies in how the listener receives it.
Why reading and listening aren’t the same
A reader controls the pace. They can go back, skip a paragraph or look at the list of options while deciding. A listener can’t do any of that: the audio arrives word by word, at the speed the voice sets, and whatever they don’t retain is lost.
Kathryn Whitenton, of Nielsen Norman Group, explained it while analyzing voice assistants: just reciting a list of options forces people to hold them in working memory while they choose (Whitenton, 2016). On a screen you recognize the option you’re looking for; by ear you have to remember it.
It’s also slower. A meta-analysis by Marc Brysbaert covering 190 studies put silent reading in English at about 238 words per minute, and reading aloud at about 183 (Brysbaert, 2019). A long text takes more time to listen to than to read, and it demands more attention.
Amazon sums it up in its Alexa design guide: prompts are heard, not read, so they have to be written for a spoken conversation (Amazon).
Six adjustments for writing for the ear
1. One idea per sentence
Long sentences with subordinate clauses work on paper because your eyes can go back to the start. By ear, by the time the main verb arrives you’ve forgotten the subject. Break sentences up and keep one idea in each. If you can’t say a sentence in one breath, it’s too long.
2. Few options at a time
This sentence is fine in writing:
Would you like to: check traffic conditions, find out about road restrictions, handle paperwork at a traffic office, or get information phone numbers?
Said out loud, it forces the listener to hold four long options without knowing which one will be last. It’s better to offer fewer, shorter options, or to ask an open question first (“How can I help?”) and guide only if needed. When options are unavoidable, say first what each one gets you.
3. Punctuation is intonation too
A synthetic voice uses punctuation to decide where to pause and how to raise or lower its intonation. Compare:
Would you like to: check traffic conditions, road restrictions or paperwork?
What would you like? Traffic conditions? Road restrictions? Paperwork?
The second version sounds like a person offering options one by one, with a pause and a question intonation on each. A period instead of a comma, or a question split into several, changes a lot about how it sounds. It’s worth trying several versions with the real voice.
4. Write numbers and abbreviations the way they’re said
“On 3/1 at 9:30 a.m. on Main St.” forces the system to guess how to read each item, and sometimes it guesses wrong (is that March 1 or January 3?). If the text is fixed, write it the way it’s said: “on March first, at nine thirty in the morning, on Main Street.” If it comes from a database, check how dates, times, prices, abbreviations and proper names are read before publishing.
5. Mark pauses and emphasis when the text isn’t enough
Some things punctuation can’t express: a slightly longer pause before an important detail, a word that should sound stronger, a number that’s read digit by digit. That’s what SSML is for, a standard W3C markup language that lets you specify pauses, emphasis, speed, or how to interpret a number or a date (W3C, 2010). Not every voice supports every tag, so check what each one does.
6. Read it out loud
It’s the cheapest and most useful test. Read the text out loud, or better yet, listen to it in the voice you’ll use. If you stumble, if you have to breathe mid-sentence or if by the end you can’t remember the first option, the listener won’t be able to either.
When a language model writes the text
Many voice assistants today generate their answers with a language model, and these models write for screens by default: long paragraphs, bulleted lists, bold text, links. None of that can be heard. A bulleted list read aloud loses its structure, and a link turns into a string of letters.
That’s why the assistant’s instructions have to explicitly ask for a spoken style: short answers, no formatting, options one at a time and numbers written the way they’re said. And you have to check it by listening to real conversations, because the model tends to slip back into its writing habits.
All of this applies the maxims philosopher Paul Grice set out for any conversation: give the information that’s needed, no more and no less, and say it clearly and in order (Stanford Encyclopedia of Philosophy). By ear, breaking them is more noticeable.
Checklist
- Can each sentence be said without taking a breath halfway through?
- Are there at most two or three options in a row?
- Does the punctuation produce the pauses and questions you want?
- Are numbers, dates and abbreviations read correctly by the real voice?
- Have you listened to the whole thing in the voice you’ll use?
Writing for the ear is a central part of conversation design. If you’re interested in the topic, back in 2018 we put together eight tips for building good Alexa skills in Spanish that still hold.
References