Glossary
Text-to-speech (TTS)
Text-to-speech turns written words into spoken audio, which is what narrates a document, reads a page aloud for somebody who cannot easily read it, or gives an assistant a voice.
In plain terms
Words in, sound out. The quality has changed enormously and that is the whole story: what used to announce itself as a machine now frequently does not, which is excellent for anybody listening and raises questions nobody had to ask when it sounded obviously artificial.
Why it matters
It makes written material available to people who cannot easily read it, which is a genuine accessibility gain rather than a convenience, and it makes long documents consumable while doing something else. It also carries the roster's sharpest disclosure question. A voice that sounds like a person will be taken for one by whoever hears it, and deciding whether to tell them is a choice your organisation makes rather than a setting a product provides.
How it works
Text is converted into audio in one pass, with the model deciding pronunciation, pacing and emphasis from the words alone. That is why it stumbles on things where the written form does not determine the spoken one: names, abbreviations, figures read as dates, and sentences whose meaning depends on which word is stressed.
Most products let you steer the delivery to some extent, whether by choosing a voice, adjusting pace, or marking up the text to indicate emphasis and pauses. How much control you get differs considerably and is worth testing on your own material, because a passable default and a good result are often a few markings apart.
Voice cloning, producing speech in a particular person's voice from samples of it, is a distinct capability with distinct obligations. Consent from the person whose voice it is belongs in writing rather than in your assumption, and that applies to a colleague who agreed cheerfully as much as to anybody else.
It is the last stage of a voice assistant and largely independent of the parts before it. That means a convincing voice can be wrapped around a poor transcript and a weak answer, and will sound more trustworthy than a plain voice saying something correct, which is worth remembering when judging a demonstration.
Two things a convincing voice does at once
Seen in the wild
Have an assistant read a reply aloud and listen for what it does with a name, an abbreviation and a figure, which is where the written form does not determine the spoken one.
ChatGPTAsk for the same paragraph read twice with different pacing and notice how much of the impression is delivery rather than content.
Google GeminiAdd a spoken step to an automation and decide, explicitly, whether the person hearing it is told it is not a colleague.
n8n
Common misconceptions
People assume
A natural voice means a capable system.
In fact
This is the last stage and is largely independent of everything before it. Convincing delivery can be wrapped around a mistaken transcript and a weak answer, and it will sound more trustworthy than a flat voice saying something correct. Judge the words, not the reading of them.
People assume
Using somebody's voice is fine if they work here.
In fact
Consent should be explicit, in writing, and specific about what the voice may be used for and for how long. People change roles and leave organisations, and an agreement made informally in a good mood is exactly the one that becomes awkward later. That is a small piece of paperwork now and a genuine problem otherwise.
Telling them apart
Text-to-speech vs Speech-to-text
Text-to-speech
Words to audio. The output stage, judged on how it sounds and on whether the listener should be told what they are hearing.
Audio to words. The input stage, judged on accuracy against difficult recordings.
Opposite directions, and their problems have nothing in common. One is an accuracy question and the other is largely a disclosure question.
Questions
- Should we tell people they are hearing a synthetic voice?
- For anything where somebody might reasonably believe they are speaking to a person, yes, and increasingly a regulator will expect it. Beyond compliance it is a straightforward matter of not misleading somebody about who they are dealing with. Decide it as a policy rather than leaving it to whoever configures the system.
- How do we stop it mispronouncing our names?
- Most products accept guidance, whether a pronunciation list or markings in the text, and applying it to your product names and the people who appear regularly is a short job. Test with the awkward ones rather than the easy ones, since a name spelled conventionally and said unconventionally is the usual problem.
- Is cloning a voice legal?
- It depends on consent, on jurisdiction and on what the voice is used for, and it is a question for somebody qualified rather than for this page. What holds everywhere is that written, specific consent from the person whose voice it is puts you in a considerably better position than any other arrangement.
Key takeaways
- It turns words into audio and decides pronunciation and pacing from the text alone.
- Names, abbreviations and figures are where the written form does not determine the spoken one.
- The voice is the last stage and independent of whether what it says is correct.
- Cloning a voice needs explicit written consent, including from colleagues who agreed cheerfully.
- Whether a listener is told they are hearing a machine is a policy decision, not a product setting.
Tools that use this
- ChatGPT
Spoken replies, where names and figures show the weak spots.
- Google Gemini
The same paragraph delivered differently, which separates delivery from content.
- n8n
A spoken step, where the disclosure decision has to be made explicitly.
Last checked July 2026