Skip to content

Glossary

Text-to-speech (TTS)

Text-to-speech turns written words into spoken audio, which is what narrates a document, reads a page aloud for somebody who cannot easily read it, or gives an assistant a voice.

In plain terms

Words in, sound out. The quality has changed enormously and that is the whole story: what used to announce itself as a machine now frequently does not, which is excellent for anybody listening and raises questions nobody had to ask when it sounded obviously artificial.

01

Why it matters

It makes written material available to people who cannot easily read it, which is a genuine accessibility gain rather than a convenience, and it makes long documents consumable while doing something else. It also carries the roster's sharpest disclosure question. A voice that sounds like a person will be taken for one by whoever hears it, and deciding whether to tell them is a choice your organisation makes rather than a setting a product provides.

02

How it works

Text is converted into audio in one pass, with the model deciding pronunciation, pacing and emphasis from the words alone. That is why it stumbles on things where the written form does not determine the spoken one: names, abbreviations, figures read as dates, and sentences whose meaning depends on which word is stressed.

Most products let you steer the delivery to some extent, whether by choosing a voice, adjusting pace, or marking up the text to indicate emphasis and pauses. How much control you get differs considerably and is worth testing on your own material, because a passable default and a good result are often a few markings apart.

Voice cloning, producing speech in a particular person's voice from samples of it, is a distinct capability with distinct obligations. Consent from the person whose voice it is belongs in writing rather than in your assumption, and that applies to a colleague who agreed cheerfully as much as to anybody else.

It is the last stage of a voice assistant and largely independent of the parts before it. That means a convincing voice can be wrapped around a poor transcript and a weak answer, and will sound more trustworthy than a plain voice saying something correct, which is worth remembering when judging a demonstration.

Two things a convincing voice does at once

Two things a convincing voice does at onceIt is worth noticing that the two columns are the same fact seen twice. Everything valuable about a convincing synthetic voice comes from it being convincing, and so does everything that now requires a decision. A listener who cannot tell may reasonably assume they are speaking to a colleague. A voice that sounds like a specific person is only usable with that person's written agreement. And delivery is persuasive in a way that has nothing to do with whether the words are right, which flatters a demonstration and misleads a customer. None of that argues against using it. It argues for making one explicit decision, early, about whether people are told what they are hearing, and for making it as a policy rather than leaving it to whoever happens to configure the system.What it earnsWritten material becomeslistenable.Access for people who cannotread it easily.Long documents consumed whiledriving.An assistant that answers outloud.What it obligesA listener may think it is aperson.A cloned voice needs writtenconsent.Delivery reads as competence.Somebody has to decide aboutdisclosure.The right-hand column arrivedwith the quality improvementrather than separately. When theoutput announced itself as amachine none of it applied, andthe better it sounds the more ofit does.
It is worth noticing that the two columns are the same fact seen twice. Everything valuable about a convincing synthetic voice comes from it being convincing, and so does everything that now requires a decision. A listener who cannot tell may reasonably assume they are speaking to a colleague. A voice that sounds like a specific person is only usable with that person's written agreement. And delivery is persuasive in a way that has nothing to do with whether the words are right, which flatters a demonstration and misleads a customer. None of that argues against using it. It argues for making one explicit decision, early, about whether people are told what they are hearing, and for making it as a policy rather than leaving it to whoever happens to configure the system.
03

Seen in the wild

  • Have an assistant read a reply aloud and listen for what it does with a name, an abbreviation and a figure, which is where the written form does not determine the spoken one.

    ChatGPT
  • Ask for the same paragraph read twice with different pacing and notice how much of the impression is delivery rather than content.

    Google Gemini
  • Add a spoken step to an automation and decide, explicitly, whether the person hearing it is told it is not a colleague.

    n8n
04

Common misconceptions

People assume

A natural voice means a capable system.

In fact

This is the last stage and is largely independent of everything before it. Convincing delivery can be wrapped around a mistaken transcript and a weak answer, and it will sound more trustworthy than a flat voice saying something correct. Judge the words, not the reading of them.

People assume

Using somebody's voice is fine if they work here.

In fact

Consent should be explicit, in writing, and specific about what the voice may be used for and for how long. People change roles and leave organisations, and an agreement made informally in a good mood is exactly the one that becomes awkward later. That is a small piece of paperwork now and a genuine problem otherwise.

05

Telling them apart

Text-to-speech vs Speech-to-text

Text-to-speech

Words to audio. The output stage, judged on how it sounds and on whether the listener should be told what they are hearing.

Speech-to-text

Audio to words. The input stage, judged on accuracy against difficult recordings.

Opposite directions, and their problems have nothing in common. One is an accuracy question and the other is largely a disclosure question.

06

Questions

Should we tell people they are hearing a synthetic voice?
For anything where somebody might reasonably believe they are speaking to a person, yes, and increasingly a regulator will expect it. Beyond compliance it is a straightforward matter of not misleading somebody about who they are dealing with. Decide it as a policy rather than leaving it to whoever configures the system.
How do we stop it mispronouncing our names?
Most products accept guidance, whether a pronunciation list or markings in the text, and applying it to your product names and the people who appear regularly is a short job. Test with the awkward ones rather than the easy ones, since a name spelled conventionally and said unconventionally is the usual problem.
Is cloning a voice legal?
It depends on consent, on jurisdiction and on what the voice is used for, and it is a question for somebody qualified rather than for this page. What holds everywhere is that written, specific consent from the person whose voice it is puts you in a considerably better position than any other arrangement.
07

Key takeaways

  • It turns words into audio and decides pronunciation and pacing from the text alone.
  • Names, abbreviations and figures are where the written form does not determine the spoken one.
  • The voice is the last stage and independent of whether what it says is correct.
  • Cloning a voice needs explicit written consent, including from colleagues who agreed cheerfully.
  • Whether a listener is told they are hearing a machine is a policy decision, not a product setting.
09

Tools that use this

  • ChatGPT

    Spoken replies, where names and figures show the weak spots.

  • Google Gemini

    The same paragraph delivered differently, which separates delivery from content.

  • n8n

    A spoken step, where the disclosure decision has to be made explicitly.

Last checked July 2026

All glossary terms