Skip to content

Glossary

Speech-to-text (STT)

Speech-to-text turns recorded or live speech into written words, which is what produces a meeting transcript, a set of live captions or a searchable record of a call.

In plain terms

Sound in, words out, and nothing else. It does not decide what to do about what was said, does not summarise it and does not answer it. That narrowness is the appeal: it is a far simpler thing to buy and to evaluate than a system that holds a conversation, and it is what most organisations actually want.

01

Why it matters

It turns speech into something you can search, quote, summarise and keep, which is a larger change than it sounds because spoken material was previously almost unusable after the fact. A meeting nobody minuted, a call nobody recorded, a site visit nobody wrote up: all of those become records. That is also the reason to be careful, since a recording of a person is a different kind of material from a document about them.

02

How it works

Audio is converted into words, and modern systems do this well enough on clear speech that the interesting questions are all about the unclear cases. Background noise, several people talking at once, unfamiliar accents and specialist vocabulary are where accuracy falls away, and unfortunately those conditions describe a great many real meetings.

Supplying a list of expected terms is the fix nobody switches on. Most systems accept your product names, your colleagues' names and your industry jargon, and bias recognition towards them. It is usually the single most effective improvement available and it is off by default in a great many deployments.

Working out who said what is a separate capability from transcribing, and products differ enormously in whether they offer it and how well it survives people interrupting each other. If your use is minutes rather than captions, that feature is the one to test rather than raw accuracy.

Live and recorded are different problems. Live transcription must commit to a word before hearing what follows, which costs accuracy; a recording can be reconsidered in full. A product excellent at one is not thereby good at the other, and vendors often quote figures from the easier case.

What a vendor measured, and what you have

What a vendor measured, and what you haveNothing about the left column is dishonest; it is simply the condition under which any system performs best, and it is the condition a demonstration will naturally be arranged in. Your material is the right column, and every line of it costs accuracy for reasons that have nothing to do with how good the system is in general. The useful response is not scepticism about the figures but a different test. Take the worst recording you actually have, the one with the poor microphone and the technical vocabulary and the people talking over each other, run it, and read the transcript against the audio. That takes twenty minutes, it is the evidence that actually transfers to your situation, and it usually also tells you which terms to put on the list that would fix half of what went wrong.The demonstration audioOne speaker, close microphone.A common accent.Ordinary vocabulary.Nobody interrupting.Your Tuesday meetingSix people, one laptopmicrophone.Four accents.Product names and industryjargon.Two of them talking at once.Accuracy figures come from theleft column and your experiencecomes from the right. Twentyminutes of your own worstrecording tells you more thanany published number, and it isthe one test that transfers toyour situation.
Nothing about the left column is dishonest; it is simply the condition under which any system performs best, and it is the condition a demonstration will naturally be arranged in. Your material is the right column, and every line of it costs accuracy for reasons that have nothing to do with how good the system is in general. The useful response is not scepticism about the figures but a different test. Take the worst recording you actually have, the one with the poor microphone and the technical vocabulary and the people talking over each other, run it, and read the transcript against the audio. That takes twenty minutes, it is the evidence that actually transfers to your situation, and it usually also tells you which terms to put on the list that would fix half of what went wrong.
03

Seen in the wild

  • Dictate a paragraph containing three proper nouns and a figure, then read the transcript rather than the result, which shows you what the system actually heard.

    Google Gemini
  • Put a transcription step into an automation and feed it a genuinely difficult recording rather than a clear one, which is the test that predicts what you will actually get.

    n8n
  • Speak to an assistant and watch the words appear as you talk, which is the live case committing to each word before hearing the next.

    ChatGPT
04

Common misconceptions

People assume

The accuracy figure tells us what to expect.

In fact

It tells you what a vendor measured, usually on clear speech in a common accent with no jargon. Your calls have background noise, people talking over each other and product names nobody outside your industry has heard. Run twenty minutes of your own worst audio, which is a better predictor than any published number.

People assume

It understands the meeting.

In fact

It produces words. Everything after that, the summary, the action points, the decisions, is a language model reading the transcript, which means an error at this stage is inherited with confidence by everything downstream. That is why reading the transcript rather than the summary is the useful check.

05

Telling them apart

Speech-to-text vs Voice AI

Speech-to-text

The first stage on its own. Produces a transcript and stops, which makes it a straightforward thing to evaluate and buy.

Voice AI

The whole conversation: listening, deciding and answering, with all the timing and interruption problems that come from three systems in sequence.

If you want a record of what was said, you want this one and the purchase is much simpler. The complications belong to the parts that come after it.

06

Questions

It keeps getting our product names wrong. What do we do?
Supply them. Most systems accept a list of expected terms, names and jargon and bias recognition towards those words, and it is usually the most effective single change available. It is also commonly not enabled by default, so it is worth asking specifically rather than assuming a vendor has done it.
Can it tell who was speaking?
Some products can and they vary considerably, particularly when people interrupt each other. If you want minutes attributed to individuals rather than a wall of text, test that capability on a real recording with several voices, because it fails differently from transcription and is often quoted less honestly.
What happens to the recordings?
That is a data question rather than a technical one, and it is sharper here than for text because a recording carries a person's identity. Establish how long audio is kept, whether transcripts are treated differently, who can access either, and what participants are told, before a tool is pointed at a real meeting.
07

Key takeaways

  • It produces words and nothing else, which makes it a far simpler purchase than a full voice system.
  • Noise, overlapping speakers, accents and jargon are where accuracy falls, and those conditions describe a great many real meetings.
  • A list of expected terms is the most effective fix available and is usually switched off.
  • Identifying who spoke is a separate capability that varies widely between products.
  • Errors here are inherited with confidence by every summary built on the transcript.
09

Tools that use this

  • Google Gemini

    Dictation with a visible transcript, so you see what was heard.

  • n8n

    A transcription step you can feed genuinely difficult audio.

  • ChatGPT

    The live case, committing to each word before the next arrives.

Last checked July 2026

All glossary terms