Glossary
Speech-to-text (STT)
Speech-to-text turns recorded or live speech into written words, which is what produces a meeting transcript, a set of live captions or a searchable record of a call.
In plain terms
Sound in, words out, and nothing else. It does not decide what to do about what was said, does not summarise it and does not answer it. That narrowness is the appeal: it is a far simpler thing to buy and to evaluate than a system that holds a conversation, and it is what most organisations actually want.
Why it matters
It turns speech into something you can search, quote, summarise and keep, which is a larger change than it sounds because spoken material was previously almost unusable after the fact. A meeting nobody minuted, a call nobody recorded, a site visit nobody wrote up: all of those become records. That is also the reason to be careful, since a recording of a person is a different kind of material from a document about them.
How it works
Audio is converted into words, and modern systems do this well enough on clear speech that the interesting questions are all about the unclear cases. Background noise, several people talking at once, unfamiliar accents and specialist vocabulary are where accuracy falls away, and unfortunately those conditions describe a great many real meetings.
Supplying a list of expected terms is the fix nobody switches on. Most systems accept your product names, your colleagues' names and your industry jargon, and bias recognition towards them. It is usually the single most effective improvement available and it is off by default in a great many deployments.
Working out who said what is a separate capability from transcribing, and products differ enormously in whether they offer it and how well it survives people interrupting each other. If your use is minutes rather than captions, that feature is the one to test rather than raw accuracy.
Live and recorded are different problems. Live transcription must commit to a word before hearing what follows, which costs accuracy; a recording can be reconsidered in full. A product excellent at one is not thereby good at the other, and vendors often quote figures from the easier case.
What a vendor measured, and what you have
Seen in the wild
Dictate a paragraph containing three proper nouns and a figure, then read the transcript rather than the result, which shows you what the system actually heard.
Google GeminiPut a transcription step into an automation and feed it a genuinely difficult recording rather than a clear one, which is the test that predicts what you will actually get.
n8nSpeak to an assistant and watch the words appear as you talk, which is the live case committing to each word before hearing the next.
ChatGPT
Common misconceptions
People assume
The accuracy figure tells us what to expect.
In fact
It tells you what a vendor measured, usually on clear speech in a common accent with no jargon. Your calls have background noise, people talking over each other and product names nobody outside your industry has heard. Run twenty minutes of your own worst audio, which is a better predictor than any published number.
People assume
It understands the meeting.
In fact
It produces words. Everything after that, the summary, the action points, the decisions, is a language model reading the transcript, which means an error at this stage is inherited with confidence by everything downstream. That is why reading the transcript rather than the summary is the useful check.
Telling them apart
Speech-to-text vs Voice AI
Speech-to-text
The first stage on its own. Produces a transcript and stops, which makes it a straightforward thing to evaluate and buy.
The whole conversation: listening, deciding and answering, with all the timing and interruption problems that come from three systems in sequence.
If you want a record of what was said, you want this one and the purchase is much simpler. The complications belong to the parts that come after it.
Questions
- It keeps getting our product names wrong. What do we do?
- Supply them. Most systems accept a list of expected terms, names and jargon and bias recognition towards those words, and it is usually the most effective single change available. It is also commonly not enabled by default, so it is worth asking specifically rather than assuming a vendor has done it.
- Can it tell who was speaking?
- Some products can and they vary considerably, particularly when people interrupt each other. If you want minutes attributed to individuals rather than a wall of text, test that capability on a real recording with several voices, because it fails differently from transcription and is often quoted less honestly.
- What happens to the recordings?
- That is a data question rather than a technical one, and it is sharper here than for text because a recording carries a person's identity. Establish how long audio is kept, whether transcripts are treated differently, who can access either, and what participants are told, before a tool is pointed at a real meeting.
Key takeaways
- It produces words and nothing else, which makes it a far simpler purchase than a full voice system.
- Noise, overlapping speakers, accents and jargon are where accuracy falls, and those conditions describe a great many real meetings.
- A list of expected terms is the most effective fix available and is usually switched off.
- Identifying who spoke is a separate capability that varies widely between products.
- Errors here are inherited with confidence by every summary built on the transcript.
Tools that use this
- Google Gemini
Dictation with a visible transcript, so you see what was heard.
- n8n
A transcription step you can feed genuinely difficult audio.
- ChatGPT
The live case, committing to each word before the next arrives.
Last checked July 2026