Glossary
Voice cloning
Voice cloning builds a synthetic copy of one person's voice from recordings of them, which makes the first question who agreed to it rather than how well it works.
In plain terms
Making a synthetic copy of one particular person's voice from recordings of them. The technology stopped being the interesting part some time ago: it works, from surprisingly little material. What remains interesting is that the thing being copied belongs to somebody, and that person may not have understood what they were agreeing to.
Why it matters
Because the usual controls do not reach it. Retention policies, access rules and residency commitments all describe what happens to data, and a voice is not really being handled as data here, it is being reproduced as a person. Somebody can consent to a recording being stored and still be astonished to hear themselves saying something they never said.
How it works
It learns a voice from samples rather than storing recordings to replay, which is why so little material is needed. A short passage of clean speech is frequently enough to produce a copy that reads new text convincingly, and quality follows recording conditions far more than length. Background noise costs more than brevity does.
Consent is the first question and it is more specific than a signature suggests. Agreeing to a recording is not agreeing to a synthetic copy, agreeing to a copy for one campaign is not agreeing to it indefinitely, and agreeing at all is not agreeing to any script somebody later writes. The serious products build permission into the process rather than treating it as paperwork alongside it.
The practical limits are about performance rather than timbre. A copy handles ordinary delivery well and struggles with the things a director would ask for: genuine emotion, comic timing, a deliberate pause that carries meaning. It sounds like the person and does not act like them, which matters most in exactly the work people want it for.
Its most valuable uses are unglamorous and repair-shaped. Fixing a misspoken line without recalling a presenter, keeping one voice consistent across a course that gets updated for years, and re-voicing finished material into other languages are all cases where the alternative is expensive rescheduling. That is where it earns its place rather than in wholesale replacement.
Provenance is becoming part of the deliverable. Where synthetic speech reaches an audience, being able to say what was generated and on whose authority is increasingly asked for, and platforms built for organisations tend to carry that machinery while the fastest consumer tools do not. Which you have chosen matters before publication rather than after.
Two questions that get answered by different departments
Seen in the wild
An audio platform whose voice work is explicitly permission-based, alongside dubbing finished material into other languages and generating effects.
ElevenLabsAn editor where the copy exists to repair a flubbed line without recalling anybody, inside a workflow that treats the transcript as the timeline.
DescriptAn enterprise video platform whose custom presenters are built with consent as part of the process rather than as paperwork beside it.
Synthesia
Common misconceptions
People assume
The consent question is answered by our data policy.
In fact
Data policies describe storage, access and deletion, and none of those is what a person objects to here. The concern is being made to say things, which is a question about a likeness rather than about a file, so a compliant retention arrangement can sit alongside a use nobody agreed to.
People assume
It needs a lot of material.
In fact
A short clean passage is frequently enough, and recording conditions matter far more than duration. That is precisely why the permission question is urgent rather than theoretical: the raw material for a convincing copy already exists for anybody who has appeared on a recorded call or a published video.
People assume
It can replace a performance.
In fact
It reproduces a voice and not an actor. Ordinary delivery is handled well while genuine emotion, comic timing and a meaningful pause remain weak, so it sounds like the person without acting like them. The gap shows up hardest in exactly the work people are most tempted to use it for.
Telling them apart
Voice cloning vs Text to speech
Voice cloning
This particular person's voice, copied.
A synthetic voice belonging to nobody.
Only one of them involves somebody who can object, which is why the same product can be uncontroversial in one mode and need a signed agreement in the other.
Questions
- What should a consent form actually cover?
- What the copy may say, for how long, in which contexts, and how it ends. A signature agreeing to a recording covers none of those, and the gap is where the disputes live. Products built for organisations tend to prompt for this because their customers discovered the hard way that a general release does not survive a specific objection.
- How much audio does it take?
- Less than people expect, often a short clean passage. Quality tracks recording conditions rather than duration, so a quiet few minutes beats a noisy hour. The practical consequence is that anybody with published video or recorded calls already has enough material in existence, which is what makes permission a live question rather than a future one.
- Where is it genuinely worth using?
- Where the alternative is getting a person back into a studio. Fixing one misspoken line, keeping a single voice consistent across material updated over years, and re-voicing finished work into other languages all replace expensive rescheduling with an edit. Those cases justify it far more comfortably than replacing a performance does.
- Do we need to disclose it?
- Increasingly yes, and it is worth deciding before publication rather than after. Being able to say what was synthetic and on whose authority is becoming an expectation for material that reaches an audience, and platforms built for organisations tend to carry that machinery while the quickest consumer tools do not.
Key takeaways
- It learns a voice from samples, so very little material is needed.
- Consent is about a likeness, which is why data policies do not answer it.
- It reproduces a voice and not a performance; timing and emotion stay weak.
- Its best uses are repair-shaped: fixing, sustaining, re-voicing.
- Disclosure machinery differs by product and matters before publication.
Tools that use this
- ElevenLabs
Permission-based voice work, dubbing and generated effects.
- Descript
A copy that exists to repair a line without recalling anybody.
- Synthesia
Custom presenters with consent built into the process.
Last checked July 2026