Glossary
Avatar video
Avatar video builds a presenter-led film from a synthetic person reading a script, which matters less because the presenter convinces than because the video can be revised by editing a sentence.
In plain terms
A video of a person who does not exist, reading what you wrote. The realism is not really the point. The point is that when the policy changes next year you fix the sentence and re-export, instead of booking a studio, a presenter and a day, which is what made this worth doing at all.
Why it matters
Because it changes what a video costs to keep rather than what it costs to make. A filmed piece is finished and expensive to revisit, so it drifts out of date and stays published anyway, while a scripted one is edited in the place the words live. Organisations adopt this for the second year, not the first.
How it works
The script is the source and the video is an export, which is the whole economic argument. Changing a line and producing the film again takes minutes and needs nobody's calendar, so material that would have aged for years because refilming was unaffordable can simply be corrected. That is a maintenance property rather than a production one.
One presenter can hold a curriculum together across time. A course built over two years with several filmed contributors looks like several projects, while a synthetic presenter is identical in the module recorded today and the one revised next spring, which quietly removes a consistency problem nobody budgets for.
Localisation is where it earns most, because the alternative is repeating the whole production. Finished material can be re-voiced into other languages with the mouth movements following, so one script reaches many markets without booking anybody again, and the cost per additional language collapses towards the translation.
Audiences accept it in some contexts and reject it firmly in others, and the boundary is about who is speaking rather than about quality. Procedural training, internal updates and product explanation are comfortable, while an apology, a message about redundancies or anything where a named leader is meant to be speaking personally are not. A synthetic presenter delivering sincerity reads as evasion.
What a video costs, and what it costs to keep
Seen in the wild
Presenter-led video without cameras, where a custom presenter can be built from a short self-recording and finished films re-voice into other languages.
HeyGenEnterprise presenter video for corporate learning, where courses update by editing text rather than refilming and one presenter runs across a curriculum.
SynthesiaAn editor treating the transcript as the timeline, so changing the words changes the film and repairs need no reshoot.
Descript
Common misconceptions
People assume
It succeeds or fails on how real the presenter looks.
In fact
Realism passed the threshold that matters some time ago and is not what teams buy it for. The purchase is editability: a video that can be corrected by changing a sentence stays accurate, and one that needs a studio to revise does not. Judge it on the second year rather than the first.
People assume
It suits any message.
In fact
The boundary is about who is speaking. Procedure, product explanation and routine updates are accepted comfortably, while apologies, difficult news and anything meant to come personally from a named leader are rejected sharply. A synthetic presenter delivering sincerity is read as an attempt to avoid delivering it.
Telling them apart
Avatar video vs Video generation
Avatar video
A person delivering a script to camera.
Any footage at all, from a description.
One is a narrow, solved, production-ready shape and the other is a broad and still-difficult capability, which is why the narrow one is in far more workplaces.
Questions
- Will people notice it is not a real person?
- Frequently yes, and in most workplace contexts they do not mind. Procedural training and internal updates carry it comfortably because nobody expected a personal performance. What fails is the message where the person is the point, and there the objection is not that it looks synthetic but that using it says something about how much the sender cared.
- Where does the money actually come back?
- In revision and in languages. A course corrected by editing a sentence stays current instead of quietly ageing, and one script re-voiced into several markets replaces repeating the whole production. Both are second-year benefits, which is why a first-year comparison against a single filmed video usually looks unconvincing.
- Can we use a real colleague as the presenter?
- Products support it, and it turns a production choice into an agreement with that person. What the likeness may be made to say, for how long, and what happens when they leave all need answering in terms they would recognise. Platforms built for organisations tend to build that into the process, which is a reason to prefer them here.
Key takeaways
- The script is the source; the video is an export you can redo.
- One presenter keeps a curriculum consistent across years.
- Localisation is where the cost per additional market collapses.
- Audiences reject it where the person speaking is the point.
Last checked July 2026