New · Veo 3.1 live — up to 4K with native audio
Intelligence Feed
industry2026-04-017 min readRichard Byrne

Mastering Digital Humans for Cinematic Narrative

AI avatars are excellent at a narrow set of jobs and poor at everything else. Where digital humans genuinely work, where they fail, and how to direct one properly.

Mastering Digital Humans for Cinematic Narrative
Tools covered
Synthesia logoSynthesia
HeyGen logoHeyGen
ElevenLabs logoElevenLabs

Let me start with the unpopular part, because it saves everyone time.

Digital humans are not good at drama. They are very good at a specific set of commercial jobs, and if you try to use them outside that set you will produce something that makes viewers uncomfortable without being able to explain why.

Knowing which side of that line your project sits on is most of the skill.

What avatars are genuinely good at

The jobs where digital humans consistently earn their place share a shape: information delivery, at volume, in more than one language, where the speaker is a presenter rather than a character.

Internal training and onboarding. Nobody needs a performance. They need clear information from a consistent presenter, and they need it updated whenever a policy changes. Re-shooting a human for a policy change is absurd. Regenerating an avatar is trivial.

Multilingual versions of the same message. This is the strongest case by a distance. One script, a dozen languages, the same presenter throughout. Doing that with human talent means a dozen shoots or a dozen actors, and the cost curve is brutal. I went into the corporate use case in more detail in the HeyGen multilingual piece.

Product explainers and documentation. Content that changes often and needs to stay current. The value is not that the avatar is convincing. It is that updating it costs almost nothing.

Content where the alternative is nothing. A great deal of corporate video simply would not be made at the cost of a shoot. Judged against not existing, an avatar explainer is a clear win.

Notice what these have in common. The viewer is there for the information. They are not there for the person.

Where they fall down

The moment a viewer is meant to care about the speaker rather than listen to them, digital humans struggle. That is not a fidelity problem and better rendering will not fix it.

Performance requires micro-behaviour: the pause before a difficult word, the glance away while thinking, the slight loss of composure. Avatars are fluent and even, which is exactly wrong for emotional content. Fluency reads as insincerity when the subject matter is serious.

Some specific traps:

Anything with emotional stakes. Apologies, bad news, testimonial, brand-purpose films. If a real person should be saying it, an avatar saying it is worse than no video.

Long unbroken takes. The uncanny effect compounds with duration. Thirty seconds is comfortable. Four unbroken minutes of an avatar talking is a hard watch, regardless of quality.

Anything implying lived experience. An avatar saying "when I started out" is a claim nobody is making truthfully, and audiences pick up on it faster than you would expect.

Where the technology got to in 2026

Two releases moved the floor, and both are worth knowing because they change what you should expect from a demo.

Synthesia 3.0 shipped in October 2025 with the Express-2 avatar engine, adding full-body movement, hand gestures and micro-expressions rather than the fixed-torso presenter that defined the category. Courses, Interactivity 2.0 and AI dubbing came with it, which tells you where the product is aimed: structured learning rather than film.

HeyGen's Avatar V arrived in early 2026 and attacks the problem that made avatars hard to use in sequences. It builds a digital twin from a short webcam recording and holds identity consistent across videos rather than drifting between them, which is the difference between one usable clip and a maintainable library.

Both narrow the uncanny gap. Neither closes the one that actually matters, which is performance rather than fidelity, and the section above on emotional stakes still applies exactly as written.

Choosing between platforms

Synthesia and HeyGen dominate this space and they suit different briefs, which I compared properly in HeyGen vs Synthesia.

The short version: Synthesia is built around enterprise workflow, with a library approach and strong governance features, which suits organisations producing training at scale. HeyGen leans toward flexibility and custom avatars, which suits marketing and creator work where the presenter is part of the brand.

Neither is better. They answer different questions, and the question worth asking first is whether you need a presenter or a person. If a stock avatar will do, you have a straightforward tooling decision. If it must be a specific individual, you are in custom-avatar territory, which means consent, likeness rights and a considerably longer setup.

Directing one properly

Assume you are in the right use case. Most avatar video is bad for reasons that have nothing to do with the avatar.

Write for speech. The commonest failure by far is feeding it a document. Written prose and spoken language are different registers, and an avatar reading corporate written English sounds exactly like what it is. Short sentences. Contractions. One idea per sentence.

Cut away constantly. This is the single biggest quality lever and it costs nothing. Do not hold on the avatar. Cut to screen recordings, product shots, diagrams, b-roll. The avatar becomes a narrator rather than a subject, and the uncanny effect largely dissolves. Most watchable avatar video is under thirty percent avatar on screen.

Fix the audio separately. Platform-native voices are serviceable. A properly directed ElevenLabs read is better, and audio carries more of the perceived quality than the face does. There is more on that in the AI audio piece.

Do not fake the background. A composited avatar in a "real" office reads as false, because the lighting never quite matches. A clean, deliberately designed background reads as a considered choice. Own the artifice and it stops being a flaw.

Keep it short. Two minutes beats eight, nearly always. If the script runs long, split it into several videos, which is better structure anyway.

The disclosure question

Worth saying plainly: tell people.

Not in a legal disclaimer nobody reads. Just do not construct the piece to imply a human presenter when there is not one. Audiences are increasingly good at spotting avatars, and the reputational cost of being caught implying otherwise is far higher than the cost of being upfront. For internal training this is a non-issue. For public-facing marketing it is a real decision, and the honest option is also the safe one.

The summary I would give a client

Digital humans are a production tool with a narrow, genuine, valuable use case. Inside it they are transformative, mostly because they collapse the cost of updating and translating. Outside it they are a false economy that makes work feel cheap.

The question is never "can we use an avatar here". It is "is the viewer here for the information or for the person". Answer that honestly and the tooling decision makes itself.

If you want help working out which side of that line a project sits on, that conversation is part of AI video production here, and you can get in touch directly. Or look at the gallery and judge the output yourself, which is the better test anyway.

digital humanssynthesiaheygennarrative
Ready to create?

Generate cinematic AI video — from €15

Five frontier models. No subscription. Buy credits, generate on demand, own the results outright.