AI Audio Engineering: ElevenLabs & the Future of Sound
Sound is where cheap AI video gives itself away. How generative voice and audio actually fit a production, where they hold up, and the failure nobody warns you about.

ElevenLabs
RunwayShow someone a piece of AI video with the sound off and they will often think it looks fine. Turn the sound on and they will tell you within four seconds that something is wrong.
Audio is where this work gives itself away. Not resolution, not hands. Sound.
Part of that is technical and part of it is that audiences are far more forgiving of a strange-looking frame than a strange-sounding room. We have all seen stylised images. We have never heard a room that does not exist.
The three audio problems, which are not the same problem
People say "AI audio" as if it is one capability. In production it is three, with quite different maturity levels.
Voice is the most solved. Generated and cloned speech is genuinely good now, good enough that the limiting factor is direction rather than fidelity.
Ambience and foley is the least solved and the most overlooked. This is the sound of the space: room tone, traffic, footsteps, cloth, the hum of a fridge. It is what makes a shot feel located somewhere. Most AI video has none of it, which is precisely why it feels like it is floating.
Music sits in between. Generative scoring is capable, and it is also where I would push back hardest on replacing a human, because music carries the emotional argument of a piece and generative tools are still weak at structure over time.
Three problems. One label. That conflation is why people are surprised when their audio does not work.
Voice: the part that is genuinely ready
ElevenLabs is the tool most productions land on, and it deserves the position. I wrote a fuller walkthrough in the voice cloning guide, so here I will stick to what matters at the production level.
The mistake nearly everyone makes is treating it as text-to-speech. You paste a script, take the first result, and move on. That output is fine, and fine is the problem, because it lands in the exact middle of the uncanny range: too smooth to be a person, too human to read as deliberately synthetic.
What fixes it is directing the read the way you would direct a voice actor. Punctuation is your main instrument, and it does more work than any setting. A comma buys a beat. A full stop buys a longer one. Sentence fragments create emphasis that no amount of parameter-tuning will.
Write for the mouth, not the page. Read your script aloud first. If you run out of breath, the sentence is too long, and the generated read will sound like it is running out of breath too.
Then generate several takes and pick, rather than accepting the first. This is the habit that separates good results from average ones, and it is the same habit as directing anything else.
The gap most people never fill
Here is the thing that will improve your audio more than any tool choice.
Add room tone.
Generated video usually arrives either silent or with a thin native audio bed. Drop a continuous, quiet ambient layer underneath the whole sequence, appropriate to the space, and the footage stops feeling like it is floating in a vacuum. It costs almost nothing and it is the single highest-return audio move available.
Cuts are the other giveaway. Real sound does not stop dead at an edit. If your ambience hard-cuts with the picture, the ear notices even when the viewer cannot say why. Overlap it across the join and the sequence knits together.
Neither of those is glamorous and neither requires generative tools at all. They are just sound editing, and they are skipped constantly.
Where native audio changes the calculation
The Veo 3 line generates audio with the picture, and that is a genuinely different proposition from generating sound separately, because the sound is synchronised to that specific shot. Footsteps land on the footfalls. It removes a sync stage that used to eat real time.
It is not a complete answer. Native audio gives you the shot's own sound; it does not give you a consistent bed across a sequence assembled from several shots, and it will not match a client's brand-audio guidelines. I use it as a foundation layer and build over it, rather than as the finished mix. The dialogue-scene comparison has more on where it holds up.
And when the picture already exists, re-voicing is its own workflow. VEED Lipsync v2 covers the case where you have footage and need a different performance on top of it.
The failures worth knowing about
Being straight about where this breaks, because it will break.
Emotional range is narrower than it sounds. Conversational, warm, authoritative: all fine. Genuine grief, real anger, complicated humour: still not there. If your script depends on an emotional turn, cast a person.
Long reads drift. Energy and pace wander over several minutes. Generate in shorter sections and assemble, rather than requesting one long take.
Names and jargon get mangled. Product names, place names and technical terms are frequent casualties. Always listen to the whole thing before delivery. Always.
Consistency across sessions is not guaranteed. A voice generated today may not exactly match one generated next month, which matters on episodic work. Generate all of a series' audio in one pass where you can.
Mixing is still mixing. Generated audio arrives at wildly inconsistent levels. Without a proper loudness pass your video will be quieter or louder than everything else in a viewer's feed, and that alone reads as amateur.
The rule I keep coming back to
Good sound is invisible. Nobody watches something and thinks the ambience was well judged. They just believe the space.
Bad sound is the opposite. It is the first thing anyone notices, and it retroactively makes the picture look worse, which is a strange effect but a real one.
So if you have a limited amount of attention to spend on a project, spend a disproportionate share of it here. It is the cheapest quality gain available in AI video, and the most commonly skipped.
If you would rather have this handled as part of a finished piece, that is included in AI video production here. Or start a project and we can talk through what the piece actually needs.
Generate cinematic AI video — from €15
Five frontier models. No subscription. Buy credits, generate on demand, own the results outright.