
ElevenLabs
HeyGenElevenLabs Eleven v3 has moved voice cloning from novelty to production tool. I use it on every commercial project that involves narration, corporate training, or multilingual delivery. Here's the complete workflow. Voice cloning is a standard part of the professional AI video production services offered here.
The Source Recording
The single biggest determinant of clone quality is the source recording. ElevenLabs requires a minimum of one minute of audio, but the minimum is not the target: I record three to five minutes for every professional clone. Here's why, and how.
What to record:
- Varied sentence structures (declarative, question, exclamatory)
- Different emotional registers (neutral, warm, authoritative, conversational)
- Fast and slow delivery sections
- Words with complex phonemes — "particularly," "specifically," "extraordinary"
Recording conditions:
- Treated room or professional booth — no reverb, no background noise
- Condenser microphone at 24bit/48kHz minimum
- Consistent distance — don't move from the mic
- No breath noise between sentences — edit these out before uploading
What kills clone quality:
- Background noise (even low-level HVAC)
- Room reverb — sounds muffled even after cloning
- Phone/laptop mic recordings
- Inconsistent distance causing volume variation
The Script Format That Works
ElevenLabs reads your text and decides how to deliver it. You can guide it with formatting:
Punctuation controls pacing:
- Full stops create natural pauses
- Commas create shorter pauses
- Em dashes — like this — create a mid-thought pause with slight emphasis
- Ellipsis... creates a trailing, uncertain pause
Capitalization for emphasis:
- ALL CAPS on a word creates heavy stress
- Avoid it for more than one word per sentence — it sounds artificial
Emotional tags (v3 feature): Use ElevenLabs' audio tags for explicit emotional direction:
<happy>Great news on that front.</happy>
<serious>This part is important.</serious>
<whisper>Just between us...</whisper>
These are available in Eleven v3 and significantly improve performance on scripts that require emotional range.
Clone Settings
In the ElevenLabs voice settings panel:
Stability: Lower stability (0.3–0.5) produces more natural variation, better for conversational tone. Higher stability (0.7–0.9) is more consistent, better for long-form narration.
Similarity: Keep at 0.75–0.85 for production use. Very high similarity can introduce artifacts.
Style exaggeration: 0 for neutral delivery. Increase carefully — it amplifies the expressive patterns in the original recording, which can over-stylise.
Speaker boost: On for voice clones that are thin or lack presence. Off for naturally warm voices.
Multilingual Delivery
Eleven v3 handles 70+ languages with the same cloned voice. The quality varies by language family:
- European Latin languages (French, Spanish, Italian, Portuguese): Excellent — natural cadence, correct phoneme mapping
- Germanic languages (German, Dutch, Swedish): Very good
- Eastern European (Polish, Czech, Romanian): Good — slight accent in complex words
- East Asian languages (Japanese, Korean, Mandarin): Functional for corporate use, not yet native-grade
For multilingual campaign work, always have a native speaker review before delivery. Not for the voice quality, but for the script itself. ElevenLabs will speak whatever you write, including grammatically wrong constructions.
Integration with HeyGen
The workflow for AI presenter videos:
- Clone voice in ElevenLabs, export as WAV
- Import to HeyGen as custom voice
- Apply to avatar of choice
- HeyGen syncs lip movement to ElevenLabs audio
For English-language content this beats HeyGen's own voice generation on lip sync, because ElevenLabs' prosody is simply more natural. The difference is audible. For multilingual content, use HeyGen's built-in translation pipeline: it handles language-specific lip sync better.
Cost Management
ElevenLabs bills by character. A 60-second script at average speaking pace is approximately 900–1,100 characters, which at Creator tier rates works out to a few cents. Pennies, effectively.
Where cost accumulates is iteration. Running 20 variations of a 1,000-character script across different emotional registers adds up quickly enough to notice on the invoice. Write the script once, correctly, then generate. The prompt engineering discipline that applies to video applies equally here.
Consent, rights and the part that gets skipped
Cloning a voice is a technical step with a legal shape around it, and this is where corporate projects come unstuck eighteen months later rather than on day one.
Get written permission for the specific uses. Not general permission to "use my voice". A clone made for internal training and later used in a paid advertisement is a different deal, and the person whose voice it is will reasonably think so too.
Agree what happens when they leave. If the presenter is an employee, decide up front whether the clone is retired, and who owns it. Nobody wants this conversation during an exit.
Never clone a voice you do not have rights to. This includes public figures, actors and anyone whose voice you sampled from a podcast. It is the fastest route to a takedown and a genuinely deserved reputational problem.
Check platform terms for your delivery channel. Disclosure requirements vary and are tightening, particularly for advertising.
None of this is difficult. It is just administrative work that feels unnecessary until the moment it is the only thing that matters, and it costs an email to get right at the start.
What v3 Still Can't Do
- Heavy regional accents: The clone will soften the source accent. If a strong Irish, Scottish, or regional American accent is required, the source recording needs to be very long (10+ minutes) and the accent very consistent.
- Singing: ElevenLabs has a separate music product. The voice clone tool is not for singing.
- Real-time: The API has a streaming mode but it's not production-ready for latency-sensitive applications.
The Bottom Line
A well-recorded source voice + a well-formatted script + correct v3 settings produces output that passes broadcast review. I've delivered ElevenLabs voice to clients who had no idea it was AI-generated. Not one of them asked. The pipeline is mature — the skill is in the setup. The free tier is enough to test the cloning quality on your own voice before committing, so you can try ElevenLabs here. Voice is one layer of a complete AI video production workflow. If you'd like it built into a full production for you, start a project.
Generate cinematic AI video — from €15
Five frontier models. No subscription. Buy credits, generate on demand, own the results outright.