Skip to content

Mariano Rodrigo

AI Solutions Engineer building production systems with artificial intelligence, automation, and full-stack architecture. This is my public engineering lab: architecture decisions, implementation reports, experiments, and lessons from real systems.

Generative video AI with HeyGen: one avatar, seven languages

The experiment in 3 minutes

My previous experiment cloned my voice across five languages with ElevenLabs. This one adds the face: a HeyGen digital twin built from a single studio recording, speaking languages I never recorded on camera. Most comparison videos use audio from my ElevenLabs clones; one demo uses my voice cloned with HeyGen’s own engine, and the Video Translate versions use HeyGen-generated dubbing. Every video below is AI-generated; none was filmed in the language you hear. The lip sync is AI-generated too: accent and fluency are where the seams show, so listen for them.

Seven languages in total: five spoken by the cloned voices (English, Spanish, Italian, Japanese, Chinese) and five produced by Video Translate (Spanish, Italian, French, German, Japanese), overlapping in the middle.

The first two plans

FreeCreator
Price (as of July 2026)USD 0USD 29/month (USD 24/month billed annually)
Videos3 per month, 1 minute each, watermarked600 monthly credits
ExportShareable 720p1080p download
Digital twin slots11+ (expandable)
Premium featuresTrial accessAvatar IV, translate, looks

Credits are feature-specific: Avatar IV, Video Translate, Precision mode and generated looks consume at different rates. The Free plan is enough to validate whether your twin looks and moves like you; I used Creator for the 1080p watermark-free exports and the paid generation features shown here.

Building the digital twin

The process is simpler than it sounds: record yourself talking to the camera for at least 2 minutes in one continuous take (HeyGen recommends up to 5 for best training), upload the footage, and record a short consent video. HeyGen quotes roughly 10 to 20 minutes to process the base avatar (non-4K footage); actual timing varies with queue and footage. The result is an avatar that models your expressions, body language and delivery style, and can be driven by supported scripts or audio, subject to moderation.

This is what HeyGen generated from my recording:

The digital twin HeyGen generated from my footage

And this is the twin speaking with my voice cloned by HeyGen’s own voice engine, before switching to my ElevenLabs clones. Three cloning engines end up compared on this page: HeyGen’s, and ElevenLabs’ two tiers:

Digital twin · my voice cloned with HeyGen's engine · English

Looks: generating environments

One avatar slot holds up to 500 looks, photo and video combined: generated variations of outfit, background and framing, so the same twin can present from an office, a studio or anywhere else without re-filming. Background removal depends on the avatar type and plan. The videos below use these generated looks; none of the environments exists physically.

The voice pairing

The twin’s voice is not HeyGen’s: the IVC and PVC comparison videos below were generated in HeyGen using audio produced with my ElevenLabs clones from the previous experiment, recorded with a Shure SM58. Same script per language; the main variable is the cloning method, Instant Voice Cloning (IVC) versus Professional Voice Cloning (PVC), though the final output also depends on synthesis and lip-sync settings. Chinese has a PVC take only, so treat it as a sample rather than a pair.

English

Instant Voice Cloning (IVC) · English
Professional Voice Cloning (PVC) · English

Spanish

Instant Voice Cloning (IVC) · Spanish
Professional Voice Cloning (PVC) · Spanish

Italian

Instant Voice Cloning (IVC) · Italian
Professional Voice Cloning (PVC) · Italian

Japanese

Instant Voice Cloning (IVC) · Japanese
Professional Voice Cloning (PVC) · Japanese

Chinese

Professional Voice Cloning (PVC) · Chinese

HeyGen Video Translate

The last piece does not require you to create and manage a separate voice clone. I gave HeyGen Translate the English video and asked for five languages: it transcribes the source audio, translates it, generates the dubbed voice and re-syncs the lips. I enabled burned-in captions for the versions shown here. One recording in, five localized videos out. Use the buttons to switch language; the first one is the original.

Translate source · English (also AI-generated: my twin with the IVC voice)

Takeaway

Two minutes of well-lit, continuous footage produced a twin that survives languages I have never spoken on camera. In my tests, Instant Voice Cloning was enough to validate the idea; Professional Voice Cloning is what I would ship. For pure localization, Video Translate was the shortcut: no separate clone to manage, captions on demand. None of it is flawless, accent and rhythm still betray the generation in places, but the production stack for a multilingual talking-head video has collapsed into one source recording, a set of reusable voice models, and a queue.