Skip to content

Mariano Rodrigo

AI Solutions Engineer building production systems with artificial intelligence, automation, and full-stack architecture. This is my public engineering lab: architecture decisions, implementation reports, experiments, and lessons from real systems.

Generative voice AI with ElevenLabs: one voice, five languages

The experiment in 3 minutes

My voice was cloned from Spanish recordings only. Everything else on this page is generated: the same clone speaking English, Italian, Chinese, and Japanese, languages it never heard me speak. This is my controlled listen to where generative voice AI stands today, using ElevenLabs, the current reference provider, as the test bench.

The clone was trained on my Spanish audio only: seconds of it for Instant Voice Cloning (IVC), around 30 minutes for Professional Voice Cloning (PVC). One script, identical generation settings, one variable: the training method. For comparison, the English and Spanish sections open with a real recording of my voice.

Test configuration

ParameterValue
Speed1.00
Stability45
Similarity80
Speaker BoostEnabled

Every sample below uses this exact configuration, so the training tier is the single variable.

Two training tiers

TierTraining audioIntended usePlan (as of July 2026)
Instant Voice Cloning (IVC)~10 secondsPrototypingUSD 6/mo
Professional Voice Cloning (PVC)~30 minutesProductionUSD 22/mo

English

Original · English (my real voice, reference)
Instant Voice Cloning (IVC) · English
Professional Voice Cloning (PVC) · English

Spanish

Original · Spanish (my real voice, training source)
Instant Voice Cloning (IVC) · Spanish
Professional Voice Cloning (PVC) · Spanish

Italian

Instant Voice Cloning (IVC) · Italian
Professional Voice Cloning (PVC) · Italian

Chinese

Instant Voice Cloning (IVC) · Chinese
Professional Voice Cloning (PVC) · Chinese

Japanese

Instant Voice Cloning (IVC) · Japanese
Professional Voice Cloning (PVC) · Japanese

How to listen

Compare each IVC/PVC pair on: identity preservation, pronunciation, intonation and rhythm, stability, naturalness. Start with the language you know best, where flaws are easiest to hear, then check whether the same flaws appear in the others.

Takeaway

Cross-lingual identity transfer works today: I can hear my own voice in languages I have never spoken, cloned from Spanish alone, from either tier. What more training data buys is consistency: steadier delivery across languages and longer passages. Seconds of audio validate an idea; minutes justify production. The limit is no longer the language: it is the consistency your use case demands.