Generative voice AI with ElevenLabs: one voice, five languages
My voice was cloned from Spanish recordings only. Everything else on this page is generated: the same clone speaking English, Italian, Chinese, and Japanese, languages it never heard me speak. This is my controlled listen to where generative voice AI stands today, using ElevenLabs, the current reference provider, as the test bench.
The clone was trained on my Spanish audio only: seconds of it for Instant Voice Cloning (IVC), around 30 minutes for Professional Voice Cloning (PVC). One script, identical generation settings, one variable: the training method. For comparison, the English and Spanish sections open with a real recording of my voice.
Test configuration
| Parameter | Value |
|---|---|
| Speed | 1.00 |
| Stability | 45 |
| Similarity | 80 |
| Speaker Boost | Enabled |
Every sample below uses this exact configuration, so the training tier is the single variable.
Two training tiers
| Tier | Training audio | Intended use | Plan (as of July 2026) |
|---|---|---|---|
| Instant Voice Cloning (IVC) | ~10 seconds | Prototyping | USD 6/mo |
| Professional Voice Cloning (PVC) | ~30 minutes | Production | USD 22/mo |
English
Spanish
Italian
Chinese
Japanese
How to listen
Compare each IVC/PVC pair on: identity preservation, pronunciation, intonation and rhythm, stability, naturalness. Start with the language you know best, where flaws are easiest to hear, then check whether the same flaws appear in the others.
Takeaway
Cross-lingual identity transfer works today: I can hear my own voice in languages I have never spoken, cloned from Spanish alone, from either tier. What more training data buys is consistency: steadier delivery across languages and longer passages. Seconds of audio validate an idea; minutes justify production. The limit is no longer the language: it is the consistency your use case demands.