ElevenLabs has released a new generation of its text-to-speech models. Eleven v4 and a low-latency variant, Eleven v4 Turbo, were announced on 28 September. For anyone who buys AI voice tools for narration, video voiceovers, dubbing or phone agents, it is the most significant product change in the category in the last two weeks. It is also a launch where most of the headline numbers come from the seller, so it is worth separating what is documented from what is promotional.
What ElevenLabs says has changed
In its launch post, ElevenLabs describes v4 as built on an entirely new architecture and as its most emotive text-to-speech model yet. The main changes it lists are:
- More control over delivery. Users can add inline tags such as laughs, accents or background sounds, and the company says v4 follows these directions more accurately than earlier models. It also says support for IPA phonemes, used to set custom pronunciations, has improved.
- Multi-speaker dialogue. The model is meant to treat a whole scene as context, so speakers respond to what was just said instead of reading isolated lines.
- 90+ languages. Both models support more than 90 languages, and a voice recorded in one language can speak others while keeping its identity.
- Voice cloning. The post says Instant Voice Clones can capture a voice with high fidelity from just 10 seconds of audio, and that Professional Voice Clones are now supported.
- A Turbo model for agents. Eleven v4 Turbo is aimed at real-time voice agents. ElevenLabs gives a median time to first speech of about 150 milliseconds.
Both models are, according to the company, available now in its agents platform, its creative apps and through its API.
TechCrunch's report on the launch gives similar details. It adds that the previous version supported 70 languages and that ElevenLabs said it saw the biggest quality jump in Japanese, Brazilian Portuguese, Mandarin and Cantonese. It also notes that rivals including Cartesia, Deepgram and WellSaid Labs, as well as Google and OpenAI, have been improving their own voice models.
Treat the benchmarks as claims
The launch post says v4 is ranked number one by Artificial Analysis and preferred by roughly 75% of listeners in blind head-to-head tests. The footnotes show where those figures come from. The preference test was run against four named competing models in September, and the latency comparison was measured against another named set. Those are the company's own comparisons, published on its own page, and we have not seen the underlying data. They are a reasonable reason to try the model, not proof that it will sound better on your script, in your accent or for your audience.
The page also gives two latency figures, about 100 milliseconds median inference latency and about 150 milliseconds median time to first speech. They measure different things, so compare like with like if you are comparing vendors for a phone or support use.
The licence and plan details matter more than the model name
The ElevenLabs pricing page is where the practical differences sit. At the time of writing, the free plan lists 10,000 credits a month. The v4 page describes the free plan as personal use with no credit card required, and says paid plans start from $6 a month. The pricing page lists a commercial licence and Instant Voice Cloning from the Starter tier upwards, with Professional Voice Cloning from the Creator tier.
If you make content for a business, a client or a monetised channel, that split is the thing to check before you build on v4 output. A free account is fine for trying the voices. It is not a licence for commercial work. Our existing guides on AI content rights and on choosing voice tools for business cover that in more depth.
Two promotions, worded differently
ElevenLabs is also running launch offers. At the time of writing, a banner on the launch post says 3x credits are included on Creator and above until 12 October. The pricing page, meanwhile, describes a v4 trial in which up to 2x of your monthly text-to-speech credits used on v4 will not count against your balance, for two weeks, on Creator plans and above, in the web and mobile apps only. The pricing page also shows a separate first-month offer on Starter that runs until 18 October.
These may be describing the same promotion in different ways, or different ones. We could not confirm which from the pages themselves. The sensible reading is to check the exact terms inside your account before you rely on extra credits, and to remember that anything that ends on 12 October ends tomorrow.
Voice cloning from ten seconds: use it carefully
A 10-second clone is convenient for creators who want a consistent voice for their own channel. It also lowers the bar for cloning someone else. The launch post stresses fidelity and does not set out consent rules, so read the provider's terms on voice rights before you upload anyone else's recording. Use your own voice, or the voice of someone who has agreed in writing, and keep the agreement. For legal questions about voice rights in your country, speak to a qualified professional.
What to do
- Trial before you pay. Run the same short script, in your own language and tone, through v4 and your current tool. Judge the result yourself rather than relying on a leaderboard.
- Check the licence. Confirm that your plan covers the way you will use the audio, and keep a note of the plan and date you generated it on.
- Read promo terms in your account. The 2x and 3x wording differs between pages, and v4 usage may consume credits differently from older models after the offer ends.
- Test agents under real conditions. If you want v4 Turbo for phone or support, test it on your own network and your own call flow, not just on published latency figures.
- Do not rebuild a working setup for novelty. If your current voice tool already meets your needs, there is no urgency to switch, and the free tier is enough to evaluate v4.
The launch is a real step for expressive, multilingual synthetic speech. As with any launch, the useful questions for a buyer are what you are allowed to do with the output, what it costs once the introductory offers end, and how it sounds on your own material.