Orply.

Eleven v4 Adds Inline Direction for More Expressive Speech

ElevenLabsMonday, September 28, 20263 min read

ElevenLabs presents v4 as a text-to-speech model that makes delivery part of the prompt: inline tags can direct emotion, pacing, reactions and style, shaping how the words are performed. The company says the model is also designed to preserve speaker identity across long-form narration, dialogue and regenerated lines. For live applications such as agents, ElevenLabs positions v4 Turbo as the low-latency option, with a reported median inference latency of about 100 milliseconds.

V4 makes the prompt part of the performance

ElevenLabs presents v4 as a text-to-speech model designed to interpret more than the words on a page. The company says its new architecture uses cues for tone, pacing, emotion, character and context to shape a spoken performance—and that creators can direct those choices with inline tags.

In the launch demonstration, tags such as “[nervous],” “[whispering nervously]” and “[voice breaking]” guide a scene from backstage chatter to a tearful monologue and then to banter. Other cues request breaths, laughter and a clapperboard snap. ElevenLabs says the audio was generated from the displayed prompts without edits or modifications.

A second example makes the control surface clearer: tags ask for a cheerful announcement, rising tension, hesitation and sarcasm in a short report about a pigeon setting off a motion sensor. The words carry the story; the tags specify how it should be performed. ElevenLabs says creators can use inline controls for delivery, emotion, pacing, reactions, sound effects and style.

That distinction matters for production: directing delivery through the prompt gives creators a way to shape a read without relying on the text alone. The company describes v4 as a model that “doesn’t just speak, it performs.” This is its characterization of the output, not a guarantee that every prompt will produce a particular emotional effect.

Consistency and cloning address different production needs

ElevenLabs says v4 can preserve speaker identity across long-form narration, dialogue and regenerated lines, including across “infinite generations.” The practical aim is to let creators revise or extend material while retaining a consistent voice. That promise is distinct from expressive control: the delivery can change with the text while the speaker identity is meant to hold.

The launch description distinguishes two cloning options. Instant Voice Clones can be made from 10 seconds of audio. Professional Voice Clones are positioned as the highest-fidelity option; in the video, ElevenLabs says they offer significantly better speaker similarity and frames cloning as embodying a voice, not merely imitating it. The accompanying examples name tone, texture and intent as qualities carried by a voice.

These options emphasize different tradeoffs in the company’s description: Instant Voice Clones require little source audio, while Professional Voice Clones prioritize fidelity. The broader v4 consistency claim concerns what happens across continued work and regenerated lines. ElevenLabs also says v4 improves multi-speaker dialogue and adds pronunciation control through IPA, addressing how exchanges sound and how particular words are spoken.

Turbo is the low-latency option, not a separate performance brief

ElevenLabs describes v4 as the expressive text-to-speech model and v4 Turbo as bringing the same research to low-latency settings, including agents and other live applications. The stated distinction is therefore about use: v4’s launch materials emphasize expressive direction and sustained voice identity, while Turbo is presented for situations where generated speech must respond quickly.

~100 ms
median inference latency for Eleven v4 Turbo

A staged phone call illustrates the kind of interaction the company has in mind. A caller asks about a medication authorization, then requests the information in Mandarin; the other speaker responds in Mandarin. The example demonstrates a language change within a single exchange, rather than establishing anything about the accuracy of the medical information.

ElevenLabs says performance has improved across more than 90 languages, with Cantonese, Mongolian and Odia among the newly supported languages. The launch description says the models are available in ElevenCreative, ElevenAgents and ElevenAPI. In practical terms, the release pairs expressive controls and voice-consistency claims with a Turbo option for live settings where latency is a central constraint.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free