ElevenLabs Launches Realtime Voice Model With Audio Tags in 70+ Languages
ElevenLabs says its generally available Eleven v3 Conversational model is built for realtime voice applications where delivery is part of the response, not merely a rendering of text. The company pitches audio tags for fine-grained expressive control, consistent streaming quality, and support for more than 70 languages, positioning the model for support agents, assistants and interactive characters that need to convey reassurance, urgency or tension as an exchange unfolds. ElevenLabs offers it through ElevenAgents and its API from $0.05 per 1,000 characters, with lower rates at scale.

ElevenLabs is selling voice as part of the interaction, not a final rendering step
ElevenLabs positions Eleven v3 Conversational for realtime voice experiences in which a response must sound appropriate to the situation as it arrives. The model’s pitch is expressive delivery under conversational conditions: a support agent responding to a delayed order, an assistant reacting to a cooking success, or a fictional character sustaining a tense scene.
The product is aimed at builders for whom intelligible text-to-speech alone is insufficient. In the exchanges shown, the words carry the task, but the intended value is in how the speech frames that task: sympathy before an update, encouragement paired with an urgent instruction, and fear or uncertainty that keeps a scene credible.
ElevenLabs describes Eleven v3 Conversational as its most expressive model for realtime speech. It says the model is optimized for realtime use, maintains consistent quality across streaming generations, and offers audio tags for fine-grained control over delivery. The demonstrations do not test those claims against production deployments; the source identifies them as exchanges generated by the model.
The generated exchanges make expression a functional requirement
The three generated demonstration exchanges treat vocal expression as part of the application’s job, rather than a layer added after the response has been written.
In the support scenario, a customer says a delivery has not arrived after three weeks. The agent apologizes, checks the status, and reports that the order has cleared customs and should arrive tomorrow. The exchange pairs an operational answer with reassurance: frustration is acknowledged without postponing the useful information. The agent opens with, “Oh wow, I’m so sorry to hear you’re having trouble with your order. I’m taking a look now.”
The cooking exchange turns on pace. After one speaker reports that a soufflé finally rose, the assistant celebrates—“Hey, great job!”—then immediately tells them to take it to the table because it will hold that height for about two minutes. Excitement is not separate from the task. It gives way quickly to a time-sensitive instruction, illustrating a response whose delivery needs to recognize success while directing the next action.
The character exchange asks for something else again. A door opens on its own, one voice hopes it was the wind, and another dismisses that possibility: “In this house? Huh, not a chance.” The final startled question—“Wait! What is that?”—leaves the scene in escalation rather than resolution. Here, the useful output is not a helpful answer or instruction, but a controlled emotional register that sustains the premise of a fictional scene.
Together, the exchanges distinguish three different demands on a voice layer. Support requires warmth alongside concrete information. Coaching requires energy without sacrificing clarity or urgency. Character performance requires dialogue that can carry suspicion, alarm, and timing rather than merely read a script aloud.
The commercial proposition combines realtime delivery with directed affect
For teams assessing Eleven v3 Conversational, the first consideration is whether speech needs to be generated while an interaction unfolds. ElevenLabs frames the model around streaming, realtime use cases rather than speech that can be fully prepared before a listener hears it. That distinction matters most in the kinds of interactions shown: a customer awaiting an answer, an assistant responding to a live event, or a character reacting within a scene.
The second consideration is whether the application needs directed affect. ElevenLabs identifies audio tags as its mechanism for fine-grained control, while the demonstrations indicate the delivery range that control is meant to support: empathy in support, excitement followed by urgency in assistance, and suspicion or alarm in an interactive character. A stable, neutral reading would not serve those situations in the same way.
Language coverage is the third part of the proposition. The company says the model supports more than 70 languages and specifically identifies German, Spanish, French, Portuguese, and Hindi as languages in which it excels. The claim is aimed at multilingual voice products, where the desired performance is not simply intelligible output in another language but expressive delivery across deployments.
The stated commercial model is character-volume pricing. ElevenLabs lists Eleven v3 Conversational from $0.05 per 1,000 characters, with lower pricing at greater scale, and makes it available through ElevenAgents and the ElevenAPI.