Orply.

Eleven v4 Turbo Brings ~100-Millisecond Voice Responses to ElevenAgents

ElevenLabsTuesday, September 29, 20264 min read

ElevenLabs says its v4 Turbo voice model, now available in ElevenAgents, combines a median inference latency of about 100 milliseconds with a wider emotional range and improved reliability on long calls. The company argues that those capabilities can help agents respond to sensitive customer needs with practical next steps, not just sympathetic language, while supporting conversations in more than 90 languages. Its demonstrations show agents offering a service-outage workaround, gathering details about a damaged order and presenting options for a parking citation.

Low latency matters when empathy has to lead somewhere

ElevenLabs’ pitch for v4 Turbo in ElevenAgents is that a voice agent should do more than sound sympathetic: it should recognize a caller’s immediate concern and offer a useful next step. The company says the model combines low latency with a wider emotional range and improved reliability over long calls. Its examples put that combination to work in situations where callers need help with a bill, a damaged order, a service outage, or a city citation.

~100 ms
median inference latency claimed for Eleven v4 Turbo

In a healthcare exchange, the agent offers to check insurance while looking up a lab schedule. When the caller says they have lost their job and cannot afford a surprise bill, an on-screen status says the agent detected a somber tone. The agent responds, “I’m so sorry to hear that. We’ll figure this out for you.” The exchange links emotional recognition to a practical constraint, though it does not show the insurance check being completed.

The clearest example of a next step addressing an immediate need comes from a service outage. The agent says a line fault cannot be fixed remotely and offers a technician appointment for the following morning. The caller objects that they have afternoon meetings and cannot be offline that long. An overlay marks the caller’s frustration and identifies a mobile-hotspot offer as the proposed solution. The agent acknowledges that it cannot send a technician sooner, then offers a free hotspot boost so the caller can tether a laptop until the appointment.

The underlying fault remains unresolved, but the offer addresses the caller’s immediate problem: staying online. That is the product’s distinctive proposition in practice—not empathy as a reassuring phrase, but a response that adapts to the caller’s circumstances when the requested fix is unavailable.

The useful response changes with the caller’s problem

The other exchanges show how that approach can take different forms. A retail customer reports a tear in a sweater delivered the previous day. The agent apologizes and asks the customer to upload a photo; an overlay identifies the order as a textured knit sweater and says the agent detected distress. Here, the next step is gathering information to assess the complaint, rather than promising a return or replacement.

In the banking example, the agent surfaces recent transactions: a $6.40 coffee-shop charge, an $18.20 rideshare charge, and a $554 charge from GPRX Digital Services. The caller says the last one is unfamiliar. The exchange shows the agent bringing the relevant charge into the conversation, but does not proceed to investigate or resolve the dispute.

A city-government example combines a calm response with a menu of options. A caller who is new to the city asks whether a parking ticket will affect their permanent driving record. The agent says a parking citation is typically not a moving violation and lists three choices: pay now, set up a payment plan, or contest the citation if it was issued in error. An overlay identifies a Parking Violations subagent and describes the response as calm and factual.

Across these cases, the action depends on the request: ask for a photo, identify a charge, offer a workaround, or lay out choices. The product claims concern the model’s latency, emotional range, language support, and reliability; the demonstrations illustrate the kinds of responses ElevenLabs intends those capabilities to support.

Turbo is presented as one part of the agent stack

ElevenLabs says v4 Turbo supports more than 90 languages and that global brands can reach customers in their own languages with native-level accuracy. In ElevenAgents, the model works alongside the company’s transcription and turn-taking models. ElevenLabs describes these components as a co-optimized stack intended to help agents respond faster, interrupt gracefully, and improve automatically as new models launch.

That system-level framing matters because a live agent must do more than produce speech: it must handle what a caller says, manage conversational turns, and respond with relevant information or options. The demonstrations offer examples of that broader task, from checking an order or identifying a transaction to providing appointment information and explaining citation options. They do not establish how every interaction concludes, but they make the intended role of the voice model clear: support a conversation that moves quickly while adapting its response to the caller’s concern.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free