ElevenLabs Adds Agent Monitoring and Context-Preserving Voice Tools
ElevenLabs’ monthly changelog argues that its latest releases are less about adding generation models than about making AI voice workflows more controllable in production. The company has introduced Spotlight to identify recurring problems in live agent conversations, Character Casting and consistent vocals to standardize creative choices, and API updates intended to preserve performance in dubbing and structure in real-time transcription.

The releases matter most where generation meets operating reality
ElevenLabs is adding controls around generation that address two different points of failure: production teams needing to make repeatable creative decisions before they publish, and operators needing to detect failure after an agent is deployed. The most consequential changes are therefore not simply new models or endpoints. They are mechanisms for preserving context, standardizing choices, and finding problems in large volumes of live interaction.
For teams running customer-facing agents, the central constraint is review capacity. ElevenAgents’ new Spotlight layer reads voice and chat conversations as they happen, groups them into topics, scores each conversation against criteria an operator writes in plain English, tracks sentiment over time, and identifies what to fix next. The product is designed for the stage at which an agent handles thousands of conversations a week and no one can realistically read every transcript.
That changes the unit of analysis from the individual transcript to a recurring operational pattern. The account-management dashboard shown in the release, for example, separates conversations into tasks including opening an account, discussing account types and features, updating details, and closing an account. It makes resolution rate and sentiment visible alongside volume. “Close an account” has the lowest resolution rate in the example, while “update account details” has the most negative sentiment score. Those are different diagnoses: one points to a task that often remains unresolved; the other indicates a particularly poor experience even when the conversation may reach an outcome.
| Subtopic | Volume | Resolution rate | Sentiment score |
|---|---|---|---|
| Open a new account | 428 | 72% | +0.05 |
| Account type and features | 312 | 84% | +0.22 |
| Update account details | 201 | 58% | -0.26 |
| Close an account | 150 | 49% | -0.05 |
Spotlight’s value, as presented, is earlier detection. Rather than learning about a broken workflow through a customer complaint, an operator can identify a weak category of conversations in the traffic already flowing through the agent.
ElevenLabs says those agents handle more than 10 million conversations each week. It has also expanded the channels through which an agent can operate: SMS, Telegram, Intercom, and Freshdesk join phone, web, Zendesk, Slack, and WhatsApp. The total is nine channels, with behavior now tunable by channel. The premise is practical rather than cosmetic: an agent should not answer a phone call in the same way it answers a text message.
Performance and production decisions are being kept closer to the original work
The creative releases address a different kind of loss of context. In audio production, a manuscript is not merely text to be narrated; it has characters, dialogue, names, invented terms, and choices that need to remain consistent. In music, repeated generation is less useful if the singer changes each time. In dubbing, a translation can be technically correct while losing the performance that made the source work.
Character Casting for audiobooks turns a manuscript into an initial casting workflow. After a user uploads a book, the system identifies its characters and proposes a voice for each. Users can hear those voices reading actual dialogue from the book before committing, rather than selecting a voice without hearing it in context. The tool also builds a pronunciation dictionary, pre-filling names, place names, and invented words.
That is aimed at reducing the setup work that can slow audiobook creation after the manuscript is already complete. The interface shown in the release offers single-cast narration or a multi-cast arrangement with distinct voices for the narrator and each character. Casting, dialogue previewing, and pronunciation handling are all placed in the same workflow.
ElevenMusic applies the same concern with continuity to generated songs. Its new vocals feature lets users generate original music with a consistent voice, either their own or one selected from the Vocal Library. The stated purpose is to keep the same singer across a track and across multiple generations, rather than producing a different voice each time a user regenerates material. Fine-tuned voices work with Music v2 through the API as well.
References add a separate control over style. A user can upload a track from 10 seconds to five minutes long, and Music v2 will generate something intended to match its style and feel. ElevenLabs says uploads receive a copyright check first, intended to help users avoid referencing material they do not have the rights to use.
The same distinction between literal content and expressive delivery underlies Dubbing v2. The system is now available through ElevenAPI, allowing developers to create dubbing projects from an audio or video file or a public URL, and to dub content into more than 90 languages. Previously, the company says, the model was available only within ElevenCreative.
ElevenLabs’ key claim is that Dubbing v2 conditions on the original performance rather than a transcript. In its account, that allows tone and emotion from the original delivery to carry into the new language rather than being flattened into a straightforward translation. For developers embedding dubbing in their own workflows, that is the substantive distinction: the API is intended to carry performance information, not only source-language words.
The API includes Create, List, Fetch, and Delete endpoints. Editing and regeneration are available on the Enterprise plan. Dubbing v1 remains available at the same price for teams already using it.
The API changes make live speech more structured and more controllable
Scribe v2 Realtime extends the same effort to preserve useful context in a live stream. It can identify entities as they are spoken, so names, places, dates, and numbers emerge as tagged elements instead of being buried in unstructured transcription. It also supports secondary languages within the same stream.
For Enterprise teams, Scribe adds a zero-retention logging switch, allowing a session to run without retaining logs. The update is available in the latest SDKs.
Taken together with Dubbing v2, the developer releases are directed at workflows where raw transcription or literal translation is not enough. A real-time system may need to isolate a person’s name or a date as an entity; a dubbed video may need to retain the source speaker’s delivery. ElevenLabs is presenting both as ways to make voice infrastructure more usable inside products rather than only inside its own creative interface.
A short set of company and community notices
ElevenLabs says it has opened its first Canadian office in Toronto and plans to double the local team this year. The careers page shown in the release lists openings in engineering and product.
The company also announced upcoming ElevenLabs Summits in Bengaluru on October 6 and New York on November 11. Its Summer of Sound music promotion runs until September 1: free users receive up to 400 generations a month, and a $50,000 prize pool is divided across the 10 most-streamed tracks.