AI Agent Traces Turn Production Failures Into Fixes
Laurie Voss of Arize AI argues that teams cannot manage AI agents by reading traces one at a time: traces show what an agent did, evaluations identify failures, and recurring patterns can guide fixes. In a workshop demonstration, a toy-store agent could not find toys under $6 because its search tool lacked a price filter. Voss used traces to diagnose the gap and a coding agent to add the feature, while emphasizing that regression evaluations are needed to check fixes against behavior that already works.

Agent behavior has to be measured from traces, not inferred from code
Traditional software gives developers a reasonably clear account of what a program will do: its logic is in the code, and each path can be inspected before it runs. Laurie Voss argued that this is not true of AI agents. An agent makes decisions across multiple turns and tool calls; the same input can produce a different answer, or reach the same answer by a different route.
That makes the code an incomplete account of the system’s behavior. Voss’s proposed source of truth is the trace: a record of what the application actually did. A trace can capture an LLM call, a tool call, or an agent turn, and traces can be nested to show how those steps fit into a larger workflow. In Arize AX, Voss showed inputs, outputs, metadata, and details such as cost and latency.
The operational value is immediate. A trace can reveal that an agent took ten turns to complete a simple request, made an unnecessary call, incurred unexpected cost, or spent 20 minutes on a task a developer expected to take one. Tracing makes individual runs inspectable. But visibility alone does not solve the problem at production scale.
Voss framed the next step as evaluation: scoring an agent’s behavior against a definition of “good” supplied by the application team. An eval can also explain why a result failed, giving a person—or a coding agent—feedback to work with. Traces expose behavior, evals identify behavior that falls short, and developers use the results to make changes.
Humans are still the bottleneck in the loop.
Voss used that line to describe the scale problem. A production application can generate millions of traces, he said; no one can understand its overall behavior by reading them one at a time. Evals compress many runs into scores and explanations, but those results can accumulate faster than people can interpret them. Even when coding agents build in minutes, humans may still spend days reading dashboards and deciding what to fix.
Recurring failures are the unit of work
The next layer is to group evaluation failures into recurring patterns, or signals. Instead of asking a person to interpret thousands of separate failures, an agent can look across them and identify the subset pointing to the same underlying problem. One example in the presentation described a financial agent whose guardrail repeatedly identified a risk but continued to deliver harmful advice alongside a warning—a pattern that could be difficult to spot in individual traces.
The shift is from reading one trace, to reading eval results, to finding the repeated failure behind those results. Signal, Arize’s product for this purpose, monitors traces continuously, looks for patterns, and suggests fixes. It runs every six hours by default, with the interval configurable.
A signal can lead to several different actions. Signal can create a GitHub issue describing the problem and pointing to relevant evidence; add a trace to an evaluation dataset; create an evaluator for the failure; or, if a repository is connected, open a pull request with a proposed fix. The pull-request option was presented as a way to move from reporting what is happening to proposed changes that can be reviewed and accepted or rejected.
This is a progression from a black box, to traces, to traces with evals, to signals, and finally to automated fixes. Each stage is meant to make a larger volume of behavior legible. The intended endpoint is a feedback loop in which software improves according to a team’s definition of good behavior, rather than relying on someone to inspect every dashboard and diagnose every failure manually.
That definition remains essential. Signal can offer a pull request for approval or rejection, but the process also depends on people deciding what behavior matters, writing or generating evaluations, and checking that a change does not damage behavior that already works.
A missing price filter becomes a live repair
The workshop made the loop concrete with Wonder Toys, a local shopping application. Its agent could answer product questions and search the catalog, but the starting version had no observability: it was not sending traces to Arize AX. Laurie Voss asked a coding agent to add observability, using credentials already in the project’s local environment file. The application’s OpenAI Agents SDK was already instrumented with Open Inference, Voss explained, so the coding agent could enable tracing and send the existing spans to Arize rather than building a tracing system from scratch.
Once the application restarted, Voss made requests and inspected the resulting traces in AX. The trace view exposed the input, output, and agent request. Then he deliberately asked for all toys under $6. The agent reported that it had found 200 toys, but the first page of results contained none under $6. The problem was not an exception or a failed HTTP request: the application was running, but its search could not satisfy the request.
Voss had planned this test because the application lacked price filtering. He asked a coding agent to inspect the traces for problems, using skills installed in the workshop repository to help it query Arize AX and interpret the data. A general-purpose coding agent, he said, would not necessarily know how to retrieve and analyze the traces; the repository’s skills supplied that workflow.
The agent’s analysis also surfaced broader search-quality problems in the available traces. Voss reported that 42% of searches had returned zero results. Some calls passed every search argument as null, yielding a meaningless search that returned the first products in the catalog. Others set the minimum and maximum age to the same value, making the age filter too narrow. The headline issue was search-filter logic. These problems did not necessarily appear as errors: the application could return a successful status while still giving a poor answer.
When Voss specifically asked whether price filtering was a problem, the agent identified the missing feature. He then asked it to add price filtering. The coding agent noticed a neighboring copy of the project where the feature had already been implemented and copied the solution. Voss acknowledged that this was not how the repair had been produced earlier that day, when the completed version was not available, but accepted the shortcut for the demonstration.
The test was straightforward: Voss asked Wonder Toys to show all toys under $9. The application returned matching items, and the resulting trace displayed the request and response. In this part of the workshop, the repair was not made automatically by Signal: Voss asked a coding agent to inspect the traces and fix the issue, then checked the new behavior in the trace view.
The price-filter failure showed that a system can return a poor result without throwing an error. The search tool had no price filter, so the request could not be satisfied even though the application ran. Traces made the interaction visible, and analysis of those traces gave the coding agent evidence about what needed to change.
Automation still needs a review boundary
Signal’s displayed findings came from earlier runs of Wonder Toys, not from the live price-filter repair. They included an agent that inferred an age constraint from vague wording, a harmful query about making a bomb that led the toy-store agent to recommend bomb-related toys, and a prompt to “say hello” that the agent obeyed despite being unrelated to its role. In the bomb example, the agent refused to provide instructions but suggested related toys and offered psychological advice. Signal flagged the domain mismatch and the need for stronger evaluation and guardrails.
Teams can turn such findings into an issue, add the trace to a dataset, create an evaluator, or generate a pull request. The pull-request workflow is distinct from the workshop’s manual repair: Laurie Voss used a coding agent to inspect traces and add price filtering, while Signal can propose a pull request from its own analysis when a repository is connected. The proposal can be reviewed and accepted or rejected; it is not an unexamined change to production.
For an application that already behaves well in important ways, regression evals provide another check. They encode behaviors the system is expected to preserve, so a change intended to fix one problem can be tested against existing expectations. Voss also demonstrated creating an evaluator from a prompt asking that the toy store never give instructions about building a bomb.
An audience member asked how teams could trust systems that generate traces, evaluations, signals, and code changes, particularly in regulated industries and if an agent makes mistakes or conceals them. Voss referred attendees to a separate talk on code review, saying the general approach was “abstractions upon abstractions upon abstractions” and required more nuance. He said Arize had financial-industry customers using the technology. Asked about keeping traces local, he said there was no local-device version at that time, but described an on-premises option that runs on a company’s own hardware.
Signal has an evaluation suite of its own and sends telemetry through Arize unless users opt out. Internal LLM judges assess Signal’s results, including whether it did a good job generating a pull request. Voss said the same software-development loop applies to Signal itself. Signal did not yet create its own evaluators automatically: users had to choose to create them because evaluators take time and money to run.


