A Financial Agent’s 13 Rejected Reports Exposed a Judge Without Research Context
Laurie Voss of Arize AI argues that reliable AI agents require more than trying a few prompts and deciding the results look good: teams need to inspect traces of what agents actually did and define success before automating evaluation. A financial-analysis agent in her workshop produces reports that expose the limits of a judge grading from its own knowledge: it rejects all 13, while a faithfulness check given the agent’s research context finds six faithful and seven unfaithful. Voss uses focused evaluators, human comparisons and repeated tests on saved failures to guide changes before and after deployment.

A polished report can conceal a failed tool call
A report can look finished and still be useless. In Laurie Voss’s financial-analysis demo, the agent searches the web, writes a report, and may tell the user it saved the report to a file. But the write fails in Colab, and the agent’s final answer can still say the report was saved. The report itself is missing.
That failure is hard to see if you inspect only the prompt and final response. A trace exposes the steps in between: the agent’s research, its tool calls, the attempted write, and the output it returned. In another run, the agent searches repeatedly for Microsoft and gets stuck in a loop. In a Rivian report, Voss finds confident financial claims that the workshop has not yet checked against source material. These are different failures, and the final answer alone does not tell you which occurred.
The agent has two turns. First it researches a ticker and focus area using web search; then it writes a concise report based on that research. The agent decides how many searches to make and when it has enough information. Even with the same input, it may search for different things, find different results, and produce a different report. Some versions may be good; others may fail.
Laurie Voss frames the basic distinction this way: traces are logs for AI, and evals are tests for AI. A trace records what happened at runtime. Its component spans represent steps such as an LLM call or tool invocation, and can include inputs, outputs, timing, token counts, and other metadata. An eval uses recorded data to judge whether the result met a defined standard. A trace answers how the agent got there; an eval asks whether what it did was any good.
In the workshop, Voss instruments the Claude Agent SDK with OpenTelemetry and OpenInference, then sends those records to Arize AX. The setup uses the Arize registration function with a space ID, API key, and project name, followed by the Claude SDK instrumentor. The registration sets up OpenTelemetry; OpenInference adds LLM-specific details such as prompt and completion text, token counts, model, and tool information. The SDK instrumentation captures LLM and tool calls without requiring tracing code around every operation. In the notebook, Voss sets batching off so telemetry is sent as it happens rather than waiting to batch it. A top-level span groups the overall financial-report operation; an Arize client can later read stored spans back for evaluation.
The trace tree shows the overall financial-report operation and nested spans for the SDK, web searches, and file-write attempt. Clicking a span reveals what the model received and what it returned. Without that view, the run can look like a prompt followed by a report. With it, the team can see the agent’s intermediate actions and locate where the execution went wrong. The input can become an evaluator’s input; the output can become the material being graded.
Agents have more opportunities to go wrong than a single LLM call. Each decision and tool call can introduce an error, and one mistake can contaminate later steps. A search for “Tesla” could return material about the eighteenth-century inventor rather than the car company. The agent might then write a coherent, well-formatted report about the wrong Tesla. Nothing about the prose necessarily signals the retrieval mistake. The user sees a polished answer and may trust it.
A multi-agent system adds handoffs to the same problem: a routing agent can send work to the wrong specialist, or a specialist can mishandle the handoff. Voss says to inspect intermediate steps as well as the final result. Did the agent choose the right tool? Did it pass the right arguments? Did it stop when it had enough? Did it get caught in a loop? Errors can cascade, so evaluation has to match the system’s actual shape.
There is a countervailing risk: an agent may find a valid solution that a rigid test does not anticipate. Voss cites an example from Anthropic’s Tau-squared benchmark. An agent was asked to reschedule an economy-class flight, something the test designers thought policy made impossible. But the agent could first upgrade the ticket to first class, which could be rescheduled. The agent succeeded by a route the benchmark did not expect and was marked as failing.
These cases put a constraint on evaluation design: distinguish an actual failure from an unexpected but valid approach. Voss recommends grading the outcome rather than demanding a particular sequence of steps. A test that checks whether a report mentions the requested stock ticker does not care how the agent found the information. It checks the result that matters. A test that insists the agent call a specific tool first may punish a different route to the same valid answer.
A few successful queries are not a test suite
Teams often ship AI features after trying a few queries and deciding the results look right. Voss calls this the “vibes” approach. A small number of successful examples can create confidence, but running a feature three times is not a test suite, and users will not limit themselves to the cases the team happened to try.
Traditional unit tests cannot simply assert an expected string when the same prompt can yield different wording on different runs—and two differently worded responses may both be correct. Human review can help, but Voss notes that it is slow, does not scale to every run, and does not provide a repeatable quality gate in continuous integration. Without evals, a prompt change that fixes tone might also introduce hallucinated product features. A team may not discover the regression until a user reports it.
Evals make it possible to compare prompt versions, test a model change against a defined set of cases, and look for regressions before release. Voss says that as models change frequently, a reusable suite can turn a model comparison that would otherwise take weeks of manual testing into a result available within hours. The point is not that an eval guarantees quality. It creates a repeatable way to measure and act on it.
Voss separates evaluation methods into three complementary kinds. Code evals are deterministic checks: does the output parse as JSON, stay within a length limit, include a required field, or mention the requested ticker? They are fast, cheap, and reproducible, but can be brittle when they mistake wording differences for quality differences. LLM judges apply a rubric to questions that require semantic interpretation, such as whether a response is faithful to its sources or whether its tone fits a customer-service context. They are more flexible, but cost money and can be wrong or nondeterministic. Human review remains useful for novel failures and for checking whether judges are reliable; it is not a perfect or scalable grader either. Voss notes that human annotators can miss defects when fatigue sets in.
The methods work as layers rather than alternatives. A deterministic check can catch a missing ticker; an LLM judge can assess whether the analysis is grounded; a human can examine an unfamiliar or ambiguous failure. Voss likens this to the Swiss cheese model: each layer has gaps, but several different layers provide more coverage than any one alone.
She also distinguishes capability evals from regression evals. A capability eval asks whether the agent can do something new and is expected to fail at first; it defines a hill to climb. Once the agent can reliably do that thing, the same test can become a regression eval, which the mature system should pass consistently. Over time, capability tests can become part of the regression suite a team runs routinely.
An eval result can include a score, a label, and, for an LLM judge, an explanation. A binary score such as zero or one is easy to compare, while the label makes the result legible—“correct” or “incorrect,” for example. The explanation is more than decoration. If several failing reports share the same explanation, the team may have a systematic issue rather than a collection of unrelated edge cases. Those explanations can inform a prompt revision or coding-agent task, provided the failure has first been inspected and the rubric is trustworthy enough to guide action.
Read the traces before deciding what to grade
The workshop generates 13 runs: an initial Tesla report and 12 test queries spanning companies and tasks. The set includes familiar large companies, a Rivian query with less public information, dividend-yield analysis for Coca-Cola, and multi-ticker comparisons. Different queries help expose differences in phrasing, intent, and complexity; related queries can also reveal how nondeterministic runs vary. Voss’s test set includes, for example, both an Amazon profitability query and a query focused on Amazon Web Services.
Before writing evaluators, Voss’s central instruction is to read actual traces. Most reports look substantial, with summaries, financial metrics, risks, and recommendations. But some outputs are much shorter because the agent says it has written the report to disk. The actual body is absent. Another run repeatedly searches the web. These are not the same defect, and neither is captured by the impression that the report “looks good.”
Voss describes a two-pass way to analyze what traces reveal. First, use open coding: read each example and write down the problem in ordinary language, without deciding on a taxonomy in advance. Notes might say “talked about the wrong ticker,” “made up a number,” or “vague recommendation.” Then use axial coding to group similar observations into broader causes. A wrong answer might result from poor retrieval, a reasoning error, a hallucination, or a scope violation. The cause matters because it points to a different remedy: better search, a prompt change, grounding checks, or clearer boundaries.
It is tempting to start with a familiar label such as “retrieval failure” and fit the evidence into it. Voss argues that the initial pass should resist that impulse: first record what happened, then group the evidence. Otherwise, teams risk measuring what is easy to categorize rather than what is actually wrong.
The workshop’s displayed frequency table is deliberately not real analysis. Voss says she assigned random categories rather than reading and coding every trace in advance. She presents the table to illustrate the method, not as evidence about the agent’s true failure distribution. Even with a real count, frequency alone should not determine priorities. A rare, high-severity failure can matter more than a common, low-severity one. Voss’s prioritization guidance is to consider frequency and severity together.
Requirements are what make the analysis actionable. Before calling a report bad, a team needs to say what a good report must do. For this financial analyst, Voss proposes that it reference the correct ticker, include recent and actionable financial data, make recommendations, and distinguish forward-looking analysis from historical summary. Those criteria turn a vague complaint into something testable.
The people who define those requirements should not be limited to engineers. Product managers, QA staff, support teams, and other colleagues may know what users need and what domain-specific failures look like. Voss presents evaluation as cross-functional work: engineers may implement the checks, but domain knowledge helps establish the standard. In the workshop’s terms, people with relevant domain knowledge should be able to read a trace or annotate whether an output meets the stated bar.
Test data must also reflect more than obvious cases. Before production traffic exists, synthetic queries can provide a starting set. Voss recommends varying phrasing even when intent is the same—for example, asking for Tesla’s financial performance, asking what is happening with Tesla stock, or asking whether Tesla is a buy. A domain expert should review generated examples, since they may cluster around conventional wording and miss how users actually ask questions. Edge cases such as nonexistent tickers, multipart requests, and jailbreak attempts belong in the set as well. Once real traffic arrives, it should increasingly replace imagined examples: the goal is to test what users ask, not what a team wishes they would ask.
Use the cheapest check that answers the question
The first evaluator in the workshop is a simple code check: did the output mention the ticker in the query? The implementation looks for likely uppercase ticker symbols in the input, excludes common acronyms, and checks whether each appears in the report. It does not require a model call. Across the 13 runs, it reports 12 passes and one failure.
The failure is revealing. The query asks about AMZN profitability trends and outlook, while also focusing on AWS. The report concentrates on Amazon Web Services and does not mention the AMZN ticker. It contains relevant research, but has narrowed the subject to a segment of the company rather than the company as a whole. The ticker check does not establish whether the analysis is financially sound. It catches a basic mismatch cheaply and directs attention to a trace worth reading.
Voss’s larger principle is to choose the evaluator for the question. A string check is enough when the question is whether a string appears. A model judge is unnecessary overhead for that. Deterministic evaluators can also check formats, required fields, length limits, or forbidden phrases. Their grading logic need not be trivial; it could, for example, query a database or an API to compare a reported price with a reference. The defining feature is that the same input produces the same evaluation result.
For semantic questions, a code check is insufficient. Code can establish that a ticker appears; it cannot determine whether the report’s data is accurate, whether the report is complete, or whether a recommendation follows from the evidence. Those require a judge and criteria for what counts as a pass.
In the workshop, the code evaluator can be run in the notebook or configured in the AX interface. In either case, its inputs need to be mapped to trace data: the query becomes the evaluator’s query parameter and the final report becomes its report parameter. The check runs against the top-level agent traces, not every nested span, because it is grading the completed report. Results can then be attached back to traces so the team can filter for failures and inspect the relevant execution.
The evaluator’s scope should match the question. A check of the final report needs the query and report. A tool-selection evaluator would need the relevant tool call and its arguments. For faithfulness, the judge needs the research context as well as the report. A judge that lacks the evidence required by its task can produce a clean-looking score that answers the wrong question.
A correctness judge can ask the wrong question
The workshop’s first built-in LLM evaluator is a correctness judge. It grades all 13 reports as incorrect. That result does not mean the reports are all bad. Voss explains that the judge is trying to assess current financial research against its own stored knowledge, while the agent has searched the live web. The report refers to information current in 2026; the judge’s knowledge does not extend to that date. It sees future-dated claims it cannot verify and rejects them.
The problem is not necessarily the model’s ability to judge. It is the question the evaluation asks. A correctness judge can be useful when the judge has the information required to check a claim. Here, the output is based on fresh web research, but the correctness evaluation relies on the judge’s prior knowledge. The two systems are not working from the same evidence.
Voss replaces that test with faithfulness: does the report stay grounded in the material the agent actually collected? Since the agent works in two turns, the research output from the first turn can be extracted from the trace and supplied as context to the judge. The evaluator then compares the report with the source material used to create it, rather than asking the judge to fact-check current events from memory.
That change produces a more useful split: six reports are labeled faithful and seven unfaithful. The result does not establish that every label is correct, but it gives the team a meaningful signal to investigate. Voss treats it as a capability eval: the reports fail often enough to identify a problem and provide a basis for improvement.
The lesson is not that faithfulness is always better than correctness. The two evaluators answer different questions. Correctness is a poor fit when the judge lacks the current information that the output relies on. Faithfulness is useful when the team wants to know whether the answer represents its supplied sources. The first step is to specify what needs to be measured, then choose the evaluator that can measure it.
A faithfulness score is not an independent verification of every claim against the world. It is a judgment about whether the report is grounded in the research supplied to the judge. If the team needs to establish factual accuracy beyond that evidence, it needs an evaluation with appropriate reference material or checking logic.
A rubric should make “actionable” observable
A built-in faithfulness check still does not cover everything the financial analyst is supposed to do. Voss wants reports to offer actionable recommendations, not merely summarize data. The workshop therefore adds a custom LLM judge for actionability.
A useful rubric needs more than an instruction to judge whether a response is “good.” Voss recommends defining the judge’s role and context, setting explicit pass/fail criteria, marking input and output with clear XML tags, and defining the available labels outside the prompt. She adds labeled examples as a practical aid, while acknowledging that some practitioners worry examples can make a judge overfit.
The role prompt, she says, is less important as a performance trick than as a way to establish what the judge is looking at. Calling it an “expert financial analyst evaluator” may make a small difference; specifying that it is evaluating a financial report for actionable investment guidance supplies the relevant context.
The criteria should describe observable features. An actionable report contains a specific recommendation—such as buy, sell, or hold—identifies concrete risks with supporting data, includes forward-looking analysis rather than only historical summary, and explains why the recommendation follows. A report is not actionable if it only summarizes data, offers no next step, lists risks without evidence, or looks only backward. A human reader should be able to assess each criterion as present or absent.
These criteria should come from the error analysis rather than an imagined ideal. Voss’s example of an actionable recommendation cites revenue growth, valuation relative to a sector median, an identified concentration risk, and a specific suggested action. The contrasting non-actionable example says that Nvidia is a major semiconductor company, has benefited from AI demand, and that investors should consider various factors. It may be broadly plausible, but it supplies no specific data, supported risk, or recommendation.
Examples help clarify the boundary. They show the judge what the labels mean, especially for borderline cases. But examples need clear delimiters, just as the query and report should be wrapped in separate XML tags: the judge must distinguish instructions from data and from illustrations of the desired output. Voss’s reason for adding examples is that models can copy a pattern more readily than they can reliably infer every unstated expectation. Her caution is that a too-rigid example may narrow the agent’s range of acceptable responses; examples should clarify the criteria, not dictate a single template for all cases.
Voss also advises against spelling out the answer format in the rubric itself. The evaluator configuration can define the labels and scores—for example, “actionable” as 1 and “not actionable” as 0. Binary choices are usually easier to apply consistently than a one-to-five scale, where the distinction between adjacent scores needs its own rules. She recommends asking the judge to explain its reasoning before assigning a score. That explanation can improve the judgment and give the team a concrete account of why the report passed or failed.
A focused evaluator makes its result useful. One “god evaluator” that simultaneously grades accuracy, tone, completeness, policy compliance, and formatting may return a failure without saying which dimension caused it. Adding more labels can make the result cumbersome rather than clarifying it. Voss recommends one evaluator per dimension: in this workshop, a ticker check, a faithfulness check, and an actionability check. Together they provide more diagnostic information than a single blended score.
Not every evaluator should have the same operational force. Some checks are guardrails: a hallucinated stock price might be a reason not to ship. Others are north-star measures that express an aspiration rather than an immediate blocker. Voss’s example of a recommendation that always mentions complementary investments is a nice-to-have, not necessarily a dealbreaker. That distinction affects how a team responds to a result and what it monitors in production. A guardrail regression may warrant an immediate alert, while a change in an aspirational metric may belong in a regular review.
The judge needs its own tests
An LLM judge is a classifier: it makes a prediction, such as “actionable” or “not actionable.” That prediction can be compared with human annotations, allowing the team to test the evaluator itself. Voss calls this meta-evaluation.
Humans can annotate traces in the interface by reading a report and selecting a label. The annotation work can be done by someone with domain expertise who does not need to write code. Those labeled examples form a reference set against which the judge can be compared. Human labels are not automatically ground truth: annotators need clear criteria, and fatigue and disagreement can affect their judgments too.
The workshop’s displayed annotations are intentionally random. Voss says she chose actionable and not actionable labels at random to create disagreement for demonstration, so the reported 46 percent agreement—six of 13—is not a meaningful assessment of the judge’s real performance. In a real evaluation, a disagreement is a prompt to inspect the judge’s explanation and the human label. Was the rubric ambiguous? Was the human label mistaken? The explanation can identify a missing distinction. For example, “includes forward-looking analysis” may be too broad if a report discusses future trends without giving any recommendation or guidance. Tightening the criterion can resolve that ambiguity.
For a real meta-evaluation, Voss recommends keeping some labeled examples out of rubric development. She suggests using about 75 percent of labels to revise criteria and examples, then holding out the remaining 25 percent to see whether the judge generalizes to examples it has not been tuned against. Otherwise, the rubric may simply fit the labels it has already seen.
Precision and recall help describe what happens when the judge flags a failure. Precision asks whether reports labeled “not actionable” really are not actionable. Recall asks how many genuinely non-actionable reports the judge catches. The measures can trade off: prioritizing recall can mean more false positives, while a high-precision judge may miss some failures. Voss says recall is often the priority because it is usually better to investigate a false alarm than to let a real failure through, though the right balance depends on the use case. A 12- or 13-example set is too small for stable estimates, but can still reveal whether the judge is obviously off target.
LLM judges also have known biases. They may favor the first or last option in a comparison, prefer longer responses even when length adds filler, and mistake confidence for correctness. They may also prefer outputs generated by the same model family. Voss uses Sonnet to judge reports produced by Haiku to reduce this self-preference effect, and says using a different model provider can reduce it further. These are tendencies to watch for, not guarantees that switching models will eliminate bias.
Voss cautions that the benchmark is human performance, not perfect agreement. She cites human inter-rater reliability—two people judging the same material—often being as low as 0.2 to 0.3 on Cohen’s kappa. Her point is that an LLM judge should be compared with human performance while recognizing that humans themselves disagree. Occasional disagreement with one annotator may reflect genuine ambiguity. A judge that disagrees consistently deserves closer scrutiny. Voss’s fairness test is practical: when a trace fails, can a reader see what the agent got wrong? If the answer looks fine, the evaluator may be the part that needs fixing.
Before an agent is judged against a task, the task itself should also be clear and solvable. Voss recommends a reference solution and tests in both directions: cases where a behavior should occur and cases where it should not. A test that rewards web search only when it is needed, for example, should not produce an agent that searches every time. If an agent repeatedly scores zero, she says to investigate whether the task or rubric is broken before concluding the agent is incapable.
Turn failures into controlled improvements
An evaluation score tells a team where to look; it does not establish that a fix worked. Voss recommends saving failing traces as a dataset and keeping passing examples as a regression set. The failure dataset provides a focused, usually cheaper set of cases for testing changes. The passing set helps reveal whether a fix damaged behavior that already worked.
In AX, the team can filter traces by an evaluation label, select the relevant failures, and save them as a dataset. The saved examples can then be run through a revised agent or prompt without having to reproduce the failure by hand. The passing cases belong in a separate dataset because a fix aimed at failures can affect other behavior. Together, failing and passing examples form a curated set for measuring the system against known problem cases and previously successful behavior.
The data should change as the product does. Synthetic examples can get an agent started before real users arrive. Once real traffic is available, genuine failures and queries should become part of the test set. Voss’s direction is to let real failures replace imagined ones as evidence accumulates. The resulting dataset captures the application’s own users, failure modes, and quality bar.
In the workshop, the judge’s explanations become input to a prompt-improvement step. The improvement request includes explicit product requirements—correct tickers, recent financial data, forward-looking analysis, recommendations, evidence-backed risks, and comparison sections for multi-ticker requests—along with the reasons reports failed. Another model is asked to identify recurring themes and rewrite the research and writing prompts. Voss emphasizes the requirements because “make the evals pass” is an unsafe objective on its own: a coding agent could optimize for the test rather than the product. Proposed changes should address the intended behavior, not just the labels.
The revised prompts are then tested in an experiment against the same failure dataset, using the same actionability evaluator. In the demonstration, the new prompts move the examples from failing to passing. Voss notes that the result is unusually clean: she had the benefit of preparing the workshop, and real iteration will not usually produce a one-shot improvement from roughly half failing to all passing. The important thing is the comparison, not the showpiece result.
An experiment holds the examples and evaluator constant while changing the task—in this case, the agent’s prompts. Because the agent is nondeterministic, a run can still vary in its searches and results. But fixing the test cases removes one major source of variation and makes it easier to attribute a score change to the modification. A task can be the whole agent or a smaller component, as long as it takes an example and returns an output that can be evaluated. The comparison does not remove all uncertainty: a live search may return different material on different runs. It does provide a controlled basis for asking whether the change improved results on the same examples.
The improvement cycle is to find failures, read their explanations, identify recurring causes, make a change, and run an experiment. The explanations connect a score to a possible remedy. A repeated complaint about missing recommendations suggests a systematic issue; it is more useful than treating each trace as an isolated defect.
For sample size, Voss offers a distinction between directional feedback and a shipping decision. She says 12 to 20 examples can provide a directional signal in a workshop-sized experiment, while a shipping decision calls for more—she suggests roughly 200 to 400 examples. She also notes that halving the margin of error requires about four times as many examples. Those figures are guidance from the workshop, not a substitute for choosing a sample size appropriate to a particular decision and failure rate.
When deciding what to change, Voss ranks better data above prompt improvements, prompt improvements above model selection, and model selection above hyperparameter tuning. If the agent is searching the wrong sources or using stale information, changing the prompt will not fix the underlying problem. Explicit instructions and examples can often be high-return changes; a more capable model may help with problems prompting cannot solve, but may add cost and latency. Temperature and top-p are easy to adjust, Voss says, but should not be the first place teams invest.
This approach can also be reversed: write the eval before building a feature. If a refund workflow must verify a customer’s identity first, specify and test that behavior before implementing it. The evaluator defines what “done” means, as a test can in conventional test-driven development. People closest to the product requirements can help define those tests in ordinary language; they need not write code.
Production traffic supplies the next test cases
Passing a development experiment is not the end of the job. The agent still has to perform on traffic the team has not seen. Online evaluations apply the same checks—such as ticker mention, faithfulness, and actionability—to incoming traces. Voss says evaluators can run at span, trace, or session scope; the scope should match the question the team is asking.
Online evaluation changes the cadence of testing. Rather than exporting a batch of traces and running a one-time notebook evaluation, the team can configure evaluators to run automatically as new production traces arrive. The result is a stream of labels and, for LLM judges, explanations attached to observed behavior. Voss describes monitors using evaluation labels to alert the team when something goes wrong. Investigation then starts with the traces and the evaluation results that prompted concern.
A team need not grade every production interaction with an LLM judge. Voss says those evaluations cost money and suggests sampling traffic to control cost while retaining a directional signal. Her slide offers 10 percent as a default and mentions 1 percent for expensive or high-volume evaluation. Those are workshop recommendations, not a guarantee that a given rate will be representative for every application. The required sample depends on traffic, evaluator cost, and the signal the team needs.
In Voss’s production loop, online evals grade traces and produce labels and explanations; monitors can alert a team when something goes wrong. The team investigates, finds failing traces, saves failures as a regression dataset, improves the agent, runs an experiment to verify the fix, and ships it. New traffic then brings new failure modes and edge cases. Monitoring extends the loop to behavior the team did not anticipate in its offline examples.
A failure caught in production can become tomorrow’s regression test. The team can save the failing trace as a dataset example and check it during a later round of changes. The next change can then be tested against a known problem rather than relying on someone to reproduce it by hand. Passing cases matter too: they help establish whether an apparent improvement has caused a regression elsewhere. Over time, the dataset grows to reflect the application’s particular users, recurring errors, and definition of acceptable quality.
Voss describes that collection as a compounding asset, provided the team continues to update and use it. The value is not simply the number of examples. It comes from preserving cases that represent the product’s own quality requirements and using them in subsequent experiments. As new traffic arrives, the team can add newly discovered edge cases and revise evaluators where the old criteria no longer capture what matters.
Coding agents can help with repair work, especially when prompts, retrieval settings, and tool definitions are spread across a repository. Voss’s workshop demonstrates a small version by feeding judge explanations to a model and asking it to rewrite two prompts. For a larger system, she describes exporting failing traces and explanations and giving them to a repository-aware coding agent such as Claude Code or Cursor. The agent can inspect the code and propose changes across the parts of the application implicated by the failures.
Voss sets two safeguards for that approach. First, give the coding agent product requirements along with failing traces and judge explanations, so it does not optimize merely to pass the tests. If the only instruction is to make the evals pass, the agent may find a way to satisfy the check without delivering the intended product behavior. Second, ask it to identify recurring themes rather than patching each trace individually. A collection of failures may share one cause; fixing that cause is preferable to accumulating special cases.
The proposed fix still needs an experiment before shipping. The coding agent’s explanation of what it changed is not evidence that the system improved. Run the revised system against the relevant dataset with the same evaluator, inspect the outputs and failures, and check the passing regression cases as well. If the results are not convincing, return to the traces and the requirements rather than accepting a score in isolation.
The resulting cycle is application, traces, online evals, labels and explanations, monitoring, investigation, dataset, code or prompt change, and experiment. Production traffic supplies new signals; those signals become tests; controlled comparisons check the proposed fixes; and verified changes return to production. Voss calls this the software development life cycle closing on itself: the software is used to generate evidence that helps improve the software.
For teams starting from scratch, Voss does not prescribe doing everything at once. Start by instrumenting the agent and reading its traces. Then write a small deterministic evaluator for the most important observable requirement. Add semantic judges where the question needs them, compare those judges with human annotations, and build datasets from real failures. The sequence matters: automate after understanding what is failing, not instead of understanding it.
