Agent Improvement Depends on Linking Production Traces to Offline Evals
Marc Klingen, Langfuse’s co-founder, argues that improving agents is becoming less a matter of rewriting prompts and more a matter of building a feedback loop around them. AI can increasingly inspect production traces, propose fixes, and test them against datasets and evaluations, reducing work that teams once did manually. But Klingen says people still need to set the goals and judge what counts as a meaningful improvement; in a Langfuse changelog-writer demo, an agent identified internal jargon and tested a revision before it was submitted for review.

The improvement loop matters more than the prompt
Marc Klingen argues that building better agents increasingly depends on designing loops around them, not continually rewriting prompts. As models have improved, they have become capable of taking on more of the work that used to require a person: examining production failures, proposing changes, and checking whether those changes improve results.
Klingen traces the shift to a broader change in what teams can ask models to do. When Langfuse began in 2023, he said, even a basic attempt to turn a GitHub issue into a pull request was limited to editing a single-page application. Multi-file changes were too difficult. As model capabilities expanded, agent workflows became more capable too: from applications using tools, to loops that evaluate their own results, and now to systems that can reason about what to fix.
The practical consequence is that teams need to connect what happens when an agent serves real users with what they test before changing it. Klingen describes a reference process with two sides. Online, teams trace and monitor agent behavior, including metrics, alerts, and user feedback. Offline, they build datasets from examples, experiment with models or prompts, and evaluate candidate changes. A revised agent is deployed, its performance is observed in production, and the results feed the next round.
The connection matters because either side alone gives an incomplete picture. An offline benchmark can drift away from what users actually do if its dataset is not refreshed from production. Production monitoring can reveal problems but does not, by itself, show whether a proposed fix works against a repeatable set of examples. In Klingen’s account, the improvement process depends on keeping traces, datasets, experiments, and evaluations in sync.
That process has traditionally taken substantial manual effort. Someone has to inspect traces, decide which failures matter, add examples to datasets, define evaluation criteria, form a hypothesis, and test changes. Klingen’s question is how much of that work can now be delegated to AI. The point is not to remove the feedback loop, but to use models to do more of the work inside it.
Automate fixes while people keep setting the target
Klingen breaks the process into nested loops. At the lowest level, a model produces the next token. Above that, an agent handles a user turn, potentially calling tools and running through multiple steps. A goal loop takes a dataset, runs the agent over it, and evaluates the results. A meta loop examines failures and proposes fixes. At the highest level, a roadmap loop uses user feedback and human judgment to decide what should improve next.
These layers describe different kinds of work, not simply increasing degrees of automation. A model can complete an agent turn without being able to judge whether the agent is meeting its broader goal. Likewise, an agent can run a dataset and report evaluation results without deciding whether the dataset captures the work the product should do. Klingen says that higher-level loops have become more feasible as models have improved, but humans remain involved in setting direction.
People still need to review which examples enter a dataset and whether evaluation criteria represent the task. They also need to inspect proposed fixes for overfitting: an agent can optimize against a dataset while making the system worse in cases the dataset does not capture. And people need to decide whether a failure is important at all. A strange production case may be outside the intended scope of the agent; fixing every unusual behavior can waste effort or distort the product.
Klingen describes proposing implementation changes as a particularly useful place to apply AI. Given production traces, existing evaluation criteria, and examples that reproduce known problems, an AI system can suggest changes and test them against the dataset. Teams might try a different model, aggregate context in a different way, or explore another implementation approach. The team still maintains the boundary of what is being optimized; AI helps generate and assess possible changes within that boundary.
Teams are also beginning to use AI to maintain evaluation criteria and datasets. If a support agent sees new kinds of questions in production, an AI system can propose additions to a dataset so it better reflects actual usage. If traces reveal recurring failures, it can propose criteria to test for them—for example, whether the response avoids naming competitors, stays concise, or answers in the language the user used.
Because these datasets and criteria define what the improvement process treats as success, Klingen says teams still need to review proposed changes. Handing that work over entirely risks producing “more slop.”
You really want to be like involved in setting the setting like the goals and how you make changes to datasets, evaluators. You want to be in that loop involved because you set thereby the direction. And then AI can automate against it.
The target itself is not fixed. Klingen uses customer support as an example: a company might begin with a broad goal of automating most support responses because customers wait too long and the company has many people answering them. But that ambition does not specify all the work support agents encounter. As teams examine real cases, they learn what the system needs to handle and where the original goal was incomplete.
People therefore do more than approve a finished plan. They set the course as the application’s requirements become clearer, and can change direction when new error classes appear. AI can take on more of the lower-level cycle of finding failures and testing fixes while people revise the goals and decide how to respond to what the system uncovers.
The aim is less manual effort without outsourcing judgment
Klingen contrasts a manual process with the risks of turning the entire improvement loop over to AI. Teams that trace agent behavior, build datasets, and evaluate changes can achieve high quality, he says, but doing so takes time: someone has to work with the data each week. At the other extreme, adding higher-level loops and asking an agent to proceed with little human direction may require less hands-on effort, but can produce output that does not make sense.
The outcome he wants is a balance: people stay involved at the points where they set direction, while AI takes on more of the tedious analysis and iteration. Klingen suggests that quality could improve as time investment falls. A team’s ability to inspect production failures is limited by how much time and patience people can devote to the task. AI can examine more error cases, but its contribution depends on having a target that people have defined and continue to revise.
User behavior can also supply feedback without requiring an agent developer to label every case manually. In customer support, Klingen points to customers saying an answer is wrong or that the agent has not understood them. In a workflow where an internal employee reviews a suggested message before sending it, that person’s approval or edits can provide another signal. Teams can use these signals alongside traces and other production data to decide what to investigate.
A changelog agent learns to stop leaking internal jargon
Marc Klingen demonstrates the approach with a changelog writer built for Langfuse. As the engineering team uses AI to ship more frequently, keeping customers informed can become a bottleneck. The changelog writer takes merged pull requests and drafts updates to documentation and release notes. A human reviews the draft through a pull request, approving it or requesting changes; the writer can revise the draft in response.
Those approvals and requested edits become feedback for improving the writer. Klingen emphasizes that the risk is not simply producing text that is inaccurate. Poor communication can misrepresent what was released and undermine the release itself.
In the demo, a coding agent uses Langfuse’s agent skills to investigate how the changelog writer could be improved. Klingen reports that the writer was factually correct, but its clarity was low, human edit requests were frequent, and it sometimes used internal jargon. Terms that made sense in the team’s code or internal comments—such as an “ingestion pipeline” or a “batch eval queue”—were implementation details rather than useful descriptions for customers.
The agent proposes changes to the dataset and evaluation criteria so they test for user-facing language and clarity. It then suggests a change to the writer’s implementation, specifically how it receives context and instructions. The revised version can be compared with the existing one on the updated dataset before a change is accepted.
| Evaluator | V1 baseline | V2 candidate | Change |
|---|---|---|---|
| Format compliance | 0.967 | 0.967 | 0.000 |
| Accuracy | 1.000 | 1.000 | 0.000 |
| User-facing language | 0.675 | 0.900 | +0.225 |
The comparison shows the targeted user-facing-language score rising from 0.675 to 0.900, while format compliance and accuracy remain unchanged. Klingen presents this as a positive result: the proposed change improves the targeted measure without changing the other two reported scores.
The demo uses an internal application and a workflow in which the agent raises pull requests rather than publishing changes directly to production. Klingen describes the setup as relatively permissive: the team is comfortable merging the proposed implementation and then acting on the next round of feedback. The example therefore shows a candidate being evaluated and submitted through a reviewable development workflow, not a system independently publishing customer-facing changes.
More automated analysis makes trace ownership more important
Marc Klingen says the pattern is already common among Langfuse’s strongest users: collect production signals, keep them alongside detailed execution traces, and run agents over batches of data on a schedule—perhaps daily or weekly, or whenever enough useful examples have accumulated. Inputs can include user feedback, internal labels from a production system, or annotations collected in an application.
As agents take on more of the analysis, the data layer has to support a different workload. Klingen describes Langfuse as historically write-intensive, because teams were ingesting traces and evaluations. With agents querying and reading much larger volumes of data, the workload is becoming more read-intensive.
He also argues that teams should own this data rather than depend on a system that forces them to sample traces or retain them only for a short period. Traces from a year earlier may become useful context for later improvement work. If agents are expected to search across that history, teams need to be able to keep it and query it at scale. Retention and query capacity therefore shape what evidence is available to the improvement loop.


