AI Evaluations Need Human Calibration Before They Can Gate Deployments
Tejas Kumar of IBM argues that AI evaluations should be treated as calibrated reliability checks, not as proof that a system works. In a customer-support example, a test passes an answer that wrongly approves a return because it contains the word “cannot”; an LLM judge can also mislead unless its verdicts are checked against human-labeled cases and the relevant policy. Kumar’s approach is to examine disagreements, refine the judge and dataset, and use measured agreement as a deployment gate.

Evals catch failures before users do
Tejas Kumar draws a distinction between two forms of reliability in AI systems. A harness supplies protections while an agent is running: it can control tool access, treat tool results as data rather than instructions, or hand off when a task hits an authentication gate. An evaluation, or eval, is meant to catch problems ahead of time, before a system is deployed.
The two work together. Kumar calls a harness “just-in-time reliability” and evals “ahead-of-time reliability.” He compares the distinction to runtime and compile-time checks in programming: a harness responds as work happens, while an eval set can stop a change from shipping. A capable harness and a good evaluation suite can also make less capable models useful, he says, and may do so at lower cost.
Harnesses tend to be more just-in-time reliability. And what I'm making today, the case I'm making about evals, is that they allow for ahead-of-time reliability.
Kumar describes an Instagram support chatbot as a warning about unsafe agent behavior. In his account, it had access to an API that could change user accounts. Attackers used a VPN to appear to be in a target’s location, he says, then asked the bot to add a secondary email address to the account. Kumar says 20,225 accounts were compromised. He says he does not know what failed inside the company; his point is that a harness, an eval, or both could have caught the dangerous behavior before it reached users.
An eval is a way to instrument reliability when exact outputs are not predictable. A conventional unit test can assert that add(1, 2) returns 3. For an AI application, the useful question is usually not whether a generated answer matches a fixed string, but whether an agent behaves acceptably across different scenarios. Kumar calls evals “fuzzy unit tests”: the underlying purpose is familiar, but the system being tested is not deterministic.
For a customer-support agent, scenarios might ask whether the answer followed policy, used the right tools, and resolved the customer’s issue. Some checks need no model judge at all. If tool calls are recorded as structured data, ordinary code can check whether the correct tool was called and with which arguments.
Kumar describes four parts of a useful eval suite: test cases, expected behavior, a judge, and a way to aggregate results. Cases provide inputs and agent responses; expected behavior describes what a good outcome should look like without requiring exact wording. A judge assigns a verdict, and aggregation turns individual verdicts into a measure such as an agreement rate. The point is to assess behavior across relevant scenarios, using model judgment only for aspects that cannot be reduced to exact assertions.
A green test can reward the wrong answer
The live return-policy example shows why simple pattern matching is not enough. A customer bought headphones 20 days ago and asks to return them; the store’s policy allows returns within 14 days. Kumar first writes a test that checks whether the response contains “cannot.” The answer “No, you cannot return them because it is outside our return window” passes.
Then he changes the other answer to: “Yes, you can return them, for free. You cannot pay for the shipping.” The test turns green, although the answer approves a return outside policy. The test output shows both answers passing after the misleading one gains the word “cannot.” The string match catches the word but not whether the answer follows the policy.
Kumar turns to an LLM judge, initially asking a model to choose which of two answers is right. The model sometimes returns the expected choice, but repeated runs can vary, and calls can time out. More importantly, a judge can have systematic preferences that distort its verdicts. Kumar identifies four:
- Position bias: In a pairwise comparison, a judge may favor the first option. Reversing the order is one way to test whether the preference follows position rather than quality.
- Sycophancy: A judge may prefer the answer that sounds more accommodating, even when it contradicts policy. In the return example, “Yes, you can” can seem more helpful than a correct refusal.
- Self-preference: A judge may favor text generated by its own model family. In Kumar’s demonstration, a GPT-generated answer was selected after being placed against another answer, even though the judge had previously rejected that response.
- Verbosity: A longer answer can win because it sounds more complete, even if it is wrong.
These are reasons to be cautious about pairwise comparisons, not just reasons to choose a better prompt. Kumar prefers pointwise judgments for many production checks: assess one response against a standard and return a pass or fail. Pairwise comparison asks a judge to choose between two answers; pointwise comparison asks whether one answer meets a defined bar. A direct standard can make the desired behavior clearer than a relative choice between two imperfect options.
Kumar warns against treating an eval’s green status as proof that the system works. The judge is itself non-deterministic, and its verdicts need to be tested against human decisions. Repeated runs can reveal instability; labeled examples can reveal systematic disagreement. Neither removes the need to decide what the judge should be measuring.
Calibrate the judge against human verdicts
Kumar’s calibration process starts with a dataset. Each record contains a customer scenario, the agent’s answer, and a verdict supplied by a human or subject-matter expert. He suggests beginning with roughly 30 examples, then expanding the collection. The human verdicts are the reference against which the judge is measured.
In the workshop’s examples, one customer reports that a coffee grinder arrived last Tuesday but its motor only buzzes. A response offering a free replacement or refund is marked as a pass. Another response says the grinder cannot be returned because it has been used, even though it is defective and within the return window; that is marked fail. The customer’s informal “pretty annoyed tbh” is intentional: Kumar wants to see whether the judge rewards a response for soothing the customer rather than following policy.
The judge sees the scenario and answer, but not the human verdict. It returns its own pass-or-fail decision, which is compared with the reference. Kumar’s target is around 80–85% agreement. He says a human team may also agree only about 80% of the time; perfect agreement can be a warning sign of overfitting or sycophancy, not evidence of a perfect judge.
His live run makes the calibration concrete. The first version of the judge gets only four of 30 examples right. After he normalizes the capitalization of its verdicts, the score improves, but it still misses the 80% threshold. He then examines disagreements. One concerns a customer who used a yoga mat for a few sessions, found it too thin, and asks for a refund after ten days. The human verdict is pass: the response correctly says that change-of-mind returns require the item to be unused and in its original packaging, while inviting the customer to report an actual defect. The judge marks it fail because it does not yet know that part of the policy.
Kumar’s response is to investigate the disagreement and supply missing context. The judge initially knows only that defective items can be returned within 14 days. The team adds the unused-item and original-packaging condition. This improves agreement, but does not immediately solve the problem; the remaining failures need to be examined too.
At that point, Kumar distinguishes a knowledge problem from a model-quality problem. First, give the judge the relevant policy. If it still fails to apply that policy, a stronger model may be warranted. In the demonstration, he switches from GPT-3.5 Turbo to GPT-4o mini after providing additional policy context. A remaining mismatch involves a customer asking for an exception six weeks after buying a speaker, on the grounds that they are loyal and the item stopped charging. The human verdict is pass for an answer that acknowledges the customer but says the 14-day defect window has passed. The judge still wants the answer to grant the exception. Kumar adds an explicit instruction that the policy is non-negotiable, even for a loyal customer.
The calibration loop is to compare the judge with human labels, inspect disagreements, add relevant context or change the model when needed, and run the dataset again. In Kumar’s final run, three of the 30 judgments disagree with the human verdicts, reaching the 80% threshold.
Use agreement as a gate, then keep the dataset alive
Once the judge reaches an acceptable agreement rate, Kumar proposes using that rate as a continuous-integration gate. In his example, a change does not ship if agreement between the judge and human labels on the evaluation dataset falls below 80%. This is a comparison on the labeled test set; it is not a claim that every production response is continuously being judged against a human verdict. The workshop’s test runner is Vitest, but he says the same approach can be implemented with other tooling.
There is a wrinkle: a conventional test runner treats any failing individual test as a reason to fail the suite. For a judge-calibration test, Kumar instead aggregates the results and fails only if overall agreement with the human labels falls below the chosen threshold. His demonstration passes at exactly 80%, or 24 out of 30 examples. He also says 100% agreement should not automatically be treated as ideal; it may signal overfitting. A production gate may need an upper bound as well as a lower one.
Kumar advises pinning the judge’s model version rather than relying on a model-name alias. If the underlying version changes, an alias could change the judge without the team knowing, and the evaluation may behave differently.
Passing the gate is not the end of the process. Customer behavior, agent responses, and policy can change. Kumar recommends sampling real traffic, collecting inputs and outputs, having an internal team score cases manually, and adding those labeled cases to the dataset. The judge can then be checked against the updated human verdicts. The dataset is “living,” rather than a set of examples written once and forgotten.
When production traffic is scarce, Kumar suggests synthetic data. Agents can generate plausible customer requests using policy and existing examples as context; they can also be asked to produce adversarial cases or variations on failures already seen. Red-team prompts—cases designed to find unsafe or policy-breaking behavior—are another form of evaluation data. Kumar suggests using known failures as prompts for generating further cases that might expose the same weakness in a different form.
A policy judge needs access to the same source as the agent
A judge cannot reliably assess policy compliance if it has not been given the policy. Kumar’s live demo makes that gap visible. Without policy context, a judge asked which of two support replies is more helpful repeatedly favors a long, upbeat answer that promises a refund, a prepaid shipping label, a quick refund, and a discount. The answer is confident and customer-friendly, but it invents a 30-day guarantee when the actual policy allows returns only within 14 days. Adding the policy to the judge’s instructions changes the verdict: the shorter answer that declines the return is preferred.
Kumar presents OpenRAG, an open-source tool from IBM, as one way to make policies retrievable rather than burying them in a prompt. He describes it as a knowledge base that can ingest company material—including documents, spreadsheets, presentations, audio, and video—and make it available to a retrieval-augmented generation system. In the demo, the document-processing pipeline converts uploaded material into formats suitable for an LLM, stores and indexes it, and returns policy passages as context.
The retrieval demonstration starts with no relevant supporting source. The system responds with generic advice and speculates that a 30-day return policy may apply. Kumar then uploads a refund-policy PDF. The displayed answer cites the file and retrieves the 14-day rule, concluding that a purchase from 20 days earlier is outside the refund period. The demonstration shows how retrieval can connect an evaluation to a changing source of policy rather than a manually maintained prompt.
Kumar says policies at a large organization can change and describes IBM’s internal HR assistant as a system that gives employees grounded answers and can handle requests such as vacation time. He attributes its performance to an evaluation pipeline based on real-world policy. OpenRAG, he says, can provide a policy retriever for evaluations as well as applications.
In response to an audience question about whether the production agent and judge should share the same policy, Kumar says they should. Ideally, both query a shared source through retrieval rather than copying policy text into separate prompts. He favors making the judge’s retrieval setup as close as possible to the agent’s: the closer the judge is to the agent, he argues, the more useful the CI signal. The judge can also be autonomous and use retrieval tools, like the production agent. Both need access to the same policy if the judge is to give a useful indication of how the agent will behave.
Use the cheapest check that can answer the question
Kumar’s cost recommendation is to use the cheapest adequate check first. Start with deterministic assertions where behavior is deterministic. If exact matching is too brittle, try pattern matching or schema validation. Inspect traces and message envelopes: tool calls, arguments, timing, and other structured events can often be checked with ordinary code. Use an LLM judge when the behavior requires a more interpretive assessment, and begin with a low-cost model before moving to a more capable one.
The checks answer different questions. Exact matching and schema validation can catch malformed or missing data without asking a model to interpret it. Traces can show whether a tool was called, how often, with which arguments, and how long it took. Those checks do not establish whether an answer handled a nuanced customer request correctly. Kumar sees model judgments as useful for that kind of assessment, calibrated against the human-labeled dataset.
The workshop’s examples also clarify where Kumar thinks evals belong. Asked whether developers should write evals for the skills they use with coding agents, he says generally no: if they are using someone else’s harness, evaluating the harness is that builder’s responsibility. He refines his rule to: build evals when you build a harness, or when you use a harness you are responsible for. Evals and harnesses complement one another—one checks behavior ahead of time, the other can protect the system at runtime.
His broader criterion is non-determinism. Whenever inputs, outputs, or both are non-deterministic, some form of evaluation is useful. The practical task is to build a representative dataset, examine failures, add missing context, and continue checking against real use.


