From Model Decisions to End-to-End Workflows
OpenAI’s Decisions API narrows a model’s task to choosing among defined actions, while Langfuse’s Marc Klingen describes using production traces and evaluations to improve agent behavior. Google’s Amin Vahdat and Microsoft’s AI for Good Lab’s Sparrow Studio examples widen the frame: useful AI depends on the infrastructure, data, and human oversight around the model.

Make the model’s job a decision
A workflow often does not need an essay from a model. It needs the next action: which team should receive a request, which lane should a vehicle take, or where should a robot turn its head. OpenAI’s Decisions API is built for that narrower job. The application supplies the input, asks a question, and defines a small set of possible answers; the model selects among them, and the application acts on the result.
In a sales demo, messages from a design agency, a large software company, and a freelancer are sorted using details such as team size, buying stage, and requested next step. The interface reports 81 milliseconds of API processing for six decisions. Romain Huet says narrowing the task to a limited set of choices makes the API nearly ten times faster. That figure comes from a launch demonstration, not a general benchmark across applications, but it illustrates the appeal: a bounded choice can fit inside an interaction where a broader response would introduce unnecessary delay.
The same pattern extends beyond text classification. In a driving-game demo, the API chooses among three lanes based on an image of the road. In a robotics example, it uses camera input and a spoken instruction to direct a small robot’s head toward an apple—or toward a game controller when asked which object would be more fun to play with. These demonstrations show how a choice can connect perception to an action, but they do not establish how the system handles every ambiguous instruction or unfamiliar environment.
The important design decision sits partly with the application developer. The developer defines the question and available actions; the model interprets the input within that frame. That can make a workflow faster and more predictable, but it also means the action set must fit the situation. A quick selection among known options is not a replacement for open-ended reasoning. Huet draws the same boundary around computer use: the API may suit basic screenshot-and-action interactions, while more sophisticated browser or computer tasks call for other capabilities.
As applications shift from generating text to taking steps, this distinction matters. A model can be one component in a larger process, with the application deciding what choices are valid and what happens next. The speed claim is most useful when read in that context—not as a claim that every task can be reduced to a menu, but as a case for choosing a smaller model job when a workflow needs a timely action.
Close the loop on real-world behavior
A fast decision handles one moment in a workflow. Improving an agent means learning from many such moments, including the ones that go wrong. Marc Klingen of Langfuse argues that the work is moving beyond repeated prompt edits toward a feedback loop connecting production behavior to datasets and offline evaluations.
Production traces can reveal what users actually asked, which steps an agent took, and where the outcome failed. But spotting a failure is not the same as showing that a proposed fix works. A team needs a repeatable set of examples and criteria to compare the old and revised systems. If those tests stop reflecting real use, they can reward improvements that matter less to users; if teams monitor production without testing candidate changes, they may have no reliable way to distinguish a fix from a regression.
Klingen’s changelog-writer example makes that loop concrete. The agent drafts customer-facing release notes from merged pull requests, and a person reviews the draft through a pull request. The team’s feedback exposed an issue beyond factual accuracy: the writing sometimes used internal engineering terms that customers would not understand. An AI coding agent proposed changes to the dataset and evaluation criteria, then suggested an implementation change and tested it against the updated cases.
| Measure | Existing version | Proposed version | Change |
|---|---|---|---|
| Format compliance | 0.967 | 0.967 | No change |
| Accuracy | 1.000 | 1.000 | No change |
| User-facing language | 0.675 | 0.900 | +0.225 |
The proposed version improved the user-facing-language score while leaving the reported accuracy and format-compliance scores unchanged. The result is evidence about this evaluation and this revision, not proof that the writer will perform better in every customer situation. Klingen’s demo keeps the change reviewable: the agent submits a proposed change rather than publishing it directly.
That boundary is central to the method. AI can inspect more traces, suggest examples, and test possible fixes, but people still decide which failures matter and what the system should optimize. They also need to judge whether an evaluation represents the task or merely encourages the agent to perform well on a narrow test set. In Klingen’s account, human oversight is not just a final approval step; people set and revise the direction as the product’s real requirements become clearer.
The practical shift, then, is not from human work to no human work. It is from manually inspecting every case toward using AI to help search and iterate—while retaining human judgment over goals, datasets, and meaningful improvements. A production signal becomes useful for improvement when it can travel into a test, and a test becomes useful when teams keep checking it against what users actually experience.
Measure what the system actually delivers
The workflow view also changes what counts as infrastructure performance. Amin Vahdat, Google’s chief technologist for AI infrastructure, distinguishes theoretical chip speed from goodput: the useful work a system completes after accounting for delays, failures, and recovery. A powerful accelerator may still deliver little useful output if the surrounding network, software, storage, or recovery process prevents a job from finishing efficiently.
At large scale, a synchronous workload can stall when one component fails and others must wait. The system has to find the fault, recover from a checkpoint, and possibly repeat work. Vahdat says that at 100,000 accelerators, failures can occur multiple times a day—and, depending on configuration, multiple times an hour. The point is not that all failures have the same cause or effect; hardware, networking, software, and runtime problems can all reduce delivered performance. Goodput puts those operational losses inside the performance measure rather than treating reliability as a separate concern.
That framing becomes more consequential as workloads become long-running agents. A person may take seconds to read an answer and ask a follow-up. An agent can process a response and make its next request in milliseconds. The added demand is not just for accelerators generating tokens: agents may also need CPUs to coordinate steps, networks to reach services, and storage to retrieve context. The system must deliver a sequence of useful operations, not simply peak compute.
Vahdat’s response is to consider the whole stack, from chips and model software to networking, cooling, and power. Co-design can improve end-to-end performance when components are tuned to work together. But specialization is a trade-off, not a universal prescription: a system designed around a durable workload may be faster or more efficient, while a more general design can preserve flexibility if demand shifts. Google’s TPU 8i and 8t, for example, are designed with different emphases on inference and training, while retaining some ability to run the other workload.
The same tension appears in facilities. Dense accelerator racks demand far more power and networking than storage racks, making a single fully interchangeable data-center design potentially wasteful. Meanwhile, training benefits from large, concentrated clusters, while serving needs to be distributed nearer to users. These choices shape not only capacity but the distance between the parts of a workflow and the reliability of their connections.
The useful comparison is therefore not simply one chip against another. It is the useful work delivered by a complete system under its actual workload and failure conditions. FLOPS can describe theoretical capacity; goodput asks whether that capacity turns into completed work at the pace applications require. Neither specialization nor flexibility wins in isolation: the right balance depends on how persistent the workload is, how components interact, and what performance the user needs.
See the workflow in an applied setting
Sparrow Studio, presented by Microsoft’s AI for Good Lab, offers a grounded example of a system organized around an ongoing pipeline rather than a single model call. An ecology project can collect uploaded historical surveys or ingest images from cameras as they arrive, then assign models to analyze the incoming data automatically. The project ties together where data comes from, how it is processed, and how it is managed.
The examples span historical camera-trap archives, near-real-time images from 4G trail cameras, and feeds from Sparrow Edge devices. The workspace also supports other forms of input, including audio, video, drone imagery, and GPS tracking. That range matters because the workflow has to handle more than a model’s interpretation of one image: users need to configure data sources, identify locations and deployments, and manage the resulting material.
The tutorial lists 53 models across detectors, classifiers, and aerial models. It shows how a project can be assigned one or more models, but does not establish which model is best for a particular species, environment, or camera. The model catalog is one part of the setup; ingestion and project configuration are the rest. Users can upload past surveys in bulk or configure a real-time project to receive and process data from connected sources.
This is a different scale of the same design question raised by decision APIs and agent evaluation. What should the model do, what information should reach it, and how does its output feed into work that continues afterward? A bounded decision API makes the next action explicit. Production traces and evaluations help teams improve an agent over time. A project workspace connects incoming data to automated analysis and ongoing management. In each case, the model’s usefulness depends on the surrounding workflow: the data, choices, tests, infrastructure, and human decisions that make its output actionable.







