Orply.

Tool Use and Feedback Define the Agentic AI Stack

Azalia MirhoseiniStanford OnlineWednesday, September 2, 20264 min read

Stanford adjunct professor Aakanksha Chowdhery and assistant professor Azalia Mirhoseini argue that agentic AI requires more than better prompting: systems must reason, use tools to act in an environment, learn from feedback and improve toward goals. Their graduate-level program frames test-time scaling, tool use and self-improvement as the core components, while stressing that evaluation must account for agents working across longer, multi-step tasks rather than isolated responses.

Agentic systems combine reasoning, action, and feedback

Stanford presents the offering as a graduate-level curriculum for working professionals, taught by ? aakanksha-chowdhery and Azalia Mirhoseini. Chowdhery frames agentic workflows as a consequential change for knowledge work in the coming months and years. The relevant distinction is not merely a model producing an answer to a prompt. Agentic systems are presented as systems that can reason, take actions, learn from feedback, and improve in pursuit of a goal.

Mirhoseini describes those capabilities directly: agents can think, act, learn from feedback, and improve themselves to achieve goals. Chowdhery identifies three components for building them: test-time scaling, the ability to use tools to take actions, and self-improvement.

In terms of building these systems, there are really three components. There's test time scaling, the ability to take actions using tools, and the ability to self-improve.

? aakanksha-chowdhery · Source

The examples shown make reasoning and action inseparable parts of the design. A slide on ReAct contrasts “Reason Only” and “Act Only” with a combined “Reason + Act” approach, in which a language model produces reasoning traces, takes actions in an environment, and receives observations. A separate planner-and-executor diagram shows a running context connecting a planner, thinking prompts, multiple executors, and an answer.

That framing puts tool use in the middle of the problem rather than at the edge of it. The model must not only generate reasoning, but use that reasoning to interact with an environment and incorporate what those interactions reveal.

Inference is presented as a new scaling frontier

Azalia Mirhoseini places inference after pre-training and fine-tuning in a three-stage scaling progression. The displayed slide calls inference “a new frontier for scaling,” while Mirhoseini introduces it as one of the three stages under discussion.

Inference scaling is presented alongside the broader agentic system rather than as a self-contained technique. Chowdhery’s account joins scaling at inference time to acting through tools and self-improvement: the system must be able to pursue work, receive signals from its environment, and use those signals to improve.

Several slides show possible feedback and optimization loops. One proposed optimizer architecture begins with generated steps for a math problem and includes an ORM labeled “P(R_i = 1),” an “Absolute Zero Reasoner,” self-play, and an ITAS optimizer for hyperparameter selection. Another contrasts standard RLHF with a constitutional-AI feedback loop: a model generates responses to red-teaming prompts, which proceed through critique and revision before a finetuned preference model.

These are not presented as a single settled recipe for self-improvement. They illustrate a recurring pattern: generate outputs, assess them through a feedback or evaluation signal, and use the result to revise or optimize later behavior.

The open problem is building agents that hold up over longer work

? aakanksha-chowdhery says the problems involved in building agentic systems are only beginning to emerge. The most concrete indication of the challenge comes from a task-suite comparison that distinguishes agency-oriented work extending for hours from single-step tasks sampled from SWE work.

Task suiteTask typeStated time rangeTask count
HCASTDiverse tasks requiring agency1 minute–30 hours97
SWAA SuiteSingle-step tasks sampled from SWE work1–30 seconds66
The displayed task suites distinguish longer agency-oriented tasks from short single-step work.

The contrast matters to the emphasis on planning, tool use, and self-improvement. HCAST is labeled as a set of diverse tasks requiring agency and spans work from one minute to 30 hours; the SWAA suite is explicitly single-step and ranges from one to 30 seconds. The slide also separates “task performance” from “time horizon analysis,” alongside human runs and time estimates.

That distinction suggests two different evaluative questions. Task performance concerns how a model performs on the suite’s tasks; time-horizon analysis is displayed as a process of finding a horizon length for each model and situating it against model release dates. The source does not report comparative results from that analysis. Instead, it presents time horizon as an additional dimension for evaluating systems expected to reason, plan, and act across many steps, rather than only produce a response on an isolated task.

Azalia Mirhoseini says participants will encounter open problems through which they can help push the frontier in agentic AI. Chowdhery names several: determining the right set of model capabilities, getting models to work with tools, and closing the loop for self-improvement.

In the arc of history, the problems that we see in building agentic systems are just starting to emerge and there's a lot of research questions to answer.

? aakanksha-chowdhery

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free