Orply.

The Best Agent Workflows Start With Context, Guardrails, and Human Judgment

Charlie HoltzAI EngineerSunday, September 27, 20267 min read

Charlie Holtz, co-founder and CEO of Conductor, argues that teams using coding agents should build around human judgment rather than treat software development as a factory line. His principles include staying close to new tools without chasing workflows that will soon become standard, protecting code and agent instructions with strict review, and giving agents access to organizational knowledge and persistent workspaces. The aim, he says, is for people to conduct teams of colleagues and agents, not simply manage automated production.

Use the frontier to find work, not to avoid it

Charlie Holtz describes Conductor as a desktop app for managing multiple coding agents in one interface, rather than juggling separate terminal windows. Building it has given him a close view of how experienced builders work: what they try, what they protect, and where they choose not to spend time.

His first principle is to stay near the frontier: try new tools and workflows as they become available. For a startup, that can reveal products the team would otherwise not have thought to build. Holtz says Conductor grew out of that kind of discovery. His company had been building a different app, Chorus, but the team were heavy users of Claude Code. They began cloning their repository repeatedly, discovered work trees, and gradually built Conductor as an internal tool around the workflow. Holtz says they would not have arrived at that product if they had not been keeping up with new ways of working.

For builders inside larger organizations, keeping up matters for a different reason. Holtz says information about useful workflows once spread through professional networks, but the pace of change now makes that route too slow: relying on what filters through can leave someone three to six months behind. He argues for being the person in the organization who knows what has changed and what might be useful.

But “near” is important. Following every new technique can turn into a way of spending time on the workflow instead of doing the work. Holtz calls that “midwit memeing” and offers a test: ask why a promising workflow is not already the default.

If a method works broadly for everyone, he says, a major model provider may eventually build it into the standard harness. In that case, a team may be better off waiting than investing heavily in its own optimization. The exception is what Holtz calls “real alpha”: information about a team’s users or codebase that the models may not have. “Why isn’t this workflow the default?” is his way of distinguishing a general improvement that may become standard from an advantage particular to the team.

Conductor, for example, is a chat app where long conversations need to render quickly. Holtz says the team spends time optimizing its React queries because that performance requirement is specific to their product; they are willing to make sacrifices elsewhere in the codebase to meet it. The distinction is not between using new tools and avoiding them. It is between adapting to a general trend and investing in a workflow advantage grounded in knowledge the model lacks.

Speed depends on knowing where not to accept slop

Charlie Holtz rejects the assumption that a team using many coding agents should accept enormous, lightly reviewed changes. At Conductor, some parts of the codebase are handled loosely, while others are “slop-free zones”: areas that require strict human review.

Holtz says Conductor has had to rewrite its app a couple of times after failing to be careful about these boundaries. One explicit example is the migrations file: any change to it in the company’s CI process requires a human reviewer.

The same principle applies to the information agents rely on. Holtz says the team assumes Slack messages are human-written, and puts substantial effort into its documentation, CLAUDE.md files, and agent skills. He has seen other strong builders do the same. These files matter because their contents are loaded into an agent’s context when it starts work. Holtz compares that to having the chance to whisper something to a new intern every time they sit down: the recurring instruction deserves careful thought.

Feed agents the organization’s own context

Charlie Holtz says Conductor’s internal agent, which the team also calls the CIA, is intended to collect information about what is happening across the organization. When a Slack message is sent, the agent picks it up and saves it to a Postgres table. Bug reports from Discord go into the same system, as do recordings of meetings.

Holtz calls this “feeding the beast.” He says agents working in a particular company need as much context as possible about how that company works. The team’s approach is to put information in a centralized place that the agent can access. A tweet shown during the talk put the idea more tersely: “the unreasonable effectiveness of just putting everything in a database and handing an agent a SQL tool.” Holtz says the tweet sums up the approach: give an agent a SQL tool and let it work with the database.

The point is not simply to give an agent more material. Holtz’s examples are the everyday records of how the organization operates: messages, bug reports, and meetings. Centralizing them is his proposed way to make that company-specific context available to agents doing work for the company.

Persistent workspaces change how teams coordinate

Charlie Holtz’s next principle is to give agents room to work: a sandbox where they can explore the codebase, tackle difficult tasks, and keep running after a developer closes a laptop. He argues that this will matter more as models improve, run for longer, and become more numerous. If agents are confined to a local machine, he says, they will not be as effective as they could be.

In a demonstration of a new version of Conductor, Holtz shows workspaces running in cloud sandboxes. Previously, he says, each task was built on a Git work tree. In the cloud version, an agent can keep working after the developer closes the laptop. Holtz frames the sandbox as more than a place to run an agent: it also makes it possible to share workspaces with teammates and start work through an API.

The shared workspace changes the unit of collaboration from an individual agent session to work that colleagues can inspect and discuss together. Holtz scrolls through work associated with different teammates, opens a colleague’s workspace, reviews a change, and asks, “can we actually use tabs not spaces.” The interface shows his message and a typing indicator for his colleague. Holtz says the workspace is shared in real time, so teammates can see and respond to one another while work is underway.

He argues that collaboration matters because great things are built by teams, not individuals. As models improve, he expects teams to attempt more ambitious projects, which will require more people and agents. The demo’s practical example is a teammate able to review changes and leave a request in the same workspace, rather than treating the agent’s work as isolated from the rest of the team.

Cloud workspaces also let people initiate tasks away from the development environment. Holtz shows an OpenClaw agent he calls Lord Cranedon, which has access to a Conductor API. From a phone, Telegram, or Slack, he says, he can ask it to create a workspace. In the demonstration, he sends the request, “hi, can you create a new workspace for me that makes all the buttons blue.” The agent creates a workspace called “All buttons blue,” which is still being prepared when Holtz shows it.

Together, the examples illustrate what Holtz means by “free-range agents”: agents that can keep working in a sandbox, collaborate with people and other agents, and be given ways to initiate work. The demo shows one specific possibility: a remote request can start a task that continues while the person is away.

The human should conduct, not manage a production line

Charlie Holtz’s final principle is a rejection of the “software factory” metaphor. He says he dislikes the term because it suggests a future organized around automation and production lines. Automation can make work more efficient and increase what people can produce, he acknowledges, but he does not want developers reduced to factory-line managers pushing buttons to make agents produce the next feature.

His alternative is a person conducting teams of agents and colleagues: directing work, moving between groups, and zooming in on details when needed without managing every action. Holtz wants software to feel “human and crafted,” and says the people building these tools have a responsibility to make them good for humans—to help people feel capable, in flow, and engaged in making things. He recalls the language of “feature factories” from a decade earlier as a warning against treating production volume as the whole point.

The images in his talk make the contrast literal: a Yu-Gi-Oh! card titled “Dark Factory of Mass Production,” with a cartoon factory assembly line, followed by a photograph of a conductor leading an orchestra. Holtz says he does not want to be a line manager in the “dark factory.” He wants the latitude to direct the larger effort and the option to turn his attention to particular details.

He describes the feeling he wants as designing alongside a team of people and AI agents, like Steve Jobs designing the Mac with a team around him. In Holtz’s orchestra metaphor, the person remains at the center of the work, shaping it while agents and colleagues contribute.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free