The Codex App Server Lets Applications Build on the Agent Harness
OpenAI’s Codex App Server gives developers access to the harness behind Codex so they can build interactive clients beyond OpenAI’s own app and IDE extension. Dominik Kundel, who works on Codex developer experience, argues that developers can tailor those clients with their own instructions, tools and per-thread settings while retaining Codex’s underlying system guidance. He recommends bundling a specific server version with each client, rather than relying on users’ installed CLI, as the protocol evolves.

The App Server exposes the harness, not just a chat window
Dominik Kundel describes the Codex App Server as a JSON-RPC-style protocol that wraps the Codex harness and lets other applications build directly on it. The same server powers OpenAI’s Codex app and VS Code extension, as well as third-party integrations including Xcode and JetBrains IDEs. Kundel said the protocol had also been used to put Codex inside Claude Code, where users can hand work to Codex or ask it to review code without leaving that environment.
The client and server exchange requests and events for initialization, configuration, threads and turns, along with streamed output. To make those exchanges concrete, Kundel built an inspector that intercepts traffic between the Codex app and the server. In his demonstration, he showed events arriving as a thread started, configuration was read and responses streamed back. The messages can look daunting at first, he said, but become more recognizable when viewed as actual requests and events rather than as a hypothetical protocol.
The server supports more than 120 client request types, with more being added. They range from initialization and thread management to goals, plugins, file search and file-system operations. Because the protocol exposes capabilities the first-party app itself uses, a client can do more than send prompts and display responses.
The comparison with the Agent Client Protocol (ACP) is about scope. Both let clients communicate with an agent harness, but Kundel characterized ACP as more generic. The Codex App Server also gives a client more control over specific parts of the Codex harness and agent configuration.
Both the harness and the App Server are open source under the Apache 2 license, Kundel said. Developers can fork them and connect the harness to other model providers if they offer a compatible Responses API. He also pointed to convenience settings for connecting to LM Studio or Ollama.
Customization works best as an extension of Codex
Dominik Kundel highlighted two ways to adapt the agent: developer instructions and tools supplied by the client. Developer instructions are appended to Codex’s system instructions. He recommended that approach over replacing the base instructions, which are available in the open-source repository. Appending guidance preserves the work that goes into tuning the system prompt for different models while letting an application add directions for its own use case. Kundel said the Codex app itself uses this to account for behavior specific to the app rather than the IDE extension or CLI.
A client can provide dynamic tools: functions the agent can call in its environment. These can be defined directly or provided through systems such as MCP. One option is to defer loading a tool. Rather than putting every tool description into the prompt from the outset, the client can make tools discoverable through a tool-search function. Kundel said this helps avoid filling the context window with tools the agent may not need while keeping them available when relevant.
The server also permits configuration changes per thread. Clients can select a model or model provider, and adjust sandbox behavior, including which files the agent may modify or whether it can make network connections. This lets an application set different operating conditions for different threads.
Authentication is another feature of the protocol. Kundel said an application can offer “Sign in with ChatGPT,” allowing users to bring their ChatGPT subscription and have Codex usage count against their own limits. A client can instead use API authentication. He pointed to T3 Code and similar tools as examples of applications using the subscription route.
Bundling the server makes the client’s capabilities predictable
Dominik Kundel said the App Server comes with the Codex CLI in the same binary. Developers who already have the CLI can use it to start the server. The Codex desktop app bundles the binary inside its application package and updates it with each app version. On Windows, the app includes both Linux and Windows binaries so that it can use the Linux version in WSL mode and the native Windows version otherwise.
Kundel’s advice was to ship a specific copy of the harness with an application rather than depend on whatever CLI a user has installed. The protocol changes over time, and experimental features can introduce breaking changes. Bundling the version a client was built against also means the developer does not have to rely on users updating their CLI before a needed feature is available.
That versioning choice is particularly useful because the server can generate TypeScript bindings and a JSON Schema for its protocol. Those tools make it easier to work with the messages, but the generated interfaces correspond to a particular server version. Bundling that version gives the client a defined set of expected capabilities.
The server can communicate over standard input and output or through WebSockets; which transport makes sense depends on the application. For developers who would rather not manage the server lifecycle themselves, Kundel pointed to the recently released Python SDK. It handles lifecycle management and provides convenience functionality for the ChatGPT sign-in flow.
Another route, he said, is to ask Codex to build a client from the App Server documentation. He described using it to create a Visual Basic client for a Windows-native app and said it largely worked in one shot. He also showed a request to build a macOS menu-bar chat app. The point was not that developers need to learn every detail of the protocol before starting, but that Codex can help implement a client when given the documentation and a specific target.
The choice between the App Server and codex exec depends on the kind of integration. codex exec runs Codex non-interactively and suits scripts, such as automating a pull-request fix in a CI/CD workflow. The App Server is for interactive, controllable clients: applications that need to show and manage the agent, or start multiple threads for work such as an evaluation harness or a goal-driven workflow. Rather than managing parallel codex exec calls independently, a client can use the server’s thread and turn events to start and coordinate work.
The interface need not resemble an agent chat
Dominik Kundel used Codex inside Doom to show how a client can expose its own application-specific capabilities to the agent. His setup used an Electron app and a WebAssembly Doom renderer; the game engine was implemented in C, and he had Codex modify the game file. A Codex interface appeared inside the game, rendered using the game engine, with Electron’s IPC protocol passing requests to the App Server running on his machine.
The client supplied dynamic tools for querying game state. In the demonstration, Codex retrieved the player’s armor and health, then listed weapons and ammunition by calling tools that returned that information. It could also run a shell command to report its current working directory. The game tools illustrate how a client can give the agent access to functions particular to its environment, rather than limiting interaction to a chat box. Kundel noted that Codex had its own workspace and that a user could, in principle, work on code from inside the game—though he did not recommend it over the Codex app.
Three implementation choices reduce avoidable friction
Dominik Kundel closed with three recommendations that follow from the protocol’s flexibility and its version changes. First, bundle the harness your client expects rather than relying on a user’s installed CLI. Second, expose application-specific functions as tools, but mark them as deferred so they remain available without unnecessarily crowding the context window. Third, start by adding developer instructions rather than replacing the full system instructions.
That last caution reflects a tradeoff. Writing a custom base prompt may look like the fastest path to a tailored agent, but it means giving up the prompt work OpenAI does to keep the harness performing well across models. Kundel recommended preserving that foundation where possible and layering application-specific guidance on top.
