Orply.

AI Gateways Let Teams Change Models Without Rebuilding Their Stack

Remy GuercioAI EngineerSunday, October 11, 20267 min read

As AI budgets face closer scrutiny, Remy Guercio of Tailscale argues that standardizing on one vendor is not a lasting answer to rising costs or AI lock-in. In his account, teams need room to change models, connectors, interfaces and execution environments without replacing the whole stack. An AI gateway, he says, can preserve that choice while giving organizations a central way to manage access, set model-specific budgets and inspect what agents do.

The answer to AI spend is not a single vendor

Remy Guercio describes the past year of AI use as “token maxing”: putting large context windows, MCP servers, command-line tools and coding agents to work finding and supplying as much context as possible. He says that combination has helped make agentic coding possible, and expects other agentic uses to grow.

But the approach has costs. Guercio says the period when teams could behave as though AI budgets were unlimited has begun to give way to closer scrutiny. In his experience operating an AI gateway, some sessions run for hundreds or thousands of messages without anyone clearing the context. A session can accumulate tens of millions of tokens while its cache stays near capacity, he says. That may sometimes be useful, but it can also mean spending heavily without a clear reason.

There is a second cost: tool chains increasingly bundle the model, agent and surrounding tools into one vertically integrated product. Guercio points to Claude Code and Codex as examples of products built across much of that stack. The more a team depends on one integrated setup, the harder it may be to change a model or workflow without changing the rest.

When finance asks why the bill is high, the instinct can be to standardize on one vendor. Guercio says he has heard that response in hundreds of discussions about AI usage and spend: if everyone uses the same tool, perhaps usage will be easier to control. He understands the appeal, especially when teams are under pressure to account for costs, but argues that consolidation is not a durable answer to lock-in.

Instead, he calls the next phase “ROI maxing.” He does not mean cost-cutting alone. Teams need to show that their AI use creates value, while making it possible to extend that use beyond a small number of expensive frontier models. If those models and their bundled ecosystems are the only available route, some uses may simply be too costly to expand across an organization.

Choice has to exist at every layer

For Guercio, the starting point is to separate an internal AI deployment into four parts: models, data connectors, interfaces, and the sandbox or environment in which an agent runs.

The model layer is the LLM or set of LLMs. The data layer is broader than MCP: it includes command-line tools, APIs, files and any other means by which an agent gets information. The interface is how a person works with the system—whether through chat, a terminal interface, an IDE, a phone or Slack. The final layer is where the agent runs: its sandbox or other execution environment.

These distinctions matter because teams do not all need the same setup. Guercio does not expect an engineering team to do all its coding in Slack, but says a marketing team might use a Slack thread to ask questions of company data. Choice is not a demand that every team build a bespoke tool at every layer. It is the ability to make a different choice when a team’s needs change.

That is what Guercio means by ROI maxing: maximizing opportunities for choice so that a team can change a model, connector, interface or environment without being forced to replace the whole system. An AI gateway, he argues, can provide a shared point of control across those layers. His case is not limited to Aperture; he says the broader value applies to AI gateways generally.

A gateway can centralize control without centralizing the tools

Aperture sits between agents and the services they use: LLM providers as well as MCP servers and other data endpoints. Guercio describes the gateway as a place to control and monitor usage for security, cost and operational understanding. Its logs, for example, can help explain why an agent failed to reach an answer or continued running when it should have stopped. He says less expensive models may be more prone to that kind of runaway behavior, and that logs let a user show the agent what happened and ask it to analyze the failure.

The dashboard shown during the talk put that monitoring proposition in concrete terms. For the displayed instance, Aperture reported 14 active users, 11,808,327,928 total tokens and an estimated cost of $13,247.20 over the last 30 days. Guercio presented the gateway as a way to inspect activity as well as apply access policies, not simply as a place to route API calls.

11.8bn
tokens reported for the displayed Aperture instance over 30 days

Aperture’s approach to identity is tied to Tailscale’s network. Guercio describes Tailscale as an identity-based mesh network built on WireGuard. Devices and services join a Tailscale network, or “tailnet,” and, in his description, connections over it carry the identity of the user or machine on the other side. A machine might carry a tag identifying it as a pull-request review bot; a connection might instead identify the person using Claude Code on a laptop. Where an organization uses Entra or Okta, Tailscale can also sync group membership, he says.

Aperture was built using tsnet, Tailscale’s open-source library for building this kind of application. Guercio says this lets Aperture see the identity of the user or agent connecting over the tailnet. In a typical gateway setup, he says, teams may exchange provider keys for a new collection of synthetic keys. Aperture’s described approach uses the identity available over the tailnet instead; it does not require another round of credentials for each user in the examples he gave.

For model access, an administrator configures providers in Aperture. Guercio showed a configuration including Anthropic, OpenAI, Vercel, Bedrock and Codex, among others. A connected user or agent can then receive access according to its identity and the organization’s rules, rather than managing a separate provider key for each tool. Budgets can also differ by model: Guercio gives the example of unlimited use for a cheap model and a very narrow budget for a frontier model.

That arrangement is intended to make provider changes less disruptive. The organization can put its provider credentials in one place and decide which models are available to which users, rather than tying access to one agent or one model vendor.

Connector permissions can move out of the agent harness

Aperture also acts as an MCP gateway. Administrators can configure which connectors are available and assign access according to teams or permissions. Guercio gives the example of allowing marketing to use Google Calendar and Gmail while giving engineering access to GitHub as well. GitHub access could, in turn, be read-only for one group and allow writes for another.

The connector configuration shown during the talk included GitHub capabilities such as reading issues and repository files, creating branches and submitting review comments. It made visible the distinction Guercio emphasized: the connector and its permissions are configured centrally, while agents can use the available integration through the gateway.

He notes that MCP has a standard, but says implementations differ, particularly in how authentication is handled. Aperture’s connector configuration is meant to make those integrations easier to manage centrally. Guercio also describes the gateway as able to handle authentication for standard APIs and expose them as MCP connections.

Once the connector and its authentication are managed in Aperture, a person can use it from Cursor, Claude Code, Codex or another agent framework by connecting to the gateway. In the setup Guercio described, the identity comes from the tailnet, rather than requiring the user to re-authenticate for each new tool. He argues that reducing these small points of friction makes it easier for teams to try another workflow, including one that might be cheaper or faster.

Interfaces can vary while the controls stay in one place

The interface can change too. Aperture includes a chat UI, but Guercio says its APIs can also support other interfaces, such as a Slack bot, a mobile app or a voice interface. Connectors are configured in Aperture rather than separately in each interface.

He also described a planned pass-through authentication feature for interfaces such as Open WebUI. His concern is that when an interface sits between a user and the gateway, requests can appear to come from the interface or system rather than preserving the person’s identity. The proposed feature would pass through user and machine identities so the gateway could retain the visibility Guercio wants. He presented this as forthcoming work, not a capability already available in the demonstration.

Sandbox choice is still an open part of the stack

Guercio says sandboxes are an area where teams already have substantial choice, in part because people differ on what a sandbox should provide. He expects Aperture to add an orchestration layer for sandboxes, bringing environment choices together with models and connectors. He presented that orchestration as planned work, not as an existing capability.

The wider argument is that a gateway can be useful even for an individual user, not just an organization trying to avoid vendor lock-in. If an agent runs in the background and does something unexpected, logs provide a way to inspect what it did and trace where it went wrong. Guercio’s example is an agent running on a tailnet: when its behavior is surprising, a user can return to the recorded activity and investigate rather than guessing.

The practical case for a gateway, then, is not only that it makes providers interchangeable. Guercio links choice to cost control and visibility: teams can set different budgets for different models, manage connector access, and inspect agent activity while trying other tools and workflows. That combination is his proposed route to proving value without making the first AI tool chain the only one an organization can use.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free