AI Gateways Aim to Pair Model Choice With Runtime Controls
Remy Guercio of Tailscale argues that teams need room to change models and tools, while Gravitee CTO Sam demonstrates a gateway denying agent requests that lack permission. The shared control point can route traffic, enforce policies, and log activity, but only for requests it can see—and its reach depends on identity and configuration.

Why a gateway enters the architecture
AI teams face two different pressures that increasingly meet at the infrastructure layer. One is economic: Remy Guercio of Tailscale says agent sessions can run for hundreds or thousands of messages, accumulating tens of millions of tokens, while model-and-agent bundles can make it harder to change providers or workflows. The other is operational: an agent with access to a consequential tool may be able to call it even when its prompt says not to.
Guercio’s proposed answer to cost and lock-in is not to standardize on one vendor. He argues that teams should be able to change models, data connectors, interfaces, or execution environments without rebuilding the rest of the system. Sam, CTO at Gravitee, approaches the same architectural question through authorization: a policy written into an agent does not prevent a technically available API call. In his view, enforcement belongs between the agent and the services it can reach.
A gateway offers a shared mediation point for both problems: routing requests, applying access rules, and recording activity. But that is a proposal about where controls can sit, not proof that costs will fall or misuse will stop. The gateway can act only on traffic it sees, and its protections depend on how access, identity, and policies are configured.
Choice across the stack, controls at the call
Guercio divides an AI deployment into four layers: models, data connectors, interfaces, and the environment where an agent runs. The point of the distinction is practical. A team might want to switch models, give one department access to different connectors, or use Slack rather than an IDE without replacing every other component. He calls this “ROI maxing”: preserving room to choose while managing access and usage centrally.
The gateway is meant to centralize control without requiring every team to use the same tools. Guercio described model-specific budgets, centralized provider credentials, and connector permissions as available capabilities. He also discussed pass-through authentication for interfaces and orchestration across sandbox environments as planned work, not features demonstrated as already available. Flexibility, in this account, is an architectural aim; not every layer is equally mature.
Sam’s hotel-booking demonstration shows the other side of the design. An agent could discover a delete_all_bookings tool, but the gateway denied the call because the agent’s identity lacked administrative permission. The agent could still cancel an individual booking it was authorized to change. In a second scenario, the gateway blocked an unauthorized request for a bulk export of guest names, emails, and card details; Sam said the displayed data was mocked.
The distinction is important: the control was not that the agent had learned to refuse. The gateway made an authorization decision on the request. Tool discovery did not itself confer permission to execute.
| Concern | Gateway mechanism described | What it does not establish |
|---|---|---|
| Choice and lock-in | Route among models and connectors | That switching will be frictionless in every setup |
| Unauthorized actions | Check identity and permissions at request time | That every relevant request passes through the gateway |
Governance includes economics and observability
Central controls can address costs as well as permissions, but those are separate jobs. Guercio described budgets that could vary by model—for example, broad access to a cheaper model and a narrow budget for a frontier model—and logs that could help teams investigate why an agent behaved unexpectedly. He presented one Aperture instance with 11.8 billion tokens over 30 days, a figure that illustrates the activity a gateway dashboard may expose, not a general measure of AI use.
Sam’s examples focus on limiting work before it reaches a model. In his demonstration, a rate limit stopped repeated requests after three per minute; semantic caching served similar questions without sending each one to an LLM; and a token-limit policy rejected an oversized prompt before inference. These controls can reduce repeated or unnecessary model work. They do not replace permission checks: a rate limit can slow a destructive call, but does not determine whether the caller is allowed to make it.
The demonstrations also make observability part of the proposed control plane. Sam showed request traces and gateway responses, while Guercio argued that logs could help explain an agent’s actions and failures. Those records may aid investigation, but their usefulness depends on which services and identities are visible to the gateway. Neither account establishes how these controls perform across organizations at scale.
The boundary of the control plane
The strongest limitation is coverage. Sam explicitly raised “shadow AI”: employees may submit documents or emails to AI tools outside organizational infrastructure, leaving those interactions beyond a gateway’s ordinary view. Guercio’s emphasis on preserving choice raises a related design problem. More options across models, connectors, interfaces, and execution environments can help teams adapt, but each path needs to remain governed if central policy is to be meaningful.
Identity and configuration are equally consequential. Sam’s denial worked because the agent had an identity without the required administrative permission. Guercio’s model depends on connecting activity to user or machine identity and assigning access accordingly. A gateway can make those rules more consistent when traffic passes through it; it cannot make an overly broad permission narrow, or compensate for a missing identity mapping.
The two accounts therefore meet at a practical tension rather than a settled result. Organizations want to change models and tools without rebuilding their stack, while also imposing dependable limits on what agents can do. Gateways are being presented as a way to serve both goals through shared routing, authorization, budgets, and logs. The hotel-booking scenarios show how a runtime denial can work; the architecture argument explains why teams might want a common control point. Neither demonstration, by itself, establishes effectiveness across a full production environment.



