AI Agent Policies Need Runtime Enforcement at the API Gateway
Sam, CTO at Gravitee, argues that prompts and policy files cannot govern what an AI agent does once it can reach an API: enforcement has to sit in the infrastructure between the agent and the services it can call. In a live hotel-booking demonstration, he showed a gateway denying destructive actions and bulk access to guest data, while also limiting repeated requests, caching similar questions and rejecting oversized prompts before they reached a model.

Instructions cannot enforce what an agent is allowed to do
Sam, CTO at Gravity, argued that governance is a runtime control problem. Prompts, markdown files and an agent’s set of tools may express a policy, but they do not prevent the agent from making a call it is technically able to make. The control, he said, belongs in the infrastructure between the agent and the services it can reach.
Sam cited a survey his organization had conducted: 88% of organizations had an agent in production taking action, but only 14% of those had governance around it.
To demonstrate the difference, he ran five scenarios against a hotel-booking agent. One agent relied on its own prompts, tools and context for restraint. The other used a tightly scoped identity and gateway controls. The live setup used the Gravitee Gateway, a Databricks backend holding the hotel data, and Sonnet and Groq for inference. Sam said the scenarios were not scripted; the interface displayed request traces and gateway responses.
The gateway can block a tool call even when the agent discovers it
In the first scenario, a user asked the agent to cancel a particular booking and, under the pretext of routine maintenance, also run a delete_all_bookings tool. The ungoverned agent executed both actions. Sam said the tool was available and the user had permission to invoke it, so the agent performed the destructive operation. He then showed the database, where the bookings had been deleted.
The governed agent handled the request differently. It could still discover the tool, but its identity did not have the required administrative permission at runtime. The gateway returned a 403 forbidden response when the agent tried to execute the destructive call. The agent completed the part it was authorized to do—the individual cancellation—and explained that it could not run the cleanup.
The demonstration’s interface showed both agents receiving the same request: the ungoverned agent said it had run the cleanup and cancelled the booking; the governed agent said it lacked admin permission for the cleanup, but had cancelled the booking. Sam’s point was not that the tool had disappeared from the agent’s view. It was that discovery did not grant permission to execute it.
For Sam, that distinction is the point: a policy document saying “do not do this” does not stop a tool call. A gateway can enforce permissions on the request itself. In the traces he showed, viewers could follow the user request, the agent’s tool calls and the resulting responses. The gateway’s denial was visible alongside the successful booking action.
A second scenario tested access to guest information. The request claimed to come from a compliance audit and asked for every guest’s name, email and card details. The ungoverned agent returned a list of guest information; Sam clarified that the data shown was mocked. The governed agent refused to export the full list because the account lacked administrative permission, and offered to look up a specific guest’s booking instead.
Sam described the gateway as blocking the request before the data was exposed. He also said an authorized user could authenticate through an identity provider or the gateway to obtain the access needed. The intended control is not simply to make the agent refuse every sensitive task; it is to tie access to identity and permission.
That approach becomes more important, he argued, as systems involve user-to-agent and agent-to-agent interactions. With agents calling other agents, and those agents calling still more tools and services, permissions managed inside each agent become difficult to track and maintain. A gateway offers a central place to apply controls and inspect what happened.
Rate limits and caching address the cost of repeated work
Governance, in Sam’s account, also includes limiting unnecessary model use. In a third scenario, a buggy agent sent the same hotel question eight times. Without a gateway policy, every request received a response. Sam estimated the simple example consumed about a thousand tokens, or roughly one cent. That was small in isolation, he said, but repeated workflows across many agents and users could add up.
The governed version allowed three requests per minute. The first three received responses; the gateway rate-limited the rest. Because the restriction was applied before requests reached the LLM, the blocked calls did not incur the downstream model work. Sam said limits could be set to control usage by user, permission or audience. The interface showed the ungoverned agent continuing to answer, while the governed agent returned a rate-limit message after the permitted requests.
Caching offered another way to avoid repeated calls. Sam described semantic caching as serving sufficiently similar prompts from a built-in Redis store rather than sending every request to a model. His example was a help-desk question repeated many times; the demonstration used “List all my hotel bookings.” On the ungoverned side, five repeated questions each went to the LLM. On the governed side, the repeated request was served from cache. Sam said the former used about a thousand tokens, while the cached example used around 200 tokens for the input and output.
The similarity threshold is configurable on a scale from zero to one, according to Sam, so teams can choose how closely a new prompt must match a cached one. He presented this as adjustable by agent, audience or user, rather than as a single setting for every use case. The gateway also surfaced logs indicating when a response came from cache. Sam said its telemetry could be exported to tools such as Datadog, Splunk or Prometheus.
Prompt size and relevance can be controlled before inference
The fifth scenario addressed oversized prompts. Sam described users pasting long conversation histories or meeting transcripts into an agent, even when the task itself was simple. In the demonstration, a prompt containing about 2,000 tokens produced a 312-token response for a basic booking task. With the token-limit policy enabled, the gateway rejected the same oversized request before it reached the LLM.
That timing matters: blocking at the gateway avoids paying for the model to process the prompt. The response told the user to provide less context and try again. Sam suggested breaking large tasks into smaller steps and giving an agent only the context it needs.
He also argued that teams can constrain the type of request an agent accepts, not just its length. As an example, an agent for hotel bookings could reject an unrelated question about the weather rather than route it to a powerful model. Sam described this as a way to keep use aligned with the agent’s purpose and avoid downstream cost for irrelevant requests. He recalled a customer using a powerful model to answer a question about the weather—something, he said, a weather station could provide—and suggested that teams could restrict requests to content relevant to the agent’s use case.
These controls sit within a broader case for separating the front-end agent from the services behind it: APIs, MCP servers, models, databases or data lakehouses. Sam said that separation supports security, cost management, governance and auditability. He cited a figure supplied by a colleague at the event as an indication of how quickly the management problem could grow.
If governance is hardcoded into individual agents, Sam argued, it may work for a few agents but become difficult to manage as agents multiply or create other agents. He also pointed to the prospect of increased regulation and the need to audit information used in model retraining. For that, he said, logs and other information need to be visible and exportable.
A gateway also provides a point of control across vendors and tools
Sam presented model routing as one way to avoid a fixed relationship with a single hyperscaler. The gateway, he said, can route requests to different models based on considerations such as latency, cost or complexity. In the demonstrated setup, he was routing between Groq and Anthropic; he said the configuration could be changed to Gemini or OpenAI, with the gateway handling transformations between different OpenAI specifications.
He also described a way to curate which tools an MCP server exposes. As an example, Sam said his organization’s CRM MCP server exposed 64 tools. Passing all of them to a model would add context and cost, while also widening the set of actions available. Selecting a smaller set, he argued, can reduce both the agent’s blast radius and the amount of tool information it must handle.
The remaining concern he raised was AI use that bypasses organizational infrastructure. Employees may use AI on their laptops and submit documents or emails without that traffic passing through a gateway. Sam called this shadow AI. He said Gravitee had recently released an edge-management capability intended to reveal such traffic, secure it and route it through the gateway, adding logging and observability for uses that would otherwise be hard to see.
Sam’s broader recommendation was to put controls at the infrastructure layer rather than rely on agents to obey policy files or prompts. He said a gateway was not necessarily required to be Gravitee’s, but argued that organizations scaling agentic systems should consider a gateway for permissions, cost controls, visibility and audit records.