Gemini Managed Sandboxes Give Agents Persistent Cloud Environments
Philipp Schmid of Google DeepMind argues that AI agents should run in managed cloud sandboxes rather than on a developer’s machine. In his account of Google’s Gemini API, the Interactions API represents work as a sequence of steps, while the Antigravity agent runs code, uses tools and retains files in a persistent environment started with an API call. Schmid’s case is that Google can handle the agent loop and infrastructure, leaving developers to provide the request, tools and working context.

A managed sandbox takes the agent loop off the developer’s machine
Philipp Schmid’s central proposition is that an agent should have its own execution environment. With the Antigravity agent in the Gemini API, a request can start a cloud sandbox where the agent runs code, creates files, installs dependencies and tools, and calls APIs. Google manages the environment; the developer sends the request without first provisioning the infrastructure.
That changes who handles the work between model calls. In a conventional tool-calling setup, Schmid said, an application may need to parse a function call, execute it somewhere, and send the result back to the model. With managed agents, he said, the agent loop runs on the server inside the sandbox, so the developer can start with one API call rather than building that orchestration.
Schmid’s distinction: The Antigravity agent in the Gemini API uses the same underlying harness as the Antigravity product, but it is not necessarily the same agent. Different products can configure the harness with different tools and system instructions. Improvements to the shared harness can benefit both.
The environment is not just a place to run a command once. Schmid described managed environments as persistent: a later interaction can reuse the environment ID from an earlier one and continue with the files and work already there. That makes it possible, for example, to ask an agent to download data in one interaction and create a chart from it in the next.
Environments can also be shared among agents. Several agents can work in the same filesystem and pass context through the files there; a sub-agent can pick up work another agent has done. Schmid presented this as a way to make the environment itself part of an agent system’s working context, rather than having every handoff depend on messages sent back through a client.
With the managed agents, it's really just a single API call, and everything runs inside that remote sandbox.
The Interactions API represents work as steps, not just turns
The Interactions API is the interface Schmid presented for both Gemini models and agents. A request can contain a simple text input or multimodal data, and can include tools or an execution environment. He described the model and agent calls as using the same general interface, with the caller specifying what to run and what input to provide.
The API also gives developers options for managing context and long-running operations. If an application does not want to keep state on the client, it can pass the ID of a previous interaction with a follow-up request. The server combines the inputs so the model receives the earlier context. For work that takes time, such as video generation, an application can submit a request, receive an ID, then stream or poll for completion rather than hold an HTTP connection open.
Schmid’s data-model argument is that agent activity does not fit neatly into alternating user and model messages. A function result comes from the environment, not the user, and an agent may reason, call tools, receive results, and continue before returning an answer. The API therefore represents activity as a timeline of typed steps. A step might be a model output, a thought, a function call, or a Google Search call; status fields can indicate whether it is done, waiting, or in progress. Text, images, audio, and video go in content fields, while structured details such as a function name and its arguments have typed fields of their own.
That structure supports combined tool use. Schmid showed a small security-agent example that checks for a critical vulnerability affecting React. The request can use Google Search and URL context to find and read an advisory, then call a custom function whose fields capture severity, a description, and a source identifier. The developer defines the available tools and the function’s shape; the model decides which tools to use. The resulting structured information could then be saved in an incident system or database.
The same API pattern can return media as well as text. Schmid showed audio generation configured with a voice and image generation configured with settings such as aspect ratio and size. Those examples illustrated the broader point: an interaction can carry different input and output modalities, while its steps make the actions and results explicit.
Sources and files give the sandbox its working context
A managed environment needs more than compute: it needs the right material and instructions. Schmid said developers can provide sources such as Google Cloud Storage buckets, GitHub repositories, or inline files. The agent can also use skills stored in a skills directory; those are picked up and included in the system instruction. An existing AGENTS.md file is incorporated in the same way.
The AI Talk Radio demonstration showed how those pieces can define a more involved workflow. Given a topic, the agent could research it, write scripts for different participants, generate speech and music, create a thumbnail, and assemble a show. Schmid entered a prompt about whether pineapple belongs on pizza, and showed a completed example on the cereal-first versus milk-first debate. It was staged as a discussion, with a host and callers advancing different arguments, rather than as a single model response.
Schmid opened the applet’s files to show where that behavior came from. Its AGENTS.md described the workflow, and the environment loaded the file at startup and included it in the system instruction. Skills supplied procedures for particular tasks. One displayed Python script used the Interactions API to generate speech from text, giving the agent a tool it could use within the larger workflow.
The example’s practical point was that an agent’s behavior can be represented in files and configuration, not only in a bespoke application. Schmid said the applet’s files could be used to create an agent in a backend setting, and that developers could remix the AI Studio applet to adapt it.
The Agents API packages a prepared environment for reuse
The Agents API addresses a different problem from running a single interaction: sharing a configured agent with a team or users. Schmid described a named agent as a way to bundle a base agent, instructions, and an environment. Developers can define one from sources or from scratch, or use an existing environment as its starting point.
For the last option, an agent can first prepare an environment—for example, by installing the GitHub CLI, adding skills, or running checks. The developer can then create a named agent based on that environment ID. Later calls to the named agent start from the prepared setup rather than repeating that work. Schmid presented this as a way to share configured agents without writing infrastructure definitions or managing systems such as Kubernetes.
He also introduced a Gemini API command-line tool, distinguishing it from a coding agent or Gemini CLI. It lets developers or coding agents call the Gemini API, run prompts against models, generate images or speech, and scaffold, test, and deploy managed agents. The examples included initializing an agent, editing its AGENTS.md or tools, testing it with a prompt, and creating it for later use. Schmid pointed to dedicated skills for the Interactions and Agents APIs as additional help for coding agents working with them.
A visible environment makes the managed setup inspectable
In a separate AI Studio playground demonstration, Schmid asked the Antigravity agent to inspect its environment. After an initial setup of roughly seven to ten seconds, it ran shell commands and reported Ubuntu 24.04.1 LTS, eight virtual CPUs, and 16 gigabytes of memory, along with Python and pip. He then continued with a follow-up question about whether Node was installed.
The playground also offered a way to download the environment, and Schmid said an API was available for doing so. Together, the file-based radio agent and the environment-inspection demo show what the sandbox model is intended to provide: a place where an agent can execute a workflow, retain working state across interactions, and leave files a developer can inspect or retrieve.


