AI agents should run in their own managed cloud sandbox rather than on a laptop, which the Gemini Interactions API provides with a single API call — server-side, stateless from the developer's perspective, persistent across turns, and shareable between agents.
AI Talk Radio is an AI Studio applet built on managed agents that generates a talk-show radio program from a single prompt. Its agent researches a topic, writes the script, and produces music, speech, and a thumbnail. The agent is implemented as an AGENTS.md file and accompanying skills that can be copied into the backend. The applet is free to use with daily usage limits.
Antigravity agent is a Google DeepMind agent harness for local use and managed remote sandboxes through the Gemini API and Google AI Studio. A single Interactions API call can provision a secure, Google-hosted Linux sandbox identified by an environment ID, where the agent uses Gemini 3.8 Flash to plan, execute Bash, Python, and Node.js code, observe results, and repeat until the task is complete. It can install packages, run tests, build applications, manage files, search Google, fetch web pages, and call custom functions or remote MCP servers. The harness supports reusable skills, configurable tools, multi-turn interactions, streaming, automatic context compaction, persistent sandbox files, synchronous hooks for intercepting and validating code-execution and filesystem operations, and sources loaded from GitHub, Google Cloud Storage, or inline files. Managed environments provide a credential-injecting proxy and named agents, allowing the same workflow to run locally and in the cloud.
Google's Agents API defines and packages reusable AI agents with a base model, instructions, and a preconfigured environment. Agents run in managed cloud sandboxes rather than on the user's laptop; the environments can be shared with users or teams and accessed through an API call.
Gemini API CLI is an experimental command-line interface for the Gemini Interactions API and managed-agent platform. It can run model interactions and managed agents, generate or analyze text and media, transcribe audio or video, manage files, models, agents, environments, triggers, and webhooks, and stream results in formats including JSON, YAML, and TOON. Commands expose machine-readable usage information and request schemas, while --dry-run validates inputs and previews redacted HTTP requests without contacting the API. Authentication can use API keys, OAuth access tokens, an operating-system keychain, or a configuration file. The repository describes the CLI as beta software with possible breaking changes, generated through Speakeasy; its software is licensed under Apache 2.0, and the repository states that it is not an official Google product.
Gemini API managed agents are configurable agent harnesses from Google that let an application provision a Linux sandbox with a single API call. The agent can reason, execute code, manage files, and browse the web autonomously; developers can customize it with instructions, skills, data, and external tools, and package reusable agents through the Agents API. Each environment is isolated at the operating-system level and supports persistent state, although inactive environments are eventually deleted and virtual machines can spin down. Managed agents are in public preview and use pay-as-you-go pricing based on model tokens and tool usage.
Google's developer API for building applications with Gemini large language models. It provides access to models like Gemini 2.0 Flash and 2.5 Pro, and also supports Gemma open models. Developers use it for tasks such as custom model testing, retrieval, memory, and tool-calling systems.
The Gemini Interactions API is Google's unified API for calling Gemini models and specialized agents. It maintains server-side interaction state through interaction IDs and the `previous_interaction_id` parameter, while also supporting stateless requests. Each Interaction resource contains a chronological sequence of execution steps, including model thoughts, tool calls and results, and final output. The API supports text and multimodal generation, structured and strongly typed outputs, built-in and custom tools, tool orchestration, observable execution steps, background execution for long-running tasks, and chained image, video, and audio generation using shared context. Google documents it as generally available and recommends it for new Gemini API projects.
Gemini TTS is an AI text-to-speech service accessed through the Gemini Interactions API. It generates spoken audio from text using the API's audio response modality.
Google AI Studio is a web-based development platform from Google for exploring and evaluating Gemini models, developing prompts, and turning natural-language ideas into code and web applications. It supports integrations including Firestore, Cloud SQL, authentication, and database schema creation, and was used to create a mobile-responsive speaker-notes website.
A built-in tool for Google's agent API that a model can call autonomously during an agent run; the supplied video describes it in a CVE-agent demonstration.
An image-generation model from Google, exposed through the Gemini Interactions API. The video describes it as a fast, low-cost model priced at about four cents per generated image.
Omni Flash is an AI model available through Google's Gemini Interactions API for programmatically creating and editing videos and building generative-media applications.
Searchable transcript of Why AI Agents Should Have Their Own Sandbox — Philipp Schmid, Google DeepMind — AI Engineer (20:42). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:12 So, hi everyone. Uh, welcome to why agents should have their own sandbox. Uh, I'm Philip. I'm part of the Google DeepMind team, mostly working on AI studio and Gemini API. Uh, we have a small crew, so I try to keep the boring stuff short and like the exciting demo stuff long. And we can also ask some questions I guess at the end. Maybe you need to sit a little bit further at the front so I can understand you.
00:33 Um so I want to start with a little bit of a background. I'm not sure if you have seen that slide but that's a very popular slide when you do like Google presentations um to show a little bit um what we have achieved in the last two year with like the Gemini models but also our open models with Gemma and the slide is already outdated since yesterday.
00:54 So yesterday we shipped Nano Banana Light uh which is our new image generation model. Super fast uh sorry super fast uh super cheap. Um it's like 4 cent per image and then also we shipped Omni flash on the interactions API. So you can now programmatically use Omni to create videos to edit videos to yeah build gen media applications. And to give a little bit of a brief history, I assume everyone has heard about agent has heard about large language models.
01:24 But to really quickly reflect where we are coming from. So I think around 2018 19 with the first bird models and GBT models, we were super excited because we could continue sentences, right? We were able to generate poems without uh with providing like a first few words and then the model coherently continue the text which was nice but not very useful or helpful for any one of us to build something.
01:50 Then we we moved into the the instruction uh way. Before we had pre-trained models which could continue text and then with instruction following we basically trained models to answer questions more or less. So we told ask a question and the model correctly responded. Um and then with with chat GPT um everything changed a little bit. I think most of the people probably got started with that and we went into this user model turnbased conversation where we had a user input, a model input, a user input, a model output.
02:18 Then afterwards we combine that with function calling because we wanted the model to take actions for us. And now we are in like this black box with goals, thinking and functions where we no longer just prompt the model. We basically give a goal to the model, tell them how to verify its goal or how to like deploy it and like hopefully have it running for multiple hours or even days and come back with the results we want to achieve.
02:43 Um, so if you haven't used Gemini API or AI studio before, you can go to AI.studio or toi.dev. So very quick, very short URLs. H, get your own API key. So if you have your computer with you and you if you want to follow along a little bit later on like trying out a few things we have shipped recently you can do that. Um no setup required can use your Gmail account.
03:04 Um can start start experimenting. If you want to use bigger models more usage enter your credit card and we improved that a lot. So you don't need to leave AI studio anymore to go into Google Cloud. You can do everything inside of AI Studio. You can set up your own budgets that nothing gets out of control. And I want to start about or talk about the interactions API which is our new uh API for Gemini models and Gemini agents.
03:30 And on the right side sorry left side we have some code snippets. So if you never have heard about the interactions API you can think a little bit about that's our answers to like the responses API from OpenI. It is much more developer focused, much more simpler, much more web developer friendly. And you can already tell in like the two snippets here uh to call a model like Gemini or to call an agent here like the anti-gravity agent, it's like literally the same interface.
03:57 And this works for all of the modalities Gemini offers. So you have a model or an agent, you have an input. The input can be a basic string or it can be multimodal data. And then you have tools or environment which uh I'm very excited and why which we are going to talk about later why every model or agent should have its own box. Uh new with the interactions API is you have serverside state management.
04:22 So if you don't want to manage the state on the client you can hand it over to the server. We have long running operations. So async requests basically you start something. If we talk about omni video generation takes quite a long time and you might not want to keep your HTTP connection open. So you send a request get back an ID can set up streaming or polling to wait uh until the the generation is completed and then um it comes with a lot of simplifications for building agents and tool use.
04:52 So you can combine built-in tools like Google search uh with custom function calling and to look at some code examples. So for serverside state it looks very similar if you have used the responses API we call it interactions because we just not get a response anymore. We interact with models and agents. You send an input and if you want to do a follow-up turn you provide a new input and basically just the ID from the previous turn and then on the server we um concatenate all of like the different inputs and make sure
05:23 that the model really reh sees what has been done so far. And this goes for all modalities. So for audio generation you have a model. We have our Gemini TTS model. We have our input and since we want audio back not text you define a response modality with audio and some generation configuration which was you for example want to use the same pattern works for image generation.
05:46 So our image generation model uh nano banana here our input some generation uh generation config and then you get back your image and for tool use uh also much better I would say than previously. You just define your tools with like Google search URL context or your own custom function and then you send your request and the model decides on it on its own.
06:11 do I first need to like run Google search or call the function and all of the code you see here is basically enough to create a very simple small CVE agent to detect if there are like any CVE um for react in this case and then like it uses Google search it uses URL context to visit those websites and then function calling to basically give structure to it so you can like save it in your database or look it up for example um and something new and maybe different to other model providers is that we moved away from this
06:46 user model term principle. So if you build an application with another LLM provider or with open models, you have always this uh user model user model turn concept. So you as a user provide an input and expect the model to return an output. This worked until like one or two years ago I would say where we just got back responses from those models. But with agents and with reasoning, all of that changed, right?
07:14 Models now think for a very long time where they might expose some thoughts directly or might just expose a signature. Function calling is not very uh a user role, right? With normal or like with previous APIs, you would need to use the user turn to return some function output, which for me always felt a little bit weird because the output or the new input is coming from the environment and not from the user.
07:39 So that's why we moved away from like those turns into steps or like to um like synchronous timeline of steps where every action or generation is its own type which makes it very easy to extend to new features and capabilities which might come in the future but also makes it still like very simple to use because you really define what it is done. So if you build a function calling agent, you return a function result and not like a a user role user with some content which might be a function result or not.
08:09 And at Google IO we ship the anti-gravity agent uh which comes with its own sandbox. So the anti-gravity remote agent inside the Gemini API is powered by the same agent harness which is used with the anti-gravity product. So if you have used the IDE or the new anti-gravity 2.0 O or the CLI, you might have used that harness already. Very important here.
08:34 When we talk about harness, we don't mean the same agent. The harness is basically the the back end of our agent, but it might be different between different products. So the anti-gravity IDE is a coding agent. So it might have a little bit of different tools or a little bit of of a different system instruction than the anti-gravity agent we have in the Gemini API, but it's still the same harness.
08:55 Meaning if we improve Gemini to become better with that harness, we benefit inside anti-gravity, but we also benefit inside of the Gemini API. And with that new agent, we also introduce this uh environment field. And an environment here remote is basically for that API call, your agent gets a sandbox managed by us. Meaning if you send a request, our back end starts a cloud sandbox where the agent can run code, can create files, write files, um, install dependencies, install your own tools or maybe call your own APIs
09:33 and it's all managed and stateless on our side, meaning that you just send the API request. You don't need to manage the infrastructure and those environments stay persistent. So if you have multi-turn conversation where the agent um you set user input I don't know like um research the latest tech news and create a file and then you want to do like a follow-up interaction you can like use the same environment from your previous turn with that new interaction but also this means that you can share environments between
10:05 different agents. So if you have like three agents working inside the same environment, they can share context through the file system or you can like even use sub agents to basically pick up a work uh another agent has been done and then of course environments are very good but the model needs the right context and the right information or the right tools to work inside those environments.
10:27 So you can provide sources uh sources can be um GCS buckets, GitHub repositories or inline files. It's very useful if you want to integrate skills. So the anti-gravity agent is skills or file system native I would say. So if you store skills inside the skills directory, it will automatically picked up and uh be pro uh included into the system instruction.
10:52 But also if you have an agents MD file from your coding agent or from your existing agents, it will also be automatically appended to the system instruction. and then um you can create an agent because like sending a request is nice and having all of the configuration stateless is also very good but what if you want to share it or what if team one builds an agent team two wants to use um therefore we created the agents API where you can define the base agent so what's the underlying brain so to speak you can define
11:22 instruction and then also the uh environment um and the environment can be a raw environment like with sources or you and uh use an existing environment. So you can think about I have some kind of a scaffolding to do where I need to install like the GitHub CLI uh add some skills and maybe do some like sanity checks. I can use the agent to create its own environment.
11:44 And after the agent is done, I get back an environment ID. I can use that environment ID to create the new agent. And every time I call this new agent, it will start from that already pre-loaded, preconfigured environment. And so it's like very easy for you to use agents to build their own environments and then like to create agents out of them which you can share with your users or with your customers without writing any Terraform without thinking about Kubernates, microservices, firecracker or anything else.
12:19 And to make it easier for you and for your coding agents, we created a CLI for the Gemini API. So it's it's not a coding agent. It's not Gemini CLI. It's really an a CLI that your agent can use to call the Gemini API. Meaning it can call Nano Banana or Omni if you want to generate some assets. But you can also use it to create those agents, iterate on those agents, and improve those agents.
12:43 The QR code should bring you to like the GitHub repository and then also we have dedicated skills for working with the interactions API and the agents API um if you want to to play around with it later. And now let's let's do some some demos. So I'm in AI Studio inside the AI talk radio. If you have seen or attended the Google IO developer keynote, you might have seen uh this already.
13:16 Um that's just a simple UI on top of managed agents. We manage agent. We are going to look at the agent in a little bit, but it's basically a talk show radio and that talk show u agent has access to all of the different Gemini models. So it can generate a thumbnail, it can generate um speech, it can generate music, it can do research and it's basically a one short radio generator and we give it a try.
13:43 And I want a talk show radio on whether pineapples belong on a pizza or not. And it's free to use. So if you go to the applet later uh you can like have free daily usage to create create your own um AI talk show radio and what happens is like first step we start our environment and then the model starts thinking on what it needs to do and then the first step is like it installs our Python dependencies and it reads our skills and then it continues to work.
14:20 Since this will take a little bit, I already ran a talk show radio on whether we should put in cereal first or milk first. And the very interesting part about a talk show radio, it's not like a podcast you listen to. It's really you have multiple um people like reasoning and like sharing their opinions on it and we give it a quick listen. Is there a correct physical order to preparing your breakfast bowl or are we flirting with culinary chaos?
14:49 [music] Welcome to AI Talk Radio. I'm Paul broadcasting from our London studio. Today we are exploring the great breakfast debate. Cereal first versus milk first. We will objectively break down the pros and cons of both methods. Let's go to our lines. We have David calling from Boston. David, you're on the air. >> Yeah, thanks Paul. Look, it's got to be cereal first.
15:17 It's basic logic, right? Cereal is the independent variable. You pour it, you see the volume, then you add milk to compliment it. Plus, like if you pour dry cereal into a bowl of milk, it splashes everywhere. It's a total mess. And the mashed potato analogy, you wouldn't put gravy in a bowl and drop a lump of potatoes on top, right? A striking analogy, David.
15:41 Cleanliness and portion control are certainly strong arguments. Let's hear the counterpoint. We have Sarah on the line from Sydney. Sarah, welcome. >> Hi, Paul. Okay, look, David has it all wrong. Milk first is it's a total game changer for texture. >> Okay, cool. So, if you want to play around with it yourself, you can go to AI.studio. And on AI. studio you have that build section on the left side and there's a gallery and gallery is basically pre-built applets use cases for Gemini and if you scroll down to it uh you
16:15 have new build experience with and there's the talk show radio if you click on it you will find the applet you can uh remix it meaning you can like adjust it so if you don't like talk show radios and you would rather have I don't know some real podcast uh application you can like very easily together with Gemini adjust it to their own needs. And since applets are all open uh in terms of code, we can also take a quick look on like how our agent really is.
16:45 And our agent is just files. So we have an agents MD file inside our environment. That file will be loaded on environment startup which is then included into the system instruction. Very basic agent file. It has all of like the instructions for our agent to do. So it should do some research, write some scripts for the different people, generate music, generate a speech, mix everything together.
17:11 And to do all of those works, we have skills for it. So if we look into like the TDS generation, we have a skill on how to run it. We have a Python file. That Python file itself allows the Gemini agent to call other Gemini models. So it uses the interactions API to generate some speech which we can then use. And that that's all it takes for our agent.
17:33 And if you want to take that agent from the UI to like a back end, it's really you just need to copy that file and create the agent and then use it. And if you want to play around with the uh agent itself uh inside the playground, we have uh inside the model pickers a new agent section. So if you want to test deep research, go test deep research. But there's also the anti-gravity agent.
17:56 And on the anti-gravity agent, we have six different uh agent templates uh which you can like very easily test, select and and see how it works. We go into like the the standard one. And what I really like to do is like explore the environment like let's see what the agent finds out about where it is running. We send our prompt and now it starts the environment should take for the first request like 7 to 10 seconds and then we see the agent like doing its work.
18:24 So we have some reasoning, we have some bash commands. So it runs so it checks what uh version west version I'm running on what CPU memory I have and then we get back a response from our agent that is runs on Ubuntu um has 8 vCPU 16 GB of memory python installed pip installed and then you can like really continue let's see um do you have do we have node installed um so you can like very easily follow up um and it just continues based on like my prompt and then also you have the option on the right side to download the
19:04 environment. So if your agent has created some files you want to extract out of the environment, share it, can very easily download it or there's an API for it as well. And then um uh what else? Uh to quickly show you again how it works, I can make that a little bit bigger. So when you when we send a request um we like start the sandbox and it loads all of our files and then we have like this agent loop completely on our server.
19:33 So when you build agents normally a lot of times you need to manage the function calling on your client right you need to passse the function call run it some somewhere and then send back the function result and with the managed agents it's really just a single API call and everything runs inside that remote sandbox you can also connect remote MCP servers or run uh local MCP servers inside of that environment and um yeah really get started so if you are interested about it uh we have great documentation inside the
20:04 Gemini API under agents with a quick start which walks you through how to send your first request, build your first agent and get started and then of course uh install the coding skills to have claude codeex or whatever model you use to to help you get started. Uh thank you for for coming. If you have some questions I'm like here