Antigravity agent is a Google DeepMind agent harness for local use and managed remote sandboxes through the Gemini API and Google AI Studio. A single Interactions API call can provision a secure, Google-hosted Linux sandbox identified by an environment ID, where the agent uses Gemini 3.8 Flash to plan, execute Bash, Python, and Node.js code, observe results, and repeat until the task is complete. It can install packages, run tests, build applications, manage files, search Google, fetch web pages, and call custom functions or remote MCP servers. The harness supports reusable skills, configurable tools, multi-turn interactions, streaming, automatic context compaction, persistent sandbox files, synchronous hooks for intercepting and validating code-execution and filesystem operations, and sources loaded from GitHub, Google Cloud Storage, or inline files. Managed environments provide a credential-injecting proxy and named agents, allowing the same workflow to run locally and in the cloud.
Gemini API CLI is an experimental command-line interface for the Gemini Interactions API and managed-agent platform. It can run model interactions and managed agents, generate or analyze text and media, transcribe audio or video, manage files, models, agents, environments, triggers, and webhooks, and stream results in formats including JSON, YAML, and TOON. Commands expose machine-readable usage information and request schemas, while --dry-run validates inputs and previews redacted HTTP requests without contacting the API. Authentication can use API keys, OAuth access tokens, an operating-system keychain, or a configuration file. The repository describes the CLI as beta software with possible breaking changes, generated through Speakeasy; its software is licensed under Apache 2.0, and the repository states that it is not an official Google product.
The Gemini Interactions API is Google's unified API for calling Gemini models and specialized agents. It maintains server-side interaction state through interaction IDs and the `previous_interaction_id` parameter, while also supporting stateless requests. Each Interaction resource contains a chronological sequence of execution steps, including model thoughts, tool calls and results, and final output. The API supports text and multimodal generation, structured and strongly typed outputs, built-in and custom tools, tool orchestration, observable execution steps, background execution for long-running tasks, and chained image, video, and audio generation using shared context. Google documents it as generally available and recommends it for new Gemini API projects.
Google AI Studio is a web-based development platform from Google for exploring and evaluating Gemini models, developing prompts, and turning natural-language ideas into code and web applications. It supports integrations including Firestore, Cloud SQL, authentication, and database schema creation, and was used to create a mobile-responsive speaker-notes website.
Managed Agents is a Google Gemini API service for running the Antigravity agent in persistent remote Linux sandboxes. A call to the Interactions API provisions a sandbox, runs the agent loop, and returns an interaction ID and environment ID; subsequent calls can use the previous interaction ID for conversation context and the environment ID to retain files, installed packages, and other sandbox state. The service supports streaming, file downloads, loadable sources such as Google Cloud Storage and GitHub, reusable skills, secure credential injection, and configurable named agents.
URL context is a Gemini API capability that allows an agent to retrieve information from web pages and use that information in its processing. It is part of Google's developer platform for building applications with Gemini models.
Searchable transcript of An Interaction Is All You Need — Ivan Leo, Google DeepMind — AI Engineer (17:04). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:12 Hey guys, my name is Ivan. I'm on the developer experience team at Google Demine. And today we're covering the brand new interactions API and our new offering manage agents. And so to take a quote from a very famous paper, an interaction is all you need. Before we move on to what the interactions API is and what the new manage agents offering is let's take a trip down memory lane and let's kind of trace how models have changed and how we've worked with models.
00:34 So when we first started working with models a lot of times you would have a single interaction you would send a message tell me a joke and then you might get a response for example why don't scientists trust atoms because they make up everything. After a while as models got more and more capable and they understood more about the world this became a bit more complicated.
00:52 We wanted to take these models, put them in different applications, but to do so, we needed them to be reliable. And so to do so, we invented function calling, where models could create individual JSON objects with predictable structures. Think like how when you want to register for a website, your web page sends a JSON payload to the back end with a username and a password.
01:12 And that enables the backend server to say, okay, we have a new user and now we can actually create a new account. that made it possible for us to actually use models in different parts of this application whether it's extracting user information, parsing data and more. But right now we're seeing this incredible cups where models are getting more capable.
01:31 We're seeing models being able to reason go for very extended periods of time before they come back to people. And so you're moving from a place where we went from completions instruction and functions where we might ask, hey, what's the weather in New York City? and the model has a call weather tool that it can get that data, ingest it, and generate a response to now a case where models have multiple tools under their belt.
01:53 They're able to call tools, think and reflect on what the response meant, interact with the environment, and ultimately generate a response. And so this just brings us back to what is an agent. Fundamentally, if you look at what an agent is, it's really a model that powers a huge chunk of it. A lot of people have reduced this to just a language model running in a loop.
02:13 And for most cases, this is true. If you have a really capable model, you'll find that as we've got models that become more and more capable, a lot of the scaffolding has fallen away. A lot of models, for example, the latest Opus models or fable models are now just immediately just using a bash tool instead of individually specialized tools like a edit file or a read file tool like we used to do.
02:34 And so really what we have is a language model that's used to understand the context, the user's intent that's given a bunch of different functions, all different resources that it then interacts with its environment in. It has some short-term memory and some long-term memory to keep track of what it's encountered and then an environment to execute functions in.
02:53 If you're building a coding agent, but you don't give it an environment where it can actually execute code, for example, it's not going to be able to do anything. And so if you'd like to try any of our models, the easiest way to do so is actually with AI Studio. And so I actually like to point out that my speaker notes today were actually coded on AI Studio.
03:11 I vived it out 2 minutes before this talk. Uh you'll see the um you can see the conversation history over here. Um I literally said I need a mobile responsive website that I can pass in speaker notes because I found out I wasn't going to be able to use presenter view. And so what I did, I just gave it some brief instructions and before you know it, you have a simple application that you can use without any hassle deployed on a website that's accessible to all.
03:36 So we have a really generous uh free tier for most of my presentations. I've just been using the free tier. I haven't been using my Google unlimited tokens thankfully. And I think it's the best way for you to try our models. Whether it's our new speechtospech translation models where you take in some you speak in let's say English and it can translate to any language or you know you want to try our latest models like the Gemini 3.5 flash etc.
04:00 It makes it very easy for you to do so with simplified UI toggles to just be able to tune all of the different parameters. So let's talk about the interactions API. Now one of the things that we've seen especially as we've moved from models to a dynamic where now we have models and agents is that the kind of workloads that we want to support have changed significantly right through the Gemini API you have nano banana you have you know Gemini 3.5 flash but you also have very capable agents like the deep research agent
04:28 that's going to run for anywhere from three minutes and more you interact with it you have some sort of research plan it's going to go out and do its own thing and it comes like a very deep and detailed research reports. Now, it's pretty difficult then to think about consuming models and agents through a whole bunch of these different endpoints. Additionally, for a lot of the older kind of API endpoints, a lot of the data was in in very deeply nested objects.
04:54 And so you can imagine the kind of difficulties that we had ourselves when we were building in first-party integrations. And we wanted these use cases where you might be able to generate a nanobanana image, do some research, generate a vo image from there or spawn out all these different agentic kind of like agents that go out and do their own thing.
05:11 And so the interactions API is our sort of solution to this as we look to build an API that can support all of these different use cases and ones that we can build on ourselves. you look to bring Gemini across the entire surface of the Google ecosystem. So the best place to get started is just with serverside state. One of the things that you'll realize especially as you use the new Gemini series of models is that we essentially have this thing called thought signatures where if you interact with the model, if it calls
05:39 a function, if it gives a response, it's going to pass back a very opaque series of these numbers. For a lot of people, manually managing this was very difficult and sometimes when we talked to different startups, we found that they would lose their cache just with a single whites space that they mistakenly added in. So now with the responses API, you can see on the first turn, hi, my name is Phil.
05:59 And in the response, we give you a interaction ID. And as long as you give us back that same interaction ID under the parameter previous interaction ID that you'll see in the second turn, all that context is preserved. And this is really important with the new Gemini series of models because if you don't give us back those thought signatures, you're going to see a decrease in performance.
06:21 But what else does this unlock? Here you see a demo built by Vampsy from our team that did an incredible demo where using a single photo, you could generate with Nanovana all these different variations of yourself in different countries and different locations. Each of these individual photos are a single interaction, but you can think of it from an initial API call.
06:38 It spawns all these different interaction calls. Using those interactions, the same interaction ID, we then generate a video using the interaction ID and the new Omni model that we just launched. Something like this would not have been possible without the interactions API built by Philip over there. So, actually, it's a really incredible piece of work, I will say.
07:01 So, to give you an example, an idea of what it would take to build something similar, this is all it would take. First, you pass in Gemini 3.1 flashlight image, generate an image of me in Venice. You take the interaction ID, you chunk it back into the client.create, and this time you specify that you want the Omni flash model instead that we launched on Wednesday.
07:20 And that allows you to now use the same Omni model with the same information, the same kind of images, and the same context throughout all these calls. It's making it possible to build more complex, rich, and just these really complex agentic applications. The same thing occurs with audio generation. If you look at the bottom, you see that now we have output.type.
07:40 In the past, you would have to take this very nested object like completions audio. And it would be very tricky for you to nest it out. I can't even remember it off the top of my head. But now, every single bit of the outputs that come out are clearly demarcated by a strong type. If it's an audio, as you see over here, all you got to do is just parse it and handle it accordingly.
08:02 Right? But what happens if now I want to generate an image? It's the same method over there, client.interaction.create. All I got to do is just modify the response modality and the generation config. And now I have an output.type of image. And this makes it very easy for me to start building and experimenting with Gemini. The same thing comes with tool use.
08:21 You can mix and match different inbuilt tools that wasn't possible before. For example, over here we're trying to get an agent to look up the internet to find the latest security reports for the React application. We've given it access to the same index that Google has with type Google search. We've given it the ability to be essentially retrieve information from web pages using the URL context tool that we ship with.
08:41 And we've given it a custom tool here called file incident. The model can reason autonomously on its own, figure out what information it needs, and then actually generate a response for you all in a single API call. This makes it very powerful because now you have models that not only can work with information at their training cut of date but beyond because now they have access to the entirety of the internet.
09:04 And all this boils down to the new steps data model that we have where instead of a world where we have a single message sent to a model and a response, we're going to have very complex things like async tool calls models working together, right? And we want to make sure we have a strong kind of data structure to support this. And so we've introduced the steps data model which replaces the original legacy outputs array with a tight discriminator that makes it very clear at each step what the model output what the dot
09:29 signatures were what the functions called were and especially when it comes to the content is being generated. Earlier you saw content.type where it was audio and video and these makes it easy to build these multimodal pipelines in a way that only the Gemini model families can. The second thing that we've launched is anti-gravity as a remote agent as what we like to call manage agent.
09:48 So we've standardized around using the anti-gravity harness across all of our Google suite of products. In AI studio, which you saw just now, websites are built with the anti-gravity agent. In the anti-gravity application, you're also using the anti-gravity agent. Now, this is a very capable harness that we've co-rained Gemini on. And moving forward, I think you're going to see many gains if you use this harness out of the box.
10:11 If you wanted to build a coding agent in the past, you would first have to tune the harness. Then you had to find a sandbox provider, manage the infrastructure, and then kind of figure out way how to preserve the context, especially in between runs. But with the anti-gravity manage agent, you get a persistent sandbox out of the box. With a single API call, you get a sandbox that you can keep hitting and treat as a personal claw without any sort of modifications.
10:37 You can see an example over here where we want to analyze what's in a GitHub repository. We then pass it over in an API call over to the anti-gravity agent and it then boots up a remote sandbox and then starts investigating what the repository is about. You can see it calls a list file to get what's in the workspace. It then reads each individual file and then it generates a final report.
10:58 All of this is happening in a remote sandbox without any sort of intervention from our site. The agent's able to research, reason and handle all of this autonomously. We also handle multi-state persistent sandboxes like what we mentioned earlier where as long as you preserve the environment ID that we returned to you, you're able to route it right back to the original sandbox that you had.
11:17 So these are two primitives that we given you. The interaction ID where you're able to preserve and work with context and ensure you're not busting your cache and the environment ID that enables you to route it to the exact same context. And so this way you can use the same sandbox. You don't need to manage any sort of persistence of state and all you need to do is just add a few lines of code.
11:37 The other way we've tried to make it easy is for you to be able to load in sources. A lot of times you might have custom dependencies. You might have EMD files, helper files that you want to throw on the on this specific sandbox. So we support Google GCS buckets. We support GitHub repositories and inline files. And so what you see over here is a demo whereby it continues off the previous call that we did.
12:00 We're passing in the environment ID and the model has the exact same files right where we left it. any sort of packages you install, any sort of files you throw and create, the model has access to it through each and every turn as long as you passes back the interaction ID and the environment ID. This isn't just the sort of constraint to very simple task with the manage agent.
12:20 What you're seeing over here is a run where the model burnt over two million tokens trying to analyze a repository that we got for a hackathon where someone built from scratch an entire programming language to code reinforcement learning environments. If you've ever used the Scratch programming language, it's a drag and drop language where you can kind of get simple characters to move.
12:36 But over here with a single API call, we're able to analyze an entire repository that has a custom DSL, a web page, a whole bunch of like schemas, and an RL training loop running inside HUD. And so it really opens up this kind of opportunities for you to build these very custom and performant agents without you having to handle the underlying infrastructure or tune the harness.
12:57 The best part about using the anti-gravity agent is that fundamentally it's the same agent running in your anti-gravity IDE, which means that if you have certain workflows that are you've tuned locally, you have skills that you've made sure are great, you can once you're happy with those skills, you just package it into a single folder. We allow you to upload the same way using the sources into into agents folder and that gives your agent the exact same skills with the same agent harness and the same prompt running
13:23 locally and in the cloud. When you're ready or when you're comfortable with what your agent is doing locally, you can then push it up and everything works as intended. The next question we often get from companies are, "Hey, I need my agent to be secure. What am I going to do if my agent leaks a network credential?" And to do so, we've implemented a man-in-the-middle proxy.
13:43 So, what does this mean? A lot of times, you need to make API calls to different providers, say the Gemini API. You might need to download private GitHub repositories, etc. We have a proxy that stands in the middle and so any call your model makes we will look at the header and we can replace items at will. What you see over here is a case where we'll transform every single outbound call to the GitHub API with an inject in the token dynamically.
14:09 And so even if the model is somehow prompt injected and leaks your GitHub API token, it's never actually going to see the token. It's only going to execute code that will run against the GitHub API and your token is never exposed. Once you're comfortable with how the agent is performing, you've set it up well, we give you two different ways for you to basically configure and set it up as a named agent, much like you've been using the anti-gravity preview agent.
14:34 The first one is to freeze it with a certain set of configured files and network sources. That's what you're seeing on the left here. I have a data analyst agent that I'm creating from the anti-gravity preview base agent with some base environment configurations here, just a GCS skill. The other way that we allow you to do so is by chatting iteratively with an agent to set up an environment as you like.
14:54 For example, you might want certain dependencies, you might want certain packages, and you can freeze an environment as a specific agent and all your subsequent calls kind of go through that. Once you're happy with that, it makes it very easy for you to scale up your workloads. At this point, we support up to a thousand named agents, which means that you can have all these different configurations and you don't pay for any of the storage.
15:15 You don't pay for you don't pay for the sandbox. you only pay for the model. We've optimized our infrastructure to be very performant to be able to handle a workload like this. And we're pretty excited to see what people are going to build. Recently, we've also out open sourced a new Gemini API CLI. This is a simple CLI that you can use to get started working with the Gemini models to test out the models locally using simple CLI commands.
15:40 The other benefit of this is that you're going to be able to iterate on your code on your coding agent or your anti-gravity agent locally. And once you're ready, it packages it as a single folder and uploads it straight and creates an agent for you. Lastly, if you're thinking of switching over to the interactions API, which is where all of our new models will be on, we've also created a new interactions API skill that you can feed to your favorite coding agent and that will then basically allow you to do the migration
16:09 without any issues. I'm sure you guys have faced issues where you've asked Gemini to implement a specific kind of uh implementation and it keeps using Gemini 2.5 flash or Gemini 2.0. And so we what we do is we make sure we keep up to date the list of all the different models. We keep up to date all the different methods. We provide a good bunch of skills and like Philip mentioned yesterday, we regularly evaluate these and make sure make sure that they're actually doing what we want them to do, which is allowing you to
16:33 use Gemini and build on Gemini well. So that's the end of my talk. Uh, thank you so much for coming by and staying around for the round of the conference. And yeah, I hope you guys get started. I'll be on the side if you guys have any questions about the new API. [music]