Google AI Edge Gallery is an app for downloading and running Gemma models locally on Android and iOS devices. It supports image description, multilingual audio transcription, on-device function calling, and built-in skills such as games, haiku generation, weather queries, and calendar events. Inference runs on the device, using supported hardware accelerators rather than requiring an API call or sending data to a remote service. Google describes AI Edge as its platform for deploying machine-learning solutions across mobile, web, and embedded applications.
Managed Agents is a Google Gemini API service for running the Antigravity agent in persistent remote Linux sandboxes. A call to the Interactions API provisions a sandbox, runs the agent loop, and returns an interaction ID and environment ID; subsequent calls can use the previous interaction ID for conversation context and the environment ID to retain files, installed packages, and other sandbox state. The service supports streaming, file downloads, loadable sources such as Google Cloud Storage and GitHub, reusable skills, secure credential injection, and configurable named agents.
Transformers.js is a JavaScript tool for running AI models locally in a web browser. The video demonstrates it running a quantized Gemma model without an API call or sending data elsewhere, with inference performed on the user's device; the example's two-billion-parameter model uses less than a gigabyte.
Searchable transcript of Research to Reality with Google DeepMind — Paige Bailey, Google DeepMind — AI Engineer (10:53). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:12 trickle in from other locations and then we'll get started pretty shortly. Um, ideally there will be some time for lots and lots of live demos. Um, so you'll ideally learn some of the things about our new models, how to use them as part of your projects. Um, and then there also should be time for questions, but it also seems like we'll have a small enough audience where if you ask questions throughout the duration of the presentation, we can also do that.
00:35 Um, so end goal is to make this as useful as possible for all of y'all. Um, my name is Paige. I work at DeepMind and excited to show you all what we've been working on too. And as mentioned, we'll get started in just another couple of minutes. It looks like people are still getting scanned at the door. All right, let's get started. Cool. So, greetings everyone.
02:41 As mentioned, my name is Paige. If you have questions throughout the duration of the presentation, please feel free to raise your hand and shout them out. Um, ideally, we'll be able to take some uh take some as I kind of carine along doing live demos. Um, I also want to give a kind of the requisite caveat at the very beginning that um, DeepMind's entire mission is to build AI responsibly and for the benefit of humanity.
03:08 So trying to cast as much light as we can um in this uh in this world that increasingly has uh increasingly has shade. Um this translates to things like alpha folds um much of the work that we do with our open models including medge gemma um and a lot of our robotics and science use cases as well as things like AI for math AI for frontier research. Um, so in general, everything that I show you today, um, is being used for those projects that are deeply deeply embedded in AI for science.
03:40 Um, and if any of y'all work in the AI for science space, um, please feel free to send questions and ask afterwards about how you can use AI to to kind of transform and accelerate that uh that good work. Um, I don't think it's a secret that Google has been a little bit busy over the course of the last few months. um feels like we've been releasing new models, new features, new products every single week.
04:05 Um most notably, uh just recently we released a computer use API which we'll be talking about a little bit. um something called managed agents which gives you the ability to take um kind of a higher order task uh describe it in a natural language and have a fleet of agents execute on it in a Linux workstation like a sandbox environment where you can add skills you can add um kind of dependencies that get pulled in along the way and a whole bunch more.
04:31 Um and then also our speechtoech translation API which we'll take a look at in a second. And one of the things that I really really love about Google is that not only are we shipping Frontier models. So you might have heard about Gemini 3.5 Flash which got released at IO. Um but we've also been focused pretty significantly on our open model releases.
04:53 Um so how many folks in the room have heard of Gemma 4? Um quite a few hands. That's excellent. Uh Gemma 4. Oh yes, absolutely. Go for it. Um, Gemma 4 is our open model family. It's the latest iteration. Um, we have many different sizes available. So, there's a two billion parameter, a 4 billion. We just released a 12 billion parameter. Um, and then we also have a 26 billion and a 31 billion um, mixture of experts and dense model um, respectively.
05:24 Um, they're useful for a lot of things and they're also Apache 2 licensed which means you can download them, use them as part of your company, fine-tune them, uh, and kind of, uh, expand on them however you feel like would be most useful. Um, and this Gemma universe is actually pretty cool to see, um, in terms of people building. I I hate slides, so we're going to see how few slides I can get through today.
05:51 Um but uh this is an example of something that someone from the community has built using Gemma 4 um and Fable 5 before it was taken off the market to rewrite some of the kernels. Um you can load the model directly within the browser. So this is loading Gemma directly in the browser using it via web assembly. It's sandboxed. Um and then it can do things that that feel pretty magical, right?
06:16 So this model is uh this model is currently running completely locally. Um and if I say something to the effect of create a table with emoji comparing and contrasting all of the uh Harry Potter books um based on which are the funniest and most exciting. um make sure to give me recommendations on which to read and also um um maybe incorporate Harry Potter Harry Potter in the methods of rationality um or and then you get kind of a a response that feels almost instantaneous.
07:11 Um, you have uh kind of the Deathly Hollows, Half Blood Prince, Sword or the Phoenix Goblet of Fire, etc. Um, I'm not sure if I agree with the ranking, but that's uh but that's a pretty interesting pretty interesting comparison. Um, and then also uh since it is local um to the machine, none of your data is getting sent elsewhere. This is not using an API.
07:33 It's just something that's running locally in the browser with transformers.js um and the Gemma 4 model. Gemma 4 is also pretty cool in the sense that you can use um Google AI edge gallery in order to analyze um some of the model capabilities that we have on device. You can download it and use it with Android and with iOS. Um and it includes everything from kind of taking images and describing them to automatically transcribing audio in multiple languages.
08:03 I believe Gemma supports over 140 different languages. Um, and then also testing it out with function calling directly on device. If you haven't had a chance to take a look at the Google AI edge gallery, um, we have a collection of skills as well. Um, so things like building games, doing haikus, um, asking queries about weather, being able to schedule events on your calendar for you, um, that are all available to use just with this AI edge gallery app completely for free and just with locally installed models.
08:34 Um, so if you haven't downloaded it, definitely try it out. There's also a way to use um the the kind of accelerator on your local device in order to power the model and do all of the inference work as opposed to just the fall back to the CPU. Um, so if you do have something like a Pixel 10 or a higherend um a higherend mobile device, you're already able to use Gemma kind of locally um for for all of that work, which is quite cool.
09:07 Gemma 4 just to to give a recap or to place how the model performance versus size shapes up. Um this is the two of the two kind of largest versions of Gemma that I had mentioned before the 31B and the 26B. Um they're quite small in comparison to some other models that are on the market but they're performing way above what you might expect. um so more than models that are in order of magnitude or larger than their size.
09:37 This is great because it kind of translates into a cost savings perspective. You don't have to worry about distributed inference for local models. Um you don't have to worry about as large of a GPU footprint. Um and it's really intended to run super super well on a single commodity GPU as opposed to needing something a little bit more fully featured.
09:56 You can also run Gemma, some of the variants on even things like Jets and Nanos. Um, and for 12b and below, you can run it locally on your laptop. Um, 2B can even fit handily on mobile devices. We've also released quantized versions. So, the the model that we saw at the very beginning that was running so so blazingly fast um was using some of the QAT checkpoints that we have for Gemma.
10:20 Um and the QAT checkpoints um are quite small when you uh when you take a look even less than a single gigabyte in size for the the uh kind of two billion parameter version. [music]