← All transcripts

Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work Transcript, AI Summary & Key Points

Y Combinator · 2 hours ago · Science & Technology · 49:24 · EN-US

📄 Transcript

Searchable transcript of Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work — Y Combinator (49:24). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Y Combinator. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:07 Good afternoon everyone. It's great to be here. Uh we talk a lot about AI that lives on your screen uh lives in the digital world. And today I'd like to talk to you about a different kind of AI that we've been building at Whimo. AI that lives in the real physical world. How many of you by the way have been in a Whimo? Just raise your arms. Wow. Okay, that is impressive.

00:39 Especially I understand many of you are out of town. Uh the folks who are visiting and have not had a chance to check out Whimo, I hope while you're here in the Bay Area, give it a try. Uh so this being a startup school, I structured this presentation as a sequence of lessons, seven lessons that we've learned over the years at Whimo around what it takes to build and safely ship today's most mature application of AI in the physical world, the Whimo driver.

01:14 Uh let me start with a short video. Uh this is a clip from a ride that I recently took in a Whimo with my kids. Uh so as you see here, you know, we're moving uh forward. We're proceeding through an intersection and a couple of human drivers just decide to cut in right in front of us and the Whimo driver reacted safely, reacted smoothly. In fact, so much so that the kids, my kids were preoccupied in the back seat.

01:44 They didn't even notice that anything happened. And to me, this was a pretty powerful moment. I've been working on this technology and this product for close to two decades, and you know, it just did something fairly important. It acted safely. It kept my kids safe. It kept everybody safe and nobody noticed. And that I think will be a bit of a theme in general when it comes to physical AI that the best AI moments will look like nothing happened.

02:13 It's just the task got done safely and smoothly. And these sort of moments where the Whimo driver kept everyone safe are happening daily across our fleet. Today the Whimo driver is serving around 500 trips per week and driving over 4 million fully autonomous miles every week in 15 cities across the United States. Just for a comparison, that's over 300 years every week of an average American driver per year.

02:51 And the Whimo driver is accomplishing that with a superhuman safety record. So what what does it take to build and deploy an AI agent in the physical world at scale? Now in Silicon Valley there's a common mantra to move fast and break things. However, when you're dealing with atoms instead of bits, breaking things is not really okay. So the thing you have to do is to move fast and ship safely.

03:29 And that's a much more difficult thing to do. You have to build systems that are robust from day one. You have to build AI models and you have to build training recipes where safety is the foundation and not an afterthought, not an add-on. And by the way, the problem itself of physical AI is different from digital AI. There are four main gaps that you have to contend with if you're building AI for the physical world versus the digital world.

03:59 First, there is the cost of error gaps. have a language model or a chatbot or a co-pilot and it makes a mistake, you know, usually it costs you a retry. In the physical world, the cost of a mistake can be measured in human lives, not tokens. There's simply not an undo and a retry button. Secondly, you have the latency gap. And typically when you're running a VLM uh or you know a digital assistant, it can take many seconds, sometimes minutes to come back with an answer to you.

04:38 A car traveling at freeway speeds moves about 100 feet in 1 second. So their milliseconds really matter. And you have to run all of your inference, make all of your decisions on board a compute that fits in a trunk of your car. Next, there's the data gap. Uh, digital AI had the internet of this wonderful immense cache of pre-labeled human knowledge and human thought that we've ever assembled.

05:10 There's no digitized version of the internet for the physical world. And lastly, there's the validation gap. In digital AI, often you can ship something that's good enough and that you let your users you uh use your product, they find the edge cases and that allows you to deploy on day one practically at unlimited scale and then you can just iterate and hill climb on quality from there.

05:40 In physical AI, the situation is different. Given the high cost of errors, you need to have a very high level of safety and a very high level of confidence on day one before you deploy your first robot, before you drive your your first autonomous mile. Now, at the same time, when you're dealing with physical AI, uh the actual experience of having your agent in the real world is invaluable and it's irreplaceable.

06:12 uh these systems are not just something that you can build in the lab you know get it perfect and then deploy at full scale overnight. So given those two factors you really need to super clearly and super crisply define the operating conditions and the deployment parameters of your agent and then build a rigorous framework to guide your deployment so that you can scale in a responsible manner.

06:38 And this is absolutely critical. Uh this is how you earn trust from your customers, from the communities, from the regulators and yourself. Uh so at Whimo, we see these gaps of course in the context of autonomous vehicles. Uh but these gaps uh will show up in practically any sort of non-trivial physical agent that we will deploy uh in some shape or form.

07:03 and driving is just simply the first domain where AI has crossed these four gaps at scale with the public interacting with our product. So let's dive into those lessons that we've learned over the years at Whimo from uh working on this problem and uh talk about how we address those gaps. Uh I have seven lessons in this talk. Um I they're all technical.

07:26 There's a lot more that goes into building a company and building a product. Uh but today I'll just focus on the technical aspects of building AI for the physical world. Uh and each one of those lessons I think by itself will not be exactly earthshattering. Uh you know a lot of it will overlap with likely things you've heard elsewhere. But I hope that the grounding of these lessons in our experience and some of the nuance that I can add about how they showed up uh in our experience of deploying a physical agent in the

07:58 and scaling it safely will be interesting and useful for many of you who are in the space as you build your product as you build your startup. Uh so let's dive in. And the first lesson has to do with this uh massive frustrating sometimes soul crushing difference between a demo and a real product. And a working demo is 1% at best of the work that you have to do.

08:25 The many nines of performance, the many nines of reliability that follow, that's where the real work happens. And if you're a founder in the room, uh, chances are you are focused on getting that first prototype, that first demo, uh, off the ground. And when you hit that first version of a system that works, that first 90%, when the demo actually works, it feels incredible.

08:45 You feel like you solved it, the sky's is the limit, you're extrapolating forward. And in our world, we hit that first milestone, that first 90% back around 2010. So when this project started uh in 2009 before we started building the system we set a couple of pretty ambitious goals for ourselves. One was to drive a 100 autonomous 100,000 miles in autonomous mode.

09:14 The second goal was to drive 10 routes. Each one was 100 miles long uh chosen to cover a v variety of conditions across the Bay Area. And we had to do each one from beginning to end without a human intervention. We had at the time a team of about a dozen engineers and we accomplished both of these goals in about a year and a half. And keep in mind this was well before any of the AI breakthroughs before connes before transformers before BLMs before any of the stuff that we talk about today.

09:45 Uh and yet you know we got it done and kind of by demo standards we driving autonomous driving was solved in 2010 right we handled everything we handled we could drive during the day during the night we handled traffic pedestrians cyclist traffic lights construction zones on freeways on surface streets so we were quote unquote capability complete and you know at the time we felt like we're on top of the world but then we quickly ran as we started building towards a product we quickly ran into a brutal reality that

10:17 there's a massive difference between doing something once or driving 10 routes once and building a scalable service with nobody behind the wheel. It took us about 10 more years to begin providing a service and then five more years to scale to half a million trips per week. So the demo took 18 months, the product took about 15 years, but now we're scaling exponentially.

10:45 To date, we've served well over 20 million fully autonomous trips, and we've driven well over 200 million fully autonomous miles. And we have rider only vehicles operating in 15 cities across the United States. And we're scaling exponentially. It took us uh 15 years to get to that first 100 million miles and about 7 months to drive the next 100 million.

11:11 It took us about 8 years to go from the time when we started our initial rider only operation to the time when we had uh when we were serving riders uh in four cities. Earlier this year we launched four cities in just one day. So why does bridging that gap from demo to product take so long? Well, because there's this harsh engineering reality that you can't really cheat.

11:43 That reliability and performance lives on this exponential ladder of nines. So getting to that first 90% or 99% that's the easy part. But then every next nine that you want to add, that takes about 10 times more effort. So you need to know upfront exactly how many nines your product actually needs. So demo might need, you know, one nine. An assist product or, you know, co-pilot might need a few, but a fully autonomous AI agent that we're going to be putting out in the physical world that engages with the public, you

12:17 know, with kids running around, that needs a whole stack of them. And at scale, the long tail is the problem space. It's your entire problem statement. When you drive millions of miles per week, a rare event that might happen once in a million miles, that just becomes your daily reality. And getting those next nines means doing something different every time.

12:42 So you don't get to say six nines of performance or reliability by doing the same thing that you did for you know to achieve the first two but longer. You have to do fundamentally different things. You requires a fundamentally different approach. For example, we can take reliability. You can, you know, get to the first couple of nines by just doing proper engineering and uh doing, you know, uh some bug fixes, but to get to the next few, you need to have you need to invest in fundamentally different approaches.

13:10 You need to build fully redundant systems, uh have tiered fallback architectures and so forth and so on. And the same thing holds for the performance of AI models. So what that actually means is that in this space it's incredibly easy to get started but it can be excruciatingly difficult to get to the real product and that effect is only amplified with every wave of technological breakthroughs and that naturally leads to hype cycles.

13:41 So every AI breakthrough from you know deep learning to condors VLAMs you name it it makes it that much easier to get started your demos your prototypes they get a 100 times easier but the tail that's where the hard problems are that moves much less it moves but the effect is muted and that's why every hype cycle produces a wave of absolutely spectacular demos and very few real products and the recurring mistake of every cycle is spending on the demo what you should be saving for the eyes.

14:16 Now, you know, this being a startup school, the last thing I want to do is throw too much cold water on the magic and the excitement of those early days. This this time is absolutely magical. It's amazing. Cherish it, leverage it. But the key is to remain honest about the product that you're building, the number of nines and performance and reliability that that product demands, and not cutting corners to get there.

14:37 uh otherwise you might be in for a pretty rude uh awakening later. So count your nines before you count your demo views. Uh and this brings us to the second lesson. Once you know how many nines your product actually needs, it fundamentally dictates the architecture and the core technical approach that you need to pursue. Now every technology has a performance versus effort curve, right?

15:05 They all tend to start fairly steep and go up and then they flatten out. And you know, as I just mentioned, every other nine gets an order of magnitude more difficult. So common failure mode is picking the tech that gives you the fastest early ramp, riding that steep curve, feeling like you're winning, projecting that, you know, steep slope into the future and feeling like the sky is the limit and then hitting the plateau and discovering that the technology path that you picked actually flattens out way before the

15:40 performance that is required by your product. Now, you might still choose to be at least for a while on that steep curve uh for a variety of practical reasons. You know, maybe you want to prototype something or demo something or build something in service of learning. But be honest with yourself where you're building for the purpose of a demo, for the purpose of learning or towards an actual product.

16:04 Uh so let's take uh an example from our domain autonomous vehicle and sensing. Uh there's been a long-standing debate about what kind of sensors do you actually need for autonomous driving. Naturally, more sensors means higher performance, but also means higher complexity. So humans of course can drive with just eyes. So there's that proof of existence.

16:26 Now, and if the goal were to just approximately match human performance or to build an assist product, that's a very reasonable way to go. However, if you're targeting full autonomy and you're targeting superhuman strongly superhuman performance, you find that weak sensing just leads to a safety curve that flattens out way too early. So, at Whimo, we've taken an approach where we use multiple sensing modalities.

16:52 We use cameras, lighters, and radars, and they all complement each other. Cameras give you high resolution and color, uh, but they're passive, and they degrade in darkness and glare. Lighter gives you a direct measurement of the 3D structure of the world uh around you and radar is very good at punching through environmental conditions and weather like fog or rain or snow and it directly can measure velocity using doubler.

17:19 Uh lighter and radar are active sensors. Uh so that means they see just as well in pitch darkness or for example when driving into a blinding sunset. And these different sensing modalities, of course, they're not backups to each other. Uh, in our stack, each modality has an encoder and the information from all of those sensors get fused into a single view of the world around us that is much more precise and um, generally vastly superior to what you get with any one sensor.

17:52 So, let me show you a few examples. Uh, here's a scene where a Whimo is driving in a dust storm in Phoenix. So what you see here is what the scene looks like to our fairly advanced high resolution and high dynamic range camera. It's very close to what a c, you know, a human would see in the same conditions, which is not much. And here on the right is what the lighter sees for the exact same frame.

18:19 And you can much more clearly see that there's a pedestrian standing on the side of the road. So if they were to step onto the road, that early detection can make a really big difference in how the situation plays out and the safety of everyone involved. Here's another example. At night driving along and there are a couple of pedestrians who are about to jump onto the road over a concrete construction barrier.

18:44 Again, at the bottom you see the camera, really can't see much. And the lighter view at the top. Again, lighter versus camera. Here's another example. A couple of dogs chasing a bull and a couple of kids chasing the dogs. And big difference. Here's what it looks like to the camera. Here's the lighter and the early detection of the uh the kids is off to the side and there are no headlights.

19:13 There are no lamps there. It's complete darkness. So, it makes a big difference. Uh or think about what happens when something physically obstructs the view of your sensors. uh you know if you don't have redundancy in sensing you can have you know a single leaf land on your sensors and bring your robot to a full stop. Uh now so you need redundancy. Redundancy of course does not necessarily mean multiple sensing modalities but if you need redundancy anyway you might as well uh benefit from the complimentary physics of

19:49 the different sensing modalities in the nominal case. So, here's a a video of uh one of our cars that picked up a leaf or actually I think a full uh branch of a tree uh that our wipers were unable to shake and the car detected that and because we have sensing redundancy, it safely was able to get back to the depot for proper cleaning. Uh so specifically when it comes to hardware uh do not anchor to today's components prices.

20:23 Uh we are on the sixth generation of the Whimo driver, the Whimo hardware suite today and with every generation the hardware not only delivered amazing capability but we're able to drastically simplify and radically reduce the cost of the hardware as well. So betting your company, betting your approach on today's hardware prices is just betting your company on a number that has a fairly short shelf life and is going to expire.

20:51 Uh so hardware will change. Uh many components will get commoditized and drop in price. So design for that future and be ready to upgrade. And that brings us to the next lesson. Lesson number three. Uh technology moves incredibly fast especially nowadays. So you need to be ready to ride those tech waves and do that repeatedly. And when you do have to not only think about the wins in performance and the wins in per capability, you have to be very mindful about uh unification and simplification.

21:32 Uh over the years we've seen a number of major breakthroughs in technology. a lot of them around AI and with every wave of innovation, we pretty much rebuild the Whimo driver around that major wave of AI breakthroughs and we often push the state-of-the-art in those areas forward ourselves. Uh we leveraged Connets around 2013 for computer vision and perception.

21:53 Then when transformers came about around 2017, we bet big on them both for perception and for the task of behavior prediction and decision- making and planning. Turns out the task of driving is not that dissimilar from the task of modeling language. Uh because of the social aspects of driving, you're kind of having a conversation with other dynamic actors in the world, but you're doing that in the space kind of body language of your agent, your car, as opposed to just the language of words.

22:23 Uh and you operate in sequences and local continuity matters. Uh but so does global context. uh and today we're leveraging uh the latest in VLMs and Frontier world models. Now using the latest tech for capability and performance wins, I don't want to say it's easy, but it can be you know reasonably straightforward. Doing applied research in isolation or starting a tiger team to you know prototype uh some new technology uh is not the most difficult part.

22:52 There's many companies, many teams that are excellent in this. The much harder muscle to build is to carry that bleeding edge research into production and deploy it in a safety critical environment without regressions and do it without breaking stride on the scaling of your product. And adding capability again is not the hardest part, but adding capability while at the same time reducing fragmentation and reducing complexity that is really important.

23:22 And finally, the hard muscle to build as a company is to be able to do that repeatedly through multiple waves of technical innovation and technical breakthroughs. So on this front, I have two bits of advice. Uh the first one, when the technology, a new technology shows up, you know, it can be very exciting, very tempting to kick off uh a new effort, a tiger team to pursue it.

23:44 And that's great. You should absolutely do that. However, when you do, it's very important that you consider what you would do after under a success scenario. Let's say that effort succeeds, you should be very clear on what the path of that new innovation is for your company, for your entire product, for your entire system. Uh, often times I've seen a failure mode where you know a project, a very difficult technical project succeeds and then there's a dead end.

24:09 So that can be very wasteful. That can be completely you know deflating. Uh the second bit of advice I have here is uh when pursuing new tech again don't just ask what does this new tech give me in terms of capability and performance also ask has it simplified my stack and has it led to fragmentation or unification. So set your launch bar to demand both breakthrough performance and at the same time radical simplification and unification.

24:41 And this exact philosophy and this muscle that we've built uh at Whimo over the years is what produced our latest core technology. Uh and this the heart of it is the Whimo foundation model. Now the Whimo foundation model is a multimodal world action language model. That's kind of a mouthful. So let me unpack the ingredients. It's a multimodal model because it is able to process these multiodal sensor inputs, cameras, lighters, and radar.

25:07 Uh it's a world model because it inherently understands how the world works, the physics, the dynamics as well as the social and semantic aspect of it. It's an action model because we are not just passively observing how the world evolves. Uh we're an active participant. So the model needs to understand the effects of our the actions of our agent on the world and be able to tell the good ones from bad ones.

25:34 And finally, it's aligned with language and that allows us to unlock general world knowledge from visual language models and that's incredibly useful in the long tail of rare semantic uh situations. So more specifically uh this is what the architecture looks like. It's kind of your typical encoder decoder architecture. The encoder part takes in the multimodal sensing and compresses it or encodes it into a efficient into an efficient representation uh that retains all of the relevant data, all of the relevant

26:10 information for the generative part or the decoder. It's an endto-end model um which has a couple of nice properties. It allows us to effectively back propagate the gradient from the task that we actually care about all the way to the early layers of the model and it allows the encoder to reach you kind of learn the right rich representations uh for what the generative part needs to solve the task.

26:35 Uh it uses a system one system two think fast think slow architecture and it leverages the general world knowledge of VLMs for efficient learning of semantic tasks. So let's dive uh deeper. Uh first the think fast path. Uh that part fuses the raw data from our cameras, our lighters, our radars and that allows for split-second safety critical decisions.

26:59 So you can think of it as kind of your driving instincts. Uh this is what allows the car to break instantly if let's say a pedestrian runs into the road or a cyclist that's nearby swers into your path. uh this is like the the if you will the lizard brain of your agent that deals with a lot of geometric tasks and can react in milliseconds. Uh second is the slow path.

27:22 Uh that's the part that's responsible for the more complex uh semantic and scene level understanding type tasks. And these sort of task these things don't typically change in milliseconds. So there you can afford a bit more latency and you can trade that off for higher capability and higher levels of uh of reasoning. So for example, if the whim driver encounters a situation when there's, you know, a vehicle, let's say it's on fire on the side of the road, the fast path might just see it as a generic uh generic obstacle

28:01 and, you know, reason that the path ahead of us is clear. And this is where the slow path comes in. And that path can use deep semantic reasoning to understand the semantics of that object, the car being on fire and the broader scene context. And that allows our driver to decide to take a very, you know, different action or a different route entirely even if geometrically the path ahead of us is very is clear.

28:27 Uh and finally there's the generate component. That's the decoder. That's the component that understands and can produce behavior. Uh it understands how other actors behave uh and it allows us to make predictions and plan our own driving decisions. And our Whimo foundation model powers the Whimo driver that runs on different generations of hardware and runs on different vehicle platforms.

28:51 You have our fifth generation and sixth generation, the JLR IPA, the Ohigh and the Hyundai Ionic. and in the future will power different products and different commercial applications like trucking and personally owned vehicles. So by uh leveraging the strategy of focusing on the high capacity foundation of uh ofboard model, we're able to move a lot of complexity upstream to that large shared foundation and that allows us to make that specialization layer uh that's running on the car uh uh pretty lightweight and that

29:28 in turn allows us to speed up the development process. So the most important muscle in this lesson uh is for your company to not just leverage the tech of the day but have the ability and build that muscle to repeatedly ride those tech waves and pull in the results of that innovation into production without regression without breaking stride in deployment and scaling and without drowning in complexity.

29:57 So let's move to the next lesson. Uh there is a well-known lesson in the AI community that general methods that leverage massive compute and massive data will always beat methods that rely on handcrafted engineered human knowledge. That's the so-called uh bitter lesson that Richard Sutton uh published and formulated in 2019. And we have lived this and we have seen this in every wave of technical breakthroughs.

30:31 Each time the bitter lesson holds methods that scale best with compute with data, they always win out. And by the way, this is uh you know one of the reasons why we bet on the approach of building the foundation model. Um uh there is a well-known property that if you bet on high-capacity model and you use your data and your compute on that you just get better scaling laws and then you distill into smaller more efficient models that are running on your agent in real time.

31:02 You just get better scaling laws as opposed to just focusing on the smaller models directly. Uh so one nuance area where uh this lesson shows up is the use of structure in your models and depending on how you use your structure you can end up on either side of the bitter lesson. Essentially structure that fights scale will always lose and structure that channel scale always wins.

31:31 In particular this comes up around the discussion of end toend models. As I mentioned an end toend model is, you know, uh has some very nice properties. You back properly gradient from the final tasks all the way through the model and it allows the API between the encoder and the decoder to learn to use rich learn representations and you know those are the easiest models to build and train.

31:55 um you know you can start uh the architectures are known you can start with doing some imitation learning and a uh kind of a black box end toend model will give you very rapid progress and you will ride that very you know initial steep part of the curve uh and for some products that's enough but if you need to reach superhuman levels of performance in a fully autonomous agent in a safety critical environment uh just doing kind of that basic vanilla end to end is not enough and this is where structure comes in and the

32:23 key question here is does the structure boost scale or does it fight it? Does it limit and constrain your solution space or does it help you scale without loss of generality? So, let me give you uh an example. Let me illustrate this point with kind of a a simple thought exercise and a a toy problem. Imagine you're building a robot that will play the game of goal and it wants you know you want it to play the game in the physical world right so you have a camera that's observing the board and you have an actuator that

32:59 will actually move the pieces around now one way you can build such a robot is to you know have an end to-end system that goes directly from pixels to actuation and maybe you train it by giving you some videos of you know how humans play the game uh and that could be a very interesting research search exercise. However, if your goal was to build the world's best playing Go robot, that's probably not the most efficient way to go.

33:22 And the reason for that is that there is a very simple intermediate representation that captures completely the state of the game, the state of the the task that you're trying to solve. as you know 19 by19 board and that gives you a fully observable and complete state of the world that you care about at least for the you know game plan laying part. So leveraging that structure it doesn't limit your model it doesn't constrain your solution space but it gives you a very helpful way to scale.

33:53 Now that of course was a toy example. Uh anything that's not trivial that you're trying to deploy in the physical world will not have that property. And the fact that such a simple clean engineered representation doesn't exist in the physical world is the whole reason why we need endto-end uh systems and learned representations and learned embeddings.

34:12 But in the physical world that does exist structure. You have laws of physics. You have rules of the road. You have objects that behave in reasonably predictable ways. uh and you can use that structure in addition to the learned representations to boost your performance, simplify validation and at the end of the day you just get better scaling loss.

34:33 Uh and this is the approach that we are pursuing at Whimo which we call structure augmented end to end. So we go beyond the basic vanilla end to end by augmenting the learn embeddings with materialized structure representations and that gives us a few very important advantages. So first is validation at inference time. Now because the model isn't just a black box where sensors go in and you know uh actuation commands go out, we can create a very powerful correctness and safety validation layer that uh you can run in

35:05 real time in uh uh when the agent is deployed on our vehicles and this is really important for any agent that's operating in the physical world. Uh secondly, we get great wins in efficiency when it comes to largecale training and evaluation of the generative part of the model, the decoder. Now, if all you have is a blackbox end to end system, you are forced to do all of your evaluation and all of your training in the end toend setup all the way from sensors to decisions to actuation.

35:38 Having that intermediate structured representation allows you to kind of mix and match. You can do some training at larger scale uh and some evaluation in the space of those compact structure representations and some in the full space of end to end from sensors to decisions. Uh and finally we get strong verifiable feedback signals for both evaluation and for training uh training recipes to support things like reinforcement learning.

36:06 that additional materialized structure just gives you uh much more powerful tools for evaluation for metrics as well as crafting your loss function or you know reinforcement learning recipes. So the lesson here is to bet on a system that's maximally learned and minimally constrained and leverage structure intentionally to boost performance and scaling laws both in training and in evaluation.

36:32 Now that raises the question of how do you actually train and evaluate your physical AI agent? And that brings us to the next lesson. To build and safely deploy an agent in the physical world, it is absolutely critical to have a good large scale realistic highfidelity simulator. Now there's two ways you can do training and evaluation. You can do open loop and you can do closed loop.

36:58 uh in open loop kind of passively observing uh input to output pairs uh and you can use that you know for evaluation for training imitation learning works like that evaluation takes usually this shape of you know if you are find yourself in this situation what would you do and then you you score that uh and that's in contrast with closed loop where in closed loop you take an action uh you see the effect that that action has on the world then uh you update your uh through your sensors the view of the world, you take

37:30 another action and so forth and so on and you evaluate and you train on those sequence of actions and sequence of world evolutions. Now the ability to take an action and evaluate that counterfactual is absolutely vital for building and deploying safety critical agents in the physical world. Uh so a real simulator is how you do that. And a real simulator isn't just some lightweight tooling that sits next to your AI.

37:59 Uh it is an uh a big AI model in of itself. And the problem of building a good realistic simulator is just as hard as building the agent itself. Uh so the AI behind the simulator really needs to understand how the world works, the physics, the semantics, you know, the traffic, the weather, and so on and so forth. And the quality of that simulator has to be high enough so that it doesn't only look good, but it's sufficient to train and evaluate with high confidence an agent that you're going to be putting in the world

38:33 in a safety critical environment. So in other words, you have to build a highly accurate generative world model. And at Whimo for years, we've been building what we called uh behavioral world models. And you were doing that way before the term world models even became popular. Uh and now in the era of endto-end models, you also need on top of behavioral realism, you need sensing realism as well.

39:00 And in fact, uh building an end toend model has been fairly easy for quite a while now. But evaluating it in closed loop that was the hard part of the problem. So we've moved on to building sensing world models. And because we're using that structure augmented representation in our models, uh we can also leverage that structure in our simulation. Our behavior world model operates in the space of structured intermediate representations and the tightly coupled sensor world model then producing produces realistic uh

39:36 sensor simulations. Our world model leverages uh the great work of Google deep minds on Gen3 and that gives us the ability to produce controllable and highly realistic scenarios both in the behavioral as well as uh sensing aspects. And that in turn allows us to not just evaluate our agent uh and train uh new versions of our agent in situations that we've previously encountered, but it allows us to train and evaluate in purely synthetic rare scenarios that we've never seen in the real world.

40:16 So what you're seeing here is not just a generated video. It's a generative a full generous simulation of the Whimo driver uh in uh operating closed loop. Uh so here we're simulating what would happen if it came across a car that was stopped in a lane on the freeway. Uh and you can go further than that. Here's a plane that's landing on a freeway in front of us.

40:39 where you can simulate an elephant on the loose walking through the intersection, snow on the Golden Gate Bridge or a dinosaur walking around. So the lesson here is that closed loop simulation is absolutely required for evaluation and is extremely val valuable for training of your physical AI agents. So you need highly realistic large scale simulation to tri train and evaluate.

41:08 And this brings us to lesson number uh six. When you're dealing with a problem of that complexity, you can't just build a model and call it a day. You have to build an entire ecosystem. And then you also need a flywheel that powers it. Because to make this work at scale, you can't just build the agent. You and one AI, you need to build three. You're building the agent for us.

41:37 That's the driver that drives the car. You also have the simulator, uh, which is that virtual playground for the agent to learn in. And then you have the critic. And the critic is what rigorously evaluates and judges the performance of the agent and tells it how to improve. And the good news is that the fundamental reasoning and the generative capabilities of all three of those are shared and that's why in our case they're based on the same foundation world model.

42:11 Now once you have these three pillars you can create an incredibly powerful flywheel to accelerate your progress. So a deployment of your agent uh in the real world generates data. That data then grounds the simulator and makes it more realistic. The simulator generates harder age cases for the critic to score and for the agent to learn from. So the agent gets smarter, gets deployed in the physical world, generate more data and that powers the flywheel and accelerates progress.

42:44 But a flywheel of course will you know spin in any direction or in place. So in order to make it go in the direction you want you need to guide it by metrics. And that brings us to the final lesson that your model is really table stakes but eval and metrics that's your most important that's your strategic mode. So build your eval before you build your technology.

43:12 build your eval and your metrics before you build your product. If you can't quantitatively define what good enough means, you're not really building a product. You're just iterating on your demo. So nowadays, the best model architectures are fairly wellnown and new ideas tend to uh proliferate fairly quickly. Data is incredibly important, but without good metrics, you're just flying blind.

43:41 you aren't leveraging the best data and you can't really evaluate the ROI on making changes to it. So really eval and metrics that that's your foundation and that what steers your whole tech stack. Uh but for physical AI agents model level evaluation is not enough. When you're putting an AI agent into the physical world, your eval and your validation needs to go much deeper and much broader.

44:11 You need to evaluate and validate every component of your system from the physical layer to the behavioral layer uh that's running on board in the physical world as well as the offboard components and uh all of the operational processes around it. So for us we call that the safety and readiness framework and we spend years building and refining it and uh that's what guides our development and our deployment and our scaling and I consider that to be one of our most important uh assets uh be and this again the reason

44:49 it's important is because in the physical world trust is everything and eval and metrics is how you go about earning that trust. You don't just win trust by talking about, you know, the clever technical solution or the clever state-of-the-art architecture of your models or doing, you know, flash a demo. You earn it gradually day by day by uh in the field by relentlessly proving that your system is safe and that your system works.

45:18 And of course, you can just prove that to yourself behind closed doors. And this is exactly why we openly publish our safety data and our uh safety ongoing safety research. So then that earned trust becomes your ultimate business advantage, right? Your your models can be leaked, algorithms can be replicated, but hundreds of millions of miles of fully autonomous operations in the real world backed by evidence-grade evaluation and publicly audited proof that is much much more difficult to replicate.

45:55 So when you zoom out and look at this playbook as a whole, you realize that none of these lessons works alone. So the nine set your bar and ensure that you pick the right technology and the right technical approach so that you don't get stuck on the local minimum. Uh then intentional use of structure to boost scaling uh and the ability to ride technical waves of innovation.

46:15 I guess helps you get to the right level of nines uh and your AI ecosystem with the agent, the simulator and the critic guided by ER eval and metrics. That's what allows you to build that powerful flywheel and that's how all of these effects uh compound. And it's this playbook that we've been refining uh over the years is what allows us to achieve the strongly superhuman safety performance of the Whimo driver.

46:41 Uh this is a snapshot of the latest safety data we've released is based on over 220 million fully autonomous miles. And we're seeing there that in the areas where we operate, the Whimo driver is about 17 times better than human drivers when it comes to crashes uh that cause serious injury. And that really matters uh because today somewhere in the world every 26 seconds someone loses their lives on a road to a crash event.

47:13 And on the current scale what that means is that Whimo is preventing a serious injury every eight days. And this isn't just a metric on a dashboard. That means that someone's loved one got to walk through the front of the door at the end of the day safe and unharmed. So these are just the early safety uh benefits of AI in the physical world and they will only grow from there.

47:37 If you look at the broader landscape, the opportunity here is absolutely massive. Uh physical AI right now is where digital AI was a few years ago and we have all of the right ingredients to go after it. We have degenerative world models. We have the architectures. We have affordable compute and sensing. We have proven scaling laws. And we have a real product operating at scale.

48:01 And the last decade of AI happened in the digital world. I think the next decade will also happen in the physical world. And for those of you who decide to build in the space, uh, good luck, have fun, and remember who you're building for, your mission, and your customers. That's what matters. Uh otherwise tech is just a science project and at the end of the day as exciting as exhilarating the tech is nothing really beats the joy of making a difference in people's lives.

48:38 >> What are we doing? >> We're in our first ever Whimo. >> And what does it mean when we're in a Whimo? >> It means that there is nobody. >> Nobody driving this thing. And uh this is a fully autonomous Whimo ride. >> I cannot believe this. The car did a better job than the if somebody was driving. >> The truck was over the yellow line. So the Whimo break and moved to the side.

49:06 >> It knew how to pronounce my name. >> Oh my god. Look at this. >> Oh, it's nice. I love it. This is not >> This is so cool. >> I'll never forget this. Never. Sorry.

💡 Answer

A working demo is only about 1% of the work; the remaining effort is achieving the many nines of performance and reliability required for a safe, scalable product.

🧠 AI Summary

Building physical AI requires moving fast while shipping safely because errors can cost human lives, latency must be measured in milliseconds, real-world data is scarce, and deployment requires high confidence from day one. Waymo's experience shows that a demo is only about 1% of the work: achieving the many nines of reliability requires different architectures, redundant sensing, repeated technology upgrades, realistic closed-loop simulation, rigorous evaluation, and an ecosystem of an agent, simulator, and critic. Evidence-based safety performance and accumulated autonomous operation become durable business advantages.

🔑 Key Points

  • Physical AI must be designed for safety from day one because errors can cost human lives and cannot be undone.
  • Physical AI faces cost-of-error, latency, data, and validation gaps that digital AI does not face in the same way.
  • A demo is about 1% of the work; scalable products require many additional nines of reliability and performance.
  • Every additional nine of reliability requires about 10 times more effort and may require fundamentally different engineering approaches.
  • Waymo combines cameras, lidars, and radars so their complementary physics creates a more precise and redundant view of the world.
  • Technology upgrades should improve capability while also reducing fragmentation and simplifying the system.
  • Structure-augmented end-to-end models combine learned representations with physical and road-rule structure to improve scaling, validation, and efficiency.
  • A physical AI ecosystem requires an agent, a realistic simulator, and a critic, all guided by metrics and evaluation.

✅ Actionable items

  • Define the operating conditions, deployment parameters, required reliability, and safety bar before deploying a physical AI agent.
  • Count the required nines of performance and reliability before optimizing for demo views.
  • Choose technology based on whether its performance-versus-effort curve can reach the product's required level, rather than only its early progress.
  • Design hardware for future upgrades instead of anchoring the business to current component prices.
  • When pursuing a new technology, define its path into the full product if the effort succeeds.
  • Set launch criteria that require both breakthrough performance and radical simplification or unification.
  • Use structure intentionally alongside learned representations when it improves scaling, validation, and evaluation without unnecessarily constraining the model.
  • Build a high-fidelity simulator that supports closed-loop training and evaluation, including counterfactual actions and rare synthetic scenarios.
  • Develop an agent, simulator, and critic whose shared foundation model supports a data-and-evaluation flywheel.
  • Build evaluation and metrics before building the technology or product, and validate every physical, behavioral, onboard, offboard, and operational component.
  • Publish safety data and ongoing safety research to build trust through evidence and repeated real-world performance.

🧭 Frameworks

Four physical AI gaps03:53
  1. Cost of error gap
  2. Latency gap
  3. Data gap
  4. Validation gap
Structure-augmented end to end34:37
  1. Process multimodal sensor inputs
  2. Combine learned embeddings with materialized structure representations
  3. Validate correctness and safety at inference time
  4. Use structured representations for training, evaluation, and feedback
Physical AI flywheel41:40
  1. Deploy the agent in the real world
  2. Use deployment data to improve simulator realism
  3. Generate harder edge cases in the simulator
  4. Use the critic to score cases and guide agent learning
  5. Deploy the improved agent and generate more data
System one and system two architecture26:36
  1. Use a fast path for split-second geometric and safety-critical decisions
  2. Use a slow path for semantic and scene-level reasoning
  3. Generate predictions and driving actions through the decoder

🧰 Tools & AI usage

  • Cameras — Provide high-resolution and color information for vehicle perception.16:58
  • Lidars — Measure the three-dimensional structure of the surrounding world and support perception in darkness and other difficult conditions.17:07
  • Radars — Operate through fog, rain, and snow and directly measure velocity.17:14
  • Closed-loop simulator — Evaluate counterfactual actions and train autonomous driving agents through sequences of actions and world evolution.37:24

AI is used for

  • Autonomous driving — Operate vehicles safely and make real-time perception, prediction, planning, and driving decisions in the physical world.01:41
  • Simulation — Train and evaluate physical AI agents in realistic closed-loop and rare synthetic scenarios.37:32
  • World modeling — Model physics, dynamics, traffic, weather, semantics, sensing, and behavior for autonomous driving.38:28

📊 Numbers mentioned

Costs

  • A car traveling at freeway speeds moves about 100 feet in 1 second.
  • Each additional nine of reliability takes about 10 times more effort.

Growth

  • It took about 15 years to reach the first 100 million autonomous miles and about 7 months to drive the next 100 million.
  • It took about 8 years to expand from initial rider-only operations to serving riders in four cities.
  • Four cities were launched in one day earlier this year.
  • The Waymo driver operates in 15 cities across the United States.
  • The Waymo driver is about 17 times better than human drivers for crashes causing serious injury.
  • At its current scale, Waymo is preventing a serious injury every eight days.

Traffic

  • Around 500 trips per week
  • Over 4 million fully autonomous miles every week
  • Over 20 million fully autonomous trips served
  • Over 200 million fully autonomous miles driven
  • Over 220 million fully autonomous miles in the latest safety data

⚖️ Advantages, risks & lessons

Advantages

  • Hundreds of millions of autonomous miles backed by evidence-grade evaluation and publicly audited proof are difficult to replicate.
  • Multimodal sensing provides redundancy and complementary information across cameras, lidars, and radars.
  • A shared foundation world model can support different hardware generations and vehicle platforms.
  • Real-world deployment data improves the simulator, which produces harder training and evaluation scenarios.
  • A safety and readiness framework guides development, deployment, and scaling.

Risks

  • Physical AI errors can cause human injury or death.
  • Onboard inference must meet millisecond-level latency constraints within vehicle compute limits.
  • Rare edge cases become frequent operational problems at millions of miles per week.
  • A technology path can plateau before reaching the performance required by the product.
  • Adding new technology without an integration path can create fragmentation, complexity, and a dead end.
  • Deploying without sufficiently realistic closed-loop evaluation creates safety and validation risks.

Lessons

  • Move fast and ship safely when building AI that operates in the physical world.
  • Define product requirements and reliability targets before selecting the technical architecture.
  • Use technology waves repeatedly, but integrate them into production without regressions or increased complexity.
  • Favor general methods that scale with compute and data while using structure that channels rather than fights scale.
  • Treat simulation as a major AI system rather than lightweight tooling.
  • Metrics and evaluation are strategic assets because they guide development and earn trust.

💬 Quotes

A working demo is 1% at best of the work that you have to do.

Central conclusion about the gap between prototypes and real products.08:20

Count your nines before you count your demo views.

Concise guidance to prioritize required reliability over early attention.14:44

Your model is really table stakes but eval and metrics that's your most important strategic moat.

Defines evaluation and metrics as the durable competitive advantage.42:59

👤 People & companies

Dmitri Dolgov

Waymo Co-CEO identified in the video title.

00:00
Richard Sutton

Formulated the bitter lesson published in 2019.

05:02
Waymo

Company developing and deploying autonomous vehicles and the Waymo driver.

00:21
Google DeepMind

Its work on Gen3 is used by Waymo's world model to produce controllable, realistic scenarios.

39:24