← All transcripts

Why Robotics Still Isn't Solved - But Could Be Soon | YC Paper Club Transcript, AI Summary & Key Points

Y Combinator · 2 hours ago · Science & Technology · 01:24:13 · EN-US

📄 Transcript

Searchable transcript of Why Robotics Still Isn't Solved - But Could Be Soon | YC Paper Club — Y Combinator (01:24:13). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Y Combinator. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:05 [music] Welcome to YC Paper Club. How you guys doing today? [laughter] Okay. How do you guys like this one? We're we're we're going to change it every time now. So So today it's actually YC Robotics Club. Okay. So this is the 10th year of next year. robotics will be solved that I've encountered in my career. Um I remember when AlphaGo came out 10 years ago, everyone said next year, you know, clearly we have the algorithm.

00:37 All we need to do is scale it up. And then Mujoko came and then we had, you know, literally in in 3,000 uh iterations where we we can train a robot to uh to walk. This is quite amazing. Clearly next year robotics is going to be solved. This time it's different. All right, we figured out Yumi uh data collection. Okay. Uh and so the Aloha was like the big breakout.

01:01 Everyone thought next year we're going to have robotics now. I mean look at this. We can water [snorts] our plants. We can fix our bike up. We can do my curig here. [laughter] Even shave me, right? We have the algorithm. All we need to do is scale it up. And then 2026 was promised. They promised me that this would be the year of of the robots. I was very convinced.

01:27 I I read the diffusion policy paper. I even played around with it a little bit myself. I'm like, this is surely the year of the robotics. We have multi-step reasoning. This is like totally going to happen. Um we have VALAS. They're amazing. Um you just need tea ops data and then you're good to go. Um and I would say honestly, we're halfway through 2026 and you still can't Yeah, you can pre-order Neo 1X.

01:50 Um, I can't buy a pie or figure robot just yet. Um, we've had some success in work cells, but Rosie the robot is still not here. Um, and I I you know, we have six months left. It's only July, so maybe maybe it will be. Maybe I'll be wrong, but it's definitely the year of the demos. I know that for certain. So, the reason why is because, you know, raise your hand if you've ever done tea ops collect data collection.

02:17 How is it easy or hard? I mean, it's remarkably hard. You especially if you have this, you know, uh uh little gripper thing, that's all you have. You have the wrist cameras and all that stuff. It's extremely difficult. It's very very finicky. Um and if we're relying on this type of data, uh as and we need to scale that up like crazy. Um then it's we're kind of doomed.

02:40 And so I would I kind of think about it in these four there's probably more but at least these four walls that we have to scale and and get over and we really haven't yet. And so the first one is physical real world modeling. you if you take these video models and you uh deploy them to let's say drive a car and you you're playing you're doing the the world models you know paper by by Jurgens Smith Huber or the the Danar uh Hafner dreamer v1 2 3 4 and you do this game and you're playing play Doom in the simulated game

03:15 in the real world it really doesn't respect physics and so those models don't respect physics all that well and so if you're driving a car in that in that simulated world and you drive into like, you know, a grocery store or Whole Foods, it magically just kind of turns into a highway and then you don't crash. And so like it doesn't really respect the the real world physics.

03:35 And then and it definitely doesn't uh and so that's all called the sim tore gap and we really haven't figured out how to solve the sim to real gap. um deformable objects. Uh uh e even worse is that's just you're determining the transition function from ST to ST+1. If you condition further on the action, then it really doesn't work and you need a lot of data for act to support that action conditioning.

03:59 So if you you're trying to estimate the dynamics function T of ST + one condition upon both figure out some course representation. Uh Nico is actually here that we worked in 2013 on uh feature pyramid networks for our little robotic policy bot war simulation way back when. Um but still that was to solve this giant Q matrix that we would have some linear thing with the pyramids.

04:25 So representation for your action space is actually like a completely unsolved uh thing if you want to learn quickly. And then this is the most important one that I don't think is talked about enough. The sensory motor issue. Uh we have these uh nerve endings that that can do so much. We can detect the normal force. We can detect the tangent force. We can detect moisture, temperature, vibration.

04:49 We can estimate the coefficient of friction. And it's everywhere all over our bodies. And these robots don't have that. They have like one little coin FT on each fingertip and that's like all you got in like the best case and maybe you have a wrist camera here and like that's kind of like the the the state-of-the-art. And so if you talk to neuroscientists about this, it's actually like incredible how good we are at building world models without vision.

05:16 So, if I if you ever try to find your charger in your backpack and you put your hand in your backpack and you're feeling around, you kind of can tell exactly what's in your backpack just by feeling around. You don't even need eyes. And so, there's no way we have robots that can do that now because we don't have an epidermis. And so, I think that that's that is a really important thing.

05:36 Um, if you ever tried to, you know, tie your hockey skates, um, when it's really, really cold out, like you start to see the policy, uh, fail and you're like, I can't even untie my skates cuz my hands are too cold. And the last one that really only robotics, robotics people will really understand, people that have deployed real robotics for long periods of time, there's this embodiment drift.

05:59 And so in this state, if I take this action, this is how much force is actually going to be applied. And the actuators get dust in them. They get corroded. They get may have been if they're in like a, you know, um near the ocean, they might have some corrosion, some rust, and they just don't work as well. And especially true in self-driving car, this is like very real issues.

06:18 When I push the gas um on my like Toyota Prius, uh it it's variable how much I'm going to get and that shifts over time. Um and the amount of power you get out of a battery over time also also shifts over time. And so uh these are then you have to retrain your entire VA and so like because it's not mapping to the teleops data is almost um uh uh stale and you have to recollect it.

06:43 And so these are very real challenges that we don't have answers to until tonight. We have some uh some great talkers some great talks tonight. Uh we have uh Marcel who is a PhD in Chelsea Finn's lab and he's going to talk about some of the cool work that he did at uh physical intelligence. Uh we have uh Milan Gennai who's a PhD student uh with Marco Pavone and Clark Barrett uh currently working at Whimo.

07:10 And then we have uh Tyler who uh is doing his PhD with Janette and Karen. and Nico who's one of my close friends from uh 2012 when I was doing my double e double e masters um and founded this cool company called rerun and he's going to talk about some of the practicalities of data uh and how do you actually uh um deploy these these robotics and then we have Bill and Guan Ming who's um uh founders of this really cool comp YC company called General Instinct uh going to be talking about uh world action models real time

07:45 world action models uh previous research experience at DeepMind and so it's a great lineup. Thank you guys very much. Please welcome some of our our speakers here [applause] and then Marcel you want to come up. I'm Marcel. I'm a PhD here at Stanford and today I'm very excited to present you some of the work that I did during my internship at physical intelligence.

08:07 Um and we call this system MAM that is multiscale embodied memory. So I want to start with like some of the these are some of the policies that we trained when I was uh doing my internship there at physical intelligence and we're trying to solve these like robot Olympic tasks and as you can see they are actually like quite dextrous task. You can see the robots um it all fully autonomous.

08:27 You can see the robots unlocking blocks um make like folding clothes that were inside out making a a butter um peanut butter sandwich. And it's actually like very impressive to me, right? Like this they are super dextrous and all. But if you see actually the longest of these task is like about 2 minutes long. But ideally when I think about what I would want my robots to do is like I would want them for example to be able to manufacture something to be able to like clean a full bed which is going to involve a full

08:56 bedroom which is going to involve like making the bed um folding clothes like a super long task right or for example cooking a full meal including like cleaning the kitchen and all. And when we think about okay but what do we need in order to solve this like very long horizon task in robotics today. I can think about like a few things is for example being able to keep track of task progress.

09:17 Um being able to keep track of time but also having like very reliable dexterity and dexterity that can actually adapt in context in case that the robot like finds a new scenario and makes a mistake it should be able to adapt to it, right? Um and a few more things probably. But my claim is that in order to obtain all of these all of these insights, we actually need to add memory into our policies.

09:41 So memory is kind of a for me is a kind of a very necessary thing for this for solving these long horizon tasks. However, if we look at most policies like PIO5, Groot, all of these policies actually don't have any memory. um which means that basically at every time step the robot will obtain like a new set of of observations and has absolutely no context of what happened before.

10:02 And here I can show a couple a couple examples of what happens when you train these policies without any memory. So on the left you're going to see a robot that is washing the dishes and just like has no context of how long it has been washing the dishes. So it just keeps like washing forever. or on the right you have one where like the robot is cooking a grilled cheese but again like has no context that is that how long it has been there and it becomes like fully burnt um but then you might ask okay like why are we

10:31 not adding memory into the into the policies right and well one of this is it is hard so I'm going to take a small detour at another paper that we wrote last year where we like try to we basically observe that there's two main problems um one is effectiveness is basically when we add memory into this policy pieces um they actually perform a little bit worse because of distribution shifts and lack of data um and an efficiency problem which basically means that when we increase the context for the robot policies um this

11:00 actually becomes much more resource intensive. So then you're going to get like much longer training time and it's going to be also much longer to run inference. Okay. So having said this what do we propose and what is our solution right? So propose compression um and we take these basically this model that is we're going to decompose our robot policies into two parts one that is going to be a highle policy that basically tells the the low-level policy which is the next step that it should take and then we have a VA

11:31 like a low-level policy that is going to actually execute the robot actions and then we're going to decompose the memory into two different types one that is going to be short context and it's going to be like some dense frames that are needed for like the actual dextrous manipulation and it's going to be fed into the low-level policy and then we're going to have a long context memory.

11:49 There's going to be basically a compressed language representation and of the last uh few minutes and it's going to go into the into the high level policy. I don't want to go too much into details but I'll I'll give an overview of like how do we add this short-term uh visual memory. So our idea here was to design a new encoder that is based on the VIT, but instead of just taking taking a single frame, we're actually also going to add some attention temporal layers.

12:14 And we're going to drop all of the all of the tokens except the current image, which should have all of the information necessary there because of this temporal tension. And then you basically get a lot of of compression. And then some of the reasons why this architecture is actually quite good is because first like you have an easy vit initialization since only now the temporal attention is new.

12:34 We actually get to compress again this image sequence on like we reduce the number of tokens. And because of this you get fast inference. And now I want to showcase like what task can we actually solve with this short-term um like with this short-term dense memory. So here we have like our main flagship task that we did when I when I was at physical intelligence and it was like making this grilled cheese.

12:55 So it's going to be able to like do all of the dextrous parts but also be able to wait for as long as needed um in order for the grilled cheese to not be burned. Um but we also can get to solve some other types of tasks such as here that maybe they seem like they wouldn't need memory such as unloading groceries from a from a grocery bag but here you actually only see the items inside with the with the wrist camera.

13:18 sometimes. So you actually need to remember where the items are and how many items there are. Or another example is like cleaning and cleaning a window where again you need memory in order to not stay there forever. So now I talked about how to add memory into the low-level policy and next I will talk about how do we add the longer longer term memory into the highle policy.

13:40 So this is the me the technique that we proposed which basically consists on kind of a um the high level policy predicting a recurrent um memories like memory scratch pad basically it's going to it's basically going to keep track of everything that has happened and then whatever its prediction was is going to be fed again into the highle policy and like this it can keep explaining like what has happened before and what it can remember also what happened.

14:04 So this is actually like much more com a much more compressed representation than images since text uses like way less tok to tokens than images. So this is great for for training. Um it is also quite physically accurate because the policy is going to be able to modify its memory with whatever has happened. Um and it is also less prone to distribution shifts.

14:25 So let me show you here like what can we actually solve with this with this long-term memory. And here we we show a task that like takes up to like tens of minutes. Um, and it actually is going to prepare all of the items for for making uh for preparing a recipe. Yeah, just for the sake of time, I'll I'll skip over it, but um yeah, something nice is that you can actually see the memory string that is predicted on the on the green box.

14:52 We also compare with a bunch of different baselines such as no memory, different types of memory, and we also get to beat here the the state-of-the-art. Something that I want to that I that I'm especially excited about when with regards with memory is it's it's that actually when we add memory into VAS, we're going to be able to get this property of in context adaptation.

15:14 And it's actually something that is really lacking right now into VAS and I think it can be super promising in the future. So, let me show you what this corresponds to. Here we have some policies no memory like they have no memory at all and you're going to see that they are going to keep making the same mistake over and over and they are incapable of reacting to the mistake they did.

15:33 So like it's just going to be stuck in a loop forever and never going to be able to to adapt. So here it's trying to open the fridge or picking this chopstick forever. But actually now when we add memory we're going to see that the robot policies are going to make a mistake in the first time right like just maybe they have some bad prior or something but they are going to actually be able to see this mistake and react to it.

15:54 So for the chopstick it like made a mistake but then it's going to go go lower to pick it up and for opening the fridge similarly it's going to switch sides um to correct its mistake. I think this is something that is very very lacking now from the robot policies and I'm super excited about um the future with like VAS with memory. Having said this um yeah this was a huge effort at physical intelligence with a bunch of collaborators I'm super thankful with and I just wanted to point out Carl who was the co-first author

16:21 here but then Homer Sergey Chelsea and Danny for all their help. Please let me know if you have any questions. [applause] >> Thank you for the presentation. I was curious about the the long-term memory that you were talking about. If it's only represented in the textual space, how did you figure out the right information to actually, you know, give to the policy?

16:46 >> Yeah, that's that's honestly that's an awesome question. So, right here we train our our highle policies with SFD. So we had to annotate all of our data but it was that's a very good point because basically we had to think beforehand what information we think is important here and we had to tell our annotators okay like you know you need to keep track of all of these things because this is what is important but I think that some of the follow-up works that I'm thinking about and I think like everyone should like

17:14 think about and all is how can for you could you for example do reinforcement learning on this memory space right in order in order to be able to know what is the right information to keep track of, but at least the first proof of concept with SFT, but >> yeah, exactly. I think that was going to be my suggestion too, you know, like figuring out the right message to store the memory, the right things to recall.

17:34 >> Yeah. >> And a follow-up question to that was how are you actually right now storing that information? It's it's represented in textual space, but like how is it actually used in inference time when you actually use the policy to take actions? So it's I mean it's a fairly short text. So this can all be kept in in RAM and then like you just feed it as normal text tokens to the to the VA.

17:53 Yeah. >> So one question uh when you extract your task into a highle text description how you make sure that the tax itself is generalized enough to handle different virus like when you make omelette there are different versions of omelette right and also how does this memory affect the number of episode you need to train for a novel task. >> Yeah thank you.

18:15 Ah great point. So for the here I don't show the the examples but we actually give some quite the the like detailed descriptions of the task. For example for the preparing the ingredients or for making a recipe and everything you give you tell exactly where the items are like what are the exact items for making for making a pizza for example and all these.

18:36 So that is very detailed. Something good about this is that the highle policy is a VLM trained with internet data. So it actually doesn't need that much data in order to generalize well. And the good thing is that the VA is like completely separate um and doesn't need to like we need to train it per task, right? To be able to do things in the kitchen and all this.

18:57 Um but if you make the task more complicated, the VA is still only receiving a small like text description of what it should do next. Um so at least we don't have to collect that much robot data with like the complexity. >> Really appreciate the talk. Uh so your embedding for the memory is textual descriptions, but it also seems like you could have solved the problem with just adding, for example, for uh doing a grilled cheese, just adding time.

19:22 And so it seems that training a puddle a policy just only based on textual descriptions like limits the representation space for what that memory can describe and potentially constrains it. Have you thought about actually expanding the memory representation or looking at other approaches for that? I I'm I try to think about how do humans right even like keep track of all their memory and it's definitely not in text space.

19:44 Um but I think like something that is quite hard and like I feel we haven't really managed to do it like very well. Ideally we would have a laten embedding that would like just keep track of all the memory. Um but just with text you can actually put a very strong bias there and be able to supervise it, be able to debug it and everything. uh which just makes it right now the most practical way to do it.

20:07 But yeah, I'm very excited to like try to explore this further and can like do some latent for example work there. >> All right, thank you Marcel. >> Thank you. [applause] >> So yeah um PhD student at Stanford uh research at AWS and Whimo and I'll be talking about how can we move toward robots that teach themselves how to reason. You're probably familiar with vision action models, but just a quick primer.

20:37 Um, VALA's are a powerful uh class of generalist policies. Probably seen them in demos for manipulation like RT2 and PI or even self-driving. Uh, maybe you've driven a W or you've been on a Whimo um or Alpomeo which is from Nvidia. So, how exactly are val trained? So, you take um a vision language model. Um these are multimodal models trained with an internet scale of um visual and textual prior and then this can be like Gemini or Quen and you continue training them on relatively scarce robotics data sets and so this

21:11 can be teleop for manipulation or maybe self-driving um someone has driven a car and recorded that and then you end up with a VA which takes as input an image or some perception feature maybe language um instruction prompts and learns to generate actions which can be executed on. So like steering commands or endeector position. And so now there's this recent trend of um leveraging and doing embodied reasoning for better action prediction.

21:38 The idea being similar to chain of thought for the LLM land where you go from question to answer by um explicitly providing some sort of logical steps. Um and so why are we interested in chain or embodied reasoning in the form of train of thought? The idea is that there's not a lot of data in robotics. um it's data scarce and so any form of signal um that you can use to augment your data set um is very valuable.

22:03 So you can start injecting different types of annotations and that can help you um improve action generation and whatnot. So you got richer training signal signal with um reasoning but also because this is in text um as a human you can go and read it. So if you're trying to decipher what was the reason for uh why a robot made a particular decision, you can actually read through uh the chain of thought trace.

22:27 Uh to this end, I'll be talking about our recent work published at robotic science and systems conference uh this year called self-supervised bootstrapping of action predictive embodied reasoning. And so we're really interested in this question of what should we be reasoning about? Specifically, what should particular embodiment and form factors reason about?

22:45 So what makes good embodied reasoning is a hard question. Um but I'll decompose it into two problem. One is the grounding problem. Um we don't have um this sort of uh oracle source of reasoning data saying that just train on this reasoning data and uh you'll get a good reason there it's not clear what we should be reasoning about. Um if there exists some traces on the internet just like there was some pre-training documents for text or captions for image that doesn't exist for robotics.

23:12 Um, second is where's the oracle source of the model? You can't just pry open a human's brain and figure out how they made a decision u from an image uh to the movements in their fingertips. So that's why I like to uh show this image of a chicken and egg problem which is where is the source of uh the model and the oracle source of the data. The second is associated with verbosity.

23:31 So there are different types of reasoning. You can be planning right. So you can say go to the pepper, pick up the pepper and so on. Um maybe you you would be reasoning in in the form of perceptual um traces. So this can look like visible object list. So list of bounding boxes of all the objects in your scene. Um or even gripper position. What's the position of the endector or what's the position of your car on the road?

23:53 But the question is should we be planning at every step or is that verbose? Because latency is a big problem in robotics. So it's not clear if we should be planning or should we analyze every single object in the scene or is that distracting and could mislead us in action prediction and similarly is gripper position reasoning action predictive um or is it misleading all of these questions can be summarized into this one research question of how does textual reasoning for specific um form factors and embodiment look

24:22 like and so that's where our approach comes in it's called R&B encore uh which is short for refine and bootstrap embodiment specific chain of thought reasoning. This is a self-improving pre-training cycle for embodied reasoning VAS. And so the insight is because we're claiming that reasoning is sort of this black box for robotics. Um we treat this as an unobserved latent variable for the observed context and action.

24:44 And by doing so, we can leverage this sort of theoretical framework um called variational inference. So how does this look? There are two highle components. So one is a reasoning proposer. You can think of this as a annotator model um which proposes for a given context for a g given demonstration um various types of proposed reasoning. So it it's asking the question is visible object and move reasoning a useful type of reasoning or should you be reasoning about plans um and gripper position or maybe visible object and

25:14 subtask reasoning what's the sort of reasoning um or annotations or good data that you should be producing before your actions. And then the second component is a reasoning validator. The idea is that this is a um scoring metric that is based on this theoretical ground of variational inference. Um there's all this theory that we've proved in the paper.

25:34 But uh to concisely summarize it um there's there's three main things that main parts to uh the score. So one is concision ensuring that your reasoning trace is um short and not you know too verbose. Um the second is um non-trivial. So this is some sort of encouragement of uh generalization uh in the reasoning behavior. The third one which is most important is action predictiveness.

25:59 This ensures whether that reasoning trace or that annotation is grounded in that embodiment. So at the end of the day you score all these reasoning traces and you end up with uh and you resample and you end up with a new data set of synthetic but action aligned and embodiment aligned reasoning. So you can analyze this reasoning data set. you can understand okay this is the good type of reasoning and this is not but more uh importantly is you can retrain your embodied reasoning VA and have a better um and more robust uh

26:28 policy so we tried our pre-training cycle on a bunch of uh embodiment so we look into obviously manipulation so we pre-trained uh embodied reasoning manipulation VAS we find that move reasoning and gripper or move and gripper position and uh position type of reasoning is very useful whereas perceptual reasoning is not very useful. So we're able to prune that out.

26:49 Um and we improve success rate because we're having this sort of action aligned um form of reasoning. We look into more about the question why perceptual reasoning is not useful. Uh we we find out that there's a lot of distracting objects in many of these scenes. And so by pruning out using our pre-training cycle, we can actually look into what is a good way to improve your annotation data quality.

27:10 So questions about improving and task salency of your traces and the data that uh you use to collect and annotate um thereby uh not only improving success rate but also this term of like object criticality rate or task salency rate. We also test this uh entire pre-training cycle for hardware uh manipulation uh VAS and so we're able to have improved out of distribution performance particularly on objects uh novel um target objects or even uh cluttered scenes and we also pre-train legged locomotion navigation models.

27:41 Uh we find that reasoning about structural affordances and movements are way more important than reasoning about terrains and counterfactuals. And then we also use our uh improve our cycle for refining human annotations for self-driving. Um and so there's a lot of garbage data that could be out there. And so the idea is that with um our approach you can remove a lot of the human uh annotations that are not very useful.

28:07 Um and so we able to observe that meta action invisible objects and perceptual reasoning is useful whereas our approach can prune out sort of hallucinated experiences. um thereby lowering the L2 path errors and uh performance metrics like collision rates. So takeaways, selective reasoning is way more important than exhaustive reasoning. Um even if that reasoning is valid, it's not necessarily useful.

28:30 Second is that self-supervised bootstrapping works. This is really important to enter into that chicken and egg problem of like the oracle source of model and reasoning data. And third, we show uh the approach generalizes across embodiment like manipulation, navigation, driving um as well as VA sizes from 1 billion to 30 billion parameter models. We basically are addressing this question, how should an embodied agent reason?

28:52 And we argue that embodied reasoning is not a fixed template um that should be applied uniformly, but rather some a resource that should be discovered, deployed, and budgeted carefully. We are very excited about this problem about how we can deploy embodied intelligence. We're very excited about problems like data quality. Um how can we recover from failures?

29:13 How do we adapt to novel scenes and specialize in them? So if you're interested in some of these questions, feel free to reach out. Um we're excited for um collaborations. Uh we would like to help you on problems on reducing the friction of deploying embody intelligence um for everyday tasks. So if you're interested in reaching out to us uh you can scan the cure code uh on the right um and more details about R&B encore um including the website models code papers on the left.

29:43 [applause] um how do you think about uh the conversation we had whereas like you know a dog is not going is figuring out all of this but not in token space and so how do you think about this without using tokens >> for oh okay so I think there's all sorts of like latent reasoning capabilities where the model can sort of navigate in its sort of continuous space of understanding um in its own unhindered u by uh you know the the textual format And this requires good architecture and theory for developing this sort of

30:22 continuous reasoning. Um but there are ways of if if I'm understanding your question correctly, non-extual reasoning. Um and I think some of these approaches can really help with that. Um the idea of here um is that we want to leverage particular priors um that exist especially with um LLMs and BLMs where there's huge amounts of data out there. We want to use those priors u for data scarce regimes um like robotics.

30:44 wanted to understand what's the thesis behind using uh reasoning for autonomous vehicles because there's already sort of uh the streaming latency problem which is even more exacerbated in the case of uh AV. >> So in that case like is your argument just that it would help uh build better uh self-supervised data >> or um is there something more there uh that can be leveraged for better performance?

31:13 >> Absolutely. So latency is a big question which we actually uh address in this paper. Um our main claim over here is that reasoning can help um introduce priors into the model. Um so this is you can think of this almost like co-raining data where instead of just training on robotics demonstrations, you can co-ra and um use re textual reasoning as good annotations to help it.

31:34 But we actually show that um we have this approach called action forcing uh where you can drop the reasoning during inference time. Um so you don't have the inference uh latency problem but at the same time you can still leverage or extract the the enhancements from the sort of textual annotations and reasoning. Uh so so you can get the best of both worlds.

31:55 >> And a follow-up to that is do you see drift um like in terms of state space? So you reason about something the state has moved now uh so the reasoning is no longer relevant or can damage uh future actions that the model might take. uh do you see situations like that as well? >> So the reasoning happens um is supposed to is it makes assumption that's happening in real time.

32:16 Um so it's based on the current image um and the current instruction and so the drift isn't necessarily a problem that's that's going to fundamentally appear in here. >> Yeah, >> thank you so much. U I just had a minor point. um when you were talking about pruning uh sort of uh things that weren't too uh useful for the model uh you made a point on how counterfactuals weren't too useful uh in terms of better reasoning if if I'm right uh could you maybe just you know explain a little bit more as to exactly where where

32:47 you were going with that I >> yeah so um what the claims that are being made um about this is one um focused on the particular benchmark that um we we're working on. Um it isn't to say particular types of reasoning shouldn't ever be done and should be removed. Um but it is to say that there are certain types of reasoning that don't need to happen at every single moment.

33:10 So to give you an example, if there's nothing novel or interesting happening in the scene, like a lot of driving, self-driving is just drive in a straight line. So and there's nothing around you. And so there's no interesting uh counterfactuals to uh that you might explore. But even plan reasoning is or I had mentioned earlier plan reasoning um that appear to be quite useful u or it seems that it might be useful to plan ahead uh but you won't want you don't want to be planning at every single step because it's quite

33:36 redundant. >> Thank you Milan. >> Thank you. [applause] >> Hi everyone. My name is Tyler Lum and I'm excited to present our work sim tool reel. So let me start with what sim tore can do the volume here. So every clip here is at 1x speed and this is a single policy that is working zero shot meaning it never saw any of these tools or tasks during training.

34:00 We do not need to retrain for a brush or hammer or a new target behavior. And many tasks like the screwdriver spinning are very dextrous. They would really require a multi-fingered hand. They just really wouldn't be possible with a parallel jaw gripper. The policy is very fast and reactive. It runs at 60 Hz and simultaneously controls both the 22 degree of freedom hand and seven degree of freedom arm.

34:24 And I also want to highlight that the level of dexterity is very difficult to demonstrate through telly operation here. So with that preview, I want to give a little bit of context. So telly operation for dextrous hands has increasingly been used to collect demonstrations for imitation learning. But highly dextrous actions remain difficult to demonstrate reliably and at scale.

34:44 In the video on the right, even this simple inhand rotation task requires slow deliberate control from the human because of the embodiment mismatch and limited force feedback which make precise contact regulation very difficult. So rather than learning the policy from telly operation, we train it entirely in simulation using simtoreal reinforcement learning.

35:05 So sim to real reinforcement learning uses GPU accelerated simulation to run tens of thousands of robots in parallel and generate experience about a thousand times faster than real time. This allows us to scale data collection with compute rather than human effort and collect decades of interaction data in only a few days. The result is extreme dexterity because it is not just imitating demonstrations but is optimizing for reward maximizing behaviors.

35:37 These policies are then deployed in the real world demonstrating impressive dextrous behaviors that would be very difficult to teleoperate. These videos are from Dexream which is one of the first works to demonstrate the effectiveness of this approach for dextrous manipulation. And this task isn't particularly particularly useful looking, but it's definitely an impressive demonstration of dexterity.

35:55 But most prior works in this space learn a separate policy for each particular skill. So one for grasping, one for reorientation, another for object spinning, and another for tool use. Each new behavior typically requires a round of new reward design, new task specific engineering, and another round of training. So we instead ask, can we train a single policy just once and have it perform all of these different tasks and skills?

36:22 So what we really want is a single policy that controls both the hand and arm through the full sequence. It first grasps the brush, reorients it within the hand, and finally uses it to sweep the objects. And we want this all from a single policy. So we don't need to do any manual switching between separate policies. Our key insight is that we can unify dextrous tool manipulation as goal reaching.

36:49 So the policy doesn't need a task label such as sweeping or hammering. it only needs to move the object from its current pose to the desired pose. So thus we train a goal condition policy that can move arbitrary objects through a sequence of desired goal poses visualized here in green. And we find that this is a very general objective for a wide range of manipulation tasks.

37:12 So at train time, we procedurally generate primitive objects in simulation, sample random goals, and train a goal condition policy with massively parallel RL and simulation to manipulate random objects to these random goals. And of course, there are many details to get right here. The system identification, domain randomization, RL or algorithm details to get exploration right.

37:33 Um, we're going to skip all of that for now, but feel free to ask questions in the Q&A. But then at inference time we need to specify the sequence of goal poses which can come from any source. And in this work we choose to condition the policy on a human video demonstration. So we use foundation pose and SAM to extract the sequence of goal poses to track.

37:52 Then the RL policy tracks these goals one by one in a 60 Hz control loop. And I really want to highlight here that the human video is not providing robot actions and it's not used to train or fine-tune the policy. It only specifies the desired object trajectory for the frozen policy to track. So concretely, the policy takes in propriception, the current object pose, a bounding box of where it should be grasped, and the current goal pose.

38:18 It runs them through an LSTM policy, and then outputs joint position targets for the full hand and arm. And this policy again is not limited to just sweeping with brushes. This single policy works zero shot across novel tools and tasks never seen during training. So what's really nice about this is that a new task simply becomes a new sequence of goalp poses rather than a new training run.

38:39 So conceptually the trajectory acts kind of like a task prompt that we can provide at inference time for a frozen policy allowing us to perform a new task on a new object in minutes instead of hours or days. So we evaluate the same frozen policy across 12 unseen tools and target behaviors and achieve substantial task progress across every tool family particularly those with long handles.

39:01 performance is weaker on heavier tools which are easier to drop and also for smaller objects. But this is really because the pose tracker has a lot of problems when it gets very oluded. So next we want to measure train and test correlation. So our goal is to see how well our training objective supports downstream test success. So on the left we show training objects consisting of primitive objects and random goals and on the right we show our test objects with human demonstrated trajectories.

39:26 So we evaluate checkpoints throughout training and as the policy improves on the generic goal reaching task with primitive objects shown on the left performance on unseen tools and human demonstrated trajectories on the right sharply rises. This validates that sim tool real's training objective with random objects and random goals is effective for generalizing to real world tools and tasks.

39:48 So next we compare sim tool real against two really common baselines. So our method successfully grasps and reorients the tool into a functional pose to complete the task. The fixed grasp baseline can grasp the object but it must rotate it using only the arm resulting in a table collision. And this I think really highlights the importance of in-hand reorientation ability as most prior approaches kind of assume that just acquiring and maintaining a fixed grasp is sufficient but it really isn't for many cases.

40:16 And lastly, kinematic retargeting. This is where we try to imitate the human video demonstration by transferring the human fingertip motion to the robot, but it doesn't reason about contact forces, so it fails to even grasp the object. Next, we analyze the failure modes of SIM to a real and find that pose tracking failures dominate. Next, object gra dropped after being grasped.

40:36 And lastly, failed grasp, but it really tries to chase it down. But the policy demonstrates really strong recovery behaviors. So here when the robot drops a hammer, it immediately regrasps it and completes the task. Lastly, please check out our website. All of our code, assets, and policy weights are open sourced and we even have an interactive demo that runs right in your browser and even works on your phone.

40:58 So it's not running on a separate server. It's running on your phone. It will drain your battery. So don't leave it running for a long time. But it's it's pretty fun. So that was a quick overview of SIM tool reel. I want to spend the last couple of minutes um talking about a follow-up work called play to perfect. So many real tasks like precise assembly require sustained contact and millimeter level precision.

41:20 But learning the skills for precise assembly from scratch is really difficult. So we argue that before we can learn the hard problem of precise assembly, we must first learn the easier problem of playing with objects in free space. So this motivates play to perfect a framework that leverages the familiar pre-training fine-tuning paradigm. We first learn a shared dextrous prior through task agnostic play which is very similar to the sim tool real task agnostic training.

41:46 We then fine-tune that prior on a sparse reward contactri assembly task and deploy the policy zero shot in the real world. This enables diverse contactrich assembly behaviors including tight insertion multi-art assembly. So in conclusion, Sim tore real enables broad reactive dexterity across novel tools and tasks while play to perfect extends this towards precise contactrich assembly.

42:18 Thank you. >> [applause] >> Yeah, I noticed that in the demos the recovery was I mean insanely impressive to say the least. Um did y'all is that something yall specifically uh aimed for uh for the robot to have good recovery or was that just an accidental byproduct of the policy? >> Yeah. Yeah, that's a really good question. So we didn't specifically train for that.

42:46 But one thing that we really importantly added like one of the details about domain randomization we added was that in simulation we added these random forces on the object that randomly pushed at it. So sometimes it would knock it out of the hand so it gets experience having to pick it back up again. If it we didn't add that I think it may not be as good at recovering because it might never drop it in uh in simulation.

43:05 >> How similar did y'all ensure the primitive tools you used were similar to like the normal brushes and like spatulas you use? >> Yeah. Yeah. Great question. So this called sim tool reel because we're trying to focus on tool use and many tools have some sort of grasible region that can be somewhat approximated with a bounding box. So that's our kind of interface here that we're using.

43:24 So we're not giving it detailed like object geometry. We're just telling it roughly like what the bounding box size is of the grassable region. So in simulation one detail here is that we're using all primitive objects just cylinders and cuboids. We could extend it to more. Then you have to generate a whole other data set and like remove stuff that's not stable.

43:42 But the one key advantage is that the simulation runs like two to three times faster when you use simple objects like that. So that's what we used here. So things that can be reasonably approximated as a cuboid. So something like a spatula or um even like a sphere would probably be fine. But if you're trying to pick up say something like a bowl or maybe scissors, it probably could pick it up in some weird way, but not in the way you'd want it to probably.

44:06 So yeah, I'm super interested in um this kind of a novel approach rather than learning from human demonstrations. Basically learn a bunch of sub goals and then use RL to um reach to those sub goals. >> But one problem for this I think is probably like how would you able to generalize right to a lot of real world um objects other than tools right because for tools you can easily define those primitives.

44:31 Have you thought about those like um for example like articulated objects right? How would you able to do those? >> Simulation has its pros and cons like a lot of some like one answer would be we can maybe just simulate all of them you know simulate scissors and simulate like all kinds of other articulated tools but there are just so many things that cannot be simulated well like water or like even a zipper or even like cracking open a can of bubbly or something.

44:55 That kind of stuff I think is not currently in the realm of simulation. So like what we do there is a really good question like can we can we transfer these priors into the real world and keep fine-tuning it. I think that's most likely the way I would do it. But that's a really good question like how to integrate this kind of great dextrous behavior but let it keep improving with real world experience.

45:15 Yeah, I think my question is more on like even if you assume you can simulate all the objects in the world but like with this paradigm are you able to like generalize right so just still RL um >> yeah it's a really good question I I think if you take a really far step back right like our policy is being given the current pose and the goal pose right and it's kind of like an inverse dynamics model in pose space but if your object can't be specified as pose maybe if it's something like like a towel.

45:45 Maybe you can specify it with key points or maybe the most general version of it is like you have a video goal and you're almost an inverse dynamics model of like the current state and like the final goal state. It's probably possible to like train something like that. But exactly how to get those details right like generating the goal at inference time is a really hard problem but I think like your question is really good one.

46:05 Yeah, like how to make that more general. >> Yeah, we're trying to basically reproduce this. >> Oh, cool. Cool. Yeah. Hey, uh, I'm trying your interactive demo right now and I think like I've been trying to poke around and find failure modes and I found that like the only failure mode that like consistently shows up is when like the arm like continuously twists like try to find different objectives and like it it fails usually when it twists the arm itself like too many times.

46:27 I was wondering if you have like a reset dynamic maybe like between actions where like maybe let's say like my arm is here and in order to reach like this state like I could turn it backwards instead of trying to twist it even further. >> Yeah. Yeah, good question. Yeah, I honestly not sure exactly what the right solution is there. Like it should in theory like if RL is really good, it should know to like spin it all the way around again and it gets really contorted, right?

46:47 Um I don't know the exact solution to that. I guess most real world applications you don't go like crazy contorting your arm, but I think yeah, it's a good it's a good question. >> Cool. Thank you. >> Um okay, first of all, this is amazing. Congratulations. Um my question is on like from the team's perspective, how do you guys iterate on such a thing?

47:05 For example, like if you have one policy and you start seeing regressions and other tasks, >> how do you guys handle that? Do you have like evals or like what does iterating even look like in this space? >> Good question. Yeah, there's a it's a bunch of things, right? Like how do you eval these kind of policies? You can eval in the real world, but it's very expensive and takes a lot of time.

47:23 And a lot of times like you don't actually get statistically significant numbers, right? You can almost get a vibe check unless you run it like truly 100 trials. It's hard to like tell if you're at 90% or 87% success rate. Um and the other thing one thing we really try to do is have a lot of automated ways to evaluate our policy across all those novel tools and tasks in simulation.

47:42 So it gives us a sense we train on the random objects and random goals and see how well it works on our real world objects and real world trajectories but we like put them in simulation. So we get some metric a feel for if it's doing way worse or way better. But um it's a good question because honestly it's really hard to tell. Sometimes it actually can do occasionally it can do better in sim but at doing some magical tossing and catching behavior that you know may work but our post tracker would probably fail.

48:07 So um you kind of need a little bit that human insight still. Yeah, it's a really really hard one. >> I have a quick question. Um so you use this uh prehistoric thing called an LSTM. Uh can you tell us about what is that? >> Yeah, it's a long short-term memory. So I can I can go into the math, but basically um there's RL packages are um how do I say it?

48:30 RL is a very finicky thing where you almost don't want to change too much of the code because once it works well once you don't want to rewrite it from scratch because any one like value you change could break the whole thing and it's really good to start from a good you know starting place and then iterate from there and that codebase already had LSTMs baked in.

48:47 We're like let's just try it out and it worked better. Yeah, there's probably other ways to integate >> LSTMs work better than transformers. Well, there's no transformer in that one. >> Yeah. But but do you think that transformers would have worked better? >> I've talked to people about this. I think that broadly um transformers are probably better data sponges if you have unlimited data.

49:02 But here I feel like we're not we're getting a lot of data, but like we're constantly updating the policy. So I feel like it's actually not in that regime where we need an enormous huge data set. And some people have actually tested this and not shown any improvement. So I haven't had the motivation. >> Yeah, agents are out now. So probably no excuse for me to not try it.

49:17 But >> the diffusion policy paper, a lot of people don't know this, but if you go through their their table, the comm actually outperforms the transformer on like half the policies. >> Yeah, unless you tune the transformer better. It's actually much more sensitive. So I think yeah, exactly that. >> And then the last question, I didn't I didn't really understand how you create random goals.

49:34 What does that mean like actually in the code? Like what do you actually >> Oh yeah, it's it's like the simple thing you can imagine like we sample a position, we sample a rotation and then we could put that at their first goal and then every subsequent goal is some delta pose of like up to 10 centimeters away and up to 90 degrees difference. Yeah, >> I see.

49:51 So that's really good for training the policy to get from A to B, but it's not good for generating the goal. So to generate it to generate the goal, you need human labels. >> Exactly. Or you could do something else. So you could actually imagine like some highle planner like looking at the scene with understanding the full context and then generating the goals for you.

50:07 I think that's probably a pretty interesting direction. Yeah. >> Uh you were mentioning uh some of the failures were due to wrong estimation of the pose. >> Oh yeah. >> What fraction of those would be attributable to that? And second question is uh are you always using the third person views or did you also perform any experiments with first person views?

50:26 >> Good question. Yeah. Um, about more than half of our failures, I think roughly 60% of our failures were purely from the post tracking. It was really like one of the big bottlenecks of the system. It's honestly not my favorite part. Um, but it really happened the most on these smaller objects. So that that marker, that one is really easy to be oluded, right?

50:44 You barely see it at some points. So I think that's like example of one that'd be really hard. So like this kind of longer leg or this like bigger brush has a lot more features. So those would be having having less pose tracking issues. So that's like roughly how that looks. Were you using like the third person view or did you also experiment with first?

51:00 >> We use a third person view roughly like where this camera is and then that's pretty much all we tried. Yeah, we just kind of found a good angle where it didn't seem too occluded and then we just use the post tracking from there. I think it could be better. Maybe you could even use like a post tracker from the camera. You probably can distill this to some image based policy but there's some details about getting the goal conditioning right.

51:15 >> Awesome. Thank you. >> Thank you. [applause] >> Yeah. Niko, I'm the one of the co-founders and the the CEO of Rerun. in a prior life. So we run started that about four years ago. In prior life we used to do machine learning computer vision for physical world applications for yeah for shipping products like that for for about a decade before this company but rerun we are building this sort of unified data layer for physical AI basically tools and infra to help you work with physical data from collection all the way

51:46 to training. So this open source SDK that's pretty uh popular. I know a bunch of people in here uh use it for basically working with physical data. So it's like logging, visualizing uh sort of generic querying and sort of uh data loading and and stuff like that for training. So basically all the kind of tools that you need to to transform and and sort of analyze data and we have a info product which is called ron hub which is basically a data catalog and sort of large scale data backend for doing all those same things

52:14 but for lots of data on sort of in the cloud. As part of building this uh we get to talk to or work with like amazing companies doing robotics from like frontier labs all the way to you know two person YC startups and uh there is a sort of there are many ways of doing robotics companies and robotics projects but there one pattern uh that we're seeing a lot of that really really working right now that I'm I'm super excited about and I want to see way more of so this talk is actually um mainly me trying to tell tell you

52:45 all to start companies like this. So basically that new category is what you might call just robotics application companies. I think some some people call these like the neo integrators and their pattern is kind of really like taking ownership of a full business problem sort of end to end. Uh right now this is working see this working a lot in like data center management and construction like warehouses, manufacturing things like that.

53:10 Just making being like really really excellent at operations like deploying support things like this kind of thing. Building as absolutely little custom hardware as possible and then just generally often starting with Telop like making sure the business works with pure Telop and like fine-tuning models. and not feeling the need to to start out with foundation models and solve like very very general problems.

53:31 I per my personal belief is that this kind of category of company is going to be the new SAS right the way SAS companies kind of came and just like took over software over the last I don't know 10 years until I guess SAS is dead now but this is not dead um this the same thing is going to happen for sort of work in physical world and like these kinds of companies are going to do a lot of the sort of transformation of the the world's economy uh so so I think yeah it's very ripe time to get into it so kind of to talk

54:00 about that I thought I'd sort of walk through a little bit like if I wasn't doing rerun how would I do it right and this is kind of the pattern that I see just how do you get started right so pretty simple pattern uh start with a single customer problem that someone will pay you for solve it with tele operation first and just offtheshelf hardware just scrappy um kind of get started the basics you need for for learning and then uh with that in place you just yeah iterate uh more and more and sort of scale and that then

54:29 will do the rest of your time with your company but that's that's sort the fun part. Um to me um I would clearly pick this very important problem. Everybody loves paper planes and like folding is super annoying. So I would you know make a robot for automating paper plane factories. First thing you do right as I I said so sell and deploy like super fast uh ideally with you know teop and offtheshelf hardware I said and like the reason for this basically like the physical world is brutal right uh everything that you do is

54:59 going to break um all the like you will not have thought of all the different failure modes up front it is not possible to think of them all in the lab um and so it's really really important that you understand kind of the end to end like real business requirements super fast because you can't fix all the theoretical things. sort of benefit if you can solve something with teaop generally you can train a model to do it right the commerce is not always the case as as we we heard about and a lot of uh but but yeah if you

55:24 can that's great so just some examples of what we might learn doing this right so maybe uh we learned that uh you need to produce a thousand perfect planes per day to be viable as a business right um maybe it's okay to to fail as long as we can sort out uh bad planes uh so we need to be efficient like discriminator paper's cheap Um customers care most turns out about the speed to onboard like new plane designs.

55:52 Um if you add the right little paper tray, maybe you reduce failures by 50% because most of the failures were actually picking up picking up the paper from a pile. Um it maybe it turns out it takes you a human 20 hours of practice to get good enough to meet like a customer's uh demand uh sort of requirement and that has huge impacts on like how you're going to run operations.

56:15 you maybe you need to now hire all the tele operators because they need training for instance you may also learn that it's like 10 more times uh valuable if your robot can also go pick up the paper and pack the boxes uh for shipping at the end right so then you have sort of a idea of your V2 product and you'll definitely learn that your arms are going to break uh you know the cheap sort of research arms that you bought they're going to break and you need to sort of um after after some use and you need to change your

56:40 supplier and so so that's that's sort of part one um part two then setting up the basics for for learning. So hello world in the space is basically fine-tuning a um let's say pi model open model of some kind uh just for simplest possible case on a few hours of demonstration like tellup demonstration and just making sure that it somehow works a little bit right so you're you're kind of up on the treadmill and like it's really important to do this early as well because like training on the data early will change how you

57:13 collect how you run operations so so that's uh that's super key. Um when you have that in place, you need to um you need to be able to make sure that you can evaluate and sort of understand performance and then obviously like collect the data that you can train well on. So number one um yeah evaluating performance first thing you need to do is to have a replica of the customer's environment uh in your own office.

57:36 I've been in a lot of robotics companies offices among the company who has who actually ship working products. I haven't seen a single office that doesn't have a replica of customer environments. Um, you just need somewhere to test and you need to test a lot. Number two is like finding a repeatable way to evaluate success and this is where you're going to do this is the backbone of all the learning that you're going to do.

58:00 And here you can really encode things that the sort of generic model companies will not do. Like you're going to encode what is important to this business, right? And that you learn on the ground with your customers. There's a lot of sort of stickiness in that. In this case, um, you know, maybe we care a lot about that the edges on the planes are sharp, right?

58:19 We care about that it matches the design, it's symmetric, maybe the weight distribution is right, I don't know, right? super important to do this yourself manually like until you really understand it and it's kind of a stabilizing definitely automate it somehow train the model outsource it but do it manually your first and then you need to be tracking you know metadata uh of all the rollouts failure classifications that kind of thing yeah second thing is yeah collecting data that is effective to train on and uh it's

58:48 sort of topological almost but like good data is data that makes the model better um but what that means in practice this view is that uh you need to be training and evaluating and kind of debugging uh your data constantly. Like if I talk to uh researchers at like big uh robotics companies with huge budgets and and so on, they'll often tell me that like one of the most common things that they'll do when they're analyzing their they're debugging their policy or analyzing their data is they actually find out that the

59:16 right thing to do is to send a different instruction to their data collectors to collect data differently or like to stop doing some mistake. So it's doing this early. You don't want to be like collecting all your data up front and then train later like um huge mistake. Um there's there's a big lot of literature and exper expertise on like what what kind of data you want.

59:37 You want the right kind of variability, no bugs. Uh so that's a deep one. But the most important thing isn't like the specific ways of doing it. It's that you are testing and iterating really fast and sort of getting your hands on real problems. Um to do all this, yeah, you need the right sort of uh data and and formats kind of tools to to work with your data through through all this.

59:56 You need to be able to record and store and kind of inspect uh your data and obviously train on it, right? Uh so just a couple sort of smaller examples. One could just be this like your customer said that they care a lot about quickly onboarding new designs. So that means you have to have some strategy to be a little bit more sample efficient. So very commonly it's like more practical about these companies.

01:00:19 What they'll do then is they'll that means okay we're we're splitting the this task into more like composable subtasks. So then you need to design the taxonomy um figure out how you want to annotate this efficiently and repeatably and uh you're now in a situation where your annotation is too complex to be doing live. Maybe in a simpler case you could actually have the operator just speak or you know have a little um foot pedal or something like that to do annotation but now you can't do that.

01:00:44 that changes your operations. Yeah. Second one on the kind of data tooling. Uh you kind of get in like why not just use Postgress or you know whatever data infra was built prior generation to do LLMs or you know feature stores for prior ML and so on. Uh there the answer is basically that physical data so all the data that you know you're going to be working with in robotics is just very very different than web data.

01:01:04 It's multimodel it's multiate um it's episodic. It has this like weird semantics of 3D and sort of deep nested structures. And that means that if you try and put that kind of data in like normal like normal data systems like table table based databases, it's incredibly hard to query and very inefficient to store and process and so on. And this this like base thing is really at the heart of a really lot of the complexity of like working with physical data.

01:01:33 Um yeah, there's a really lot lot to say of that, but the end the end effect is that most teams if you don't like set the right storage sort of layer at the bottom and have building like a really lot of workarounds and allows a huge amount of friction, but you have these like very very basic things uh in place. Uh you're ready to hill climb, right? So you're you're deploying and learning from like real valuable uh like robotic service that that's giving like doing something worthwhile.

01:02:01 uh you know how to evaluate performance, you can improve, you know how to collect data and and kind of train on it well and you can debug across the stack, right? Super important. You don't know where the problems are and robotics like real robotics uh is kind of a death by a thousand cuts kind of industry. So you I just had to find all the problems and then after that it's just iterate and scale.

01:02:21 This is what you're doing the whole company. Super super fun. Includes you know improving intelligence. So, you know, scaling up crazy amounts of data perhaps if you need it, like iterating on algorithms and and sort of more advanced use of data, maybe adding in France's favorite with tactile or or whatnot or depth or sound. Um, but you're also like really going to have to get excellent at sales and like assembly of your robot, shipping it like really fast, having a what's your unboxing experience like operation,

01:02:48 support, um, and kind of everything everything in between. And like importantly these other areas that are not just modeling is like a lot of the source of your moat. This is the kind of stuff that the model pure model companies will not do right. Yeah. Just on that like just kind of the thing that we see even the absolute the teams that we've seen like really really succeed and there are some companies taking this approach that raised reasonably little small amounts of capital that have are making already making a lot

01:03:20 of money doing very very well and growing super fast. uh is just like yeah basically iterating super fast. So on the like frontier lab side that tends to mean that they're like investing huge amounts of compute for every researcher. Everybody knows they spend a lot on GPUs of course for model experiments but also uh like quite a big um spend on CPU for like just really turning down latency on like searching and exploring data.

01:03:45 Uh but these like robotics application startups instead they really focus on um sort of very very pragmatic and simple like flexible systems uh to to have like very very minimizing moving parts. Super important to have like fast turnaround the new data um full stack debugging uh and just like making sure that you can understand all moving parts. Um yeah so this is my pitch to all of you.

01:04:10 uh you should at least someone in here should go start a robotics application company. Uh the market or markets are enormous. Uh the base models will keep getting better. There is actually enough friction in the physical world to build real business modes. So means you can you know stick around which is great and you can do this with a relatively small amount of capital.

01:04:33 You will not need to raise a billion dollar seed and but you still need like great AI, you need great engineering to win. So that means that like all of you here and like I guess people listening uh you have a leg up and it will still be super fun. I'm still going to do a quick plug or like reiterate what Rerun does. So if you're building a company like this um definitely check out Rerun or you know talk to me.

01:04:54 As I said we have a open source SDK. It's basically meant for you to iterate super fast with robotics data. Um has all the pieces you need. It works really well with agents if you want to um you know make it work exactly like you like it and uh sort of a production catalog and sort of storage engine to make it fast and easy to use when at some point you need to start scaling.

01:05:14 You have a lot of this data for production or for training or whatever it is. All right, that's me and you can you know find us here. [applause] So I think where we see a lot of early success is it tends to be in uh things that you can tell off basically. So it will be um yeah data centers is a quite a significant category. Uh but a lot of warehouse robotics there are so many pieces that go into just moving things around in the world and a lot of them are quite repeatable and labor is fairly cheap.

01:05:55 Uh but it's also hard to manage, right? And it's there's too little labor out there. Uh we found with the companies that we work with and generally know that can build like a reliable robot, they basically are 100% supply constrained. They they have a very very easy time like filling their demand. So that would be one area. Small scale manufacturing of different kinds.

01:06:17 Uh we see a lot of action there. Uh food as well. Hey, Nico. Thank you. Um, why haven't there been a bunch of these robot application companies yet that have been, you know, at billions of revenue? >> I think it's, you know, robotics is this like death by a thousand cuts kind of thing. And I think in any area like this, it matters that you can kind of get try out an application fairly cheaply.

01:06:49 And that's actually quite new, right? It's it's there are a lot more arms on the market now, right? The base models are way better now than two years ago. Um, and so it just it's there's been this lack of the basics that you need to kind of do this fairly cheaply. So everybody's had to go out and raise and go for much more general things to start with.

01:07:10 I think like one dilemma one might face while building this kind of company is how to estimate the scale of data you would need to solve a particular problem before the model is deployable. So like how do you go about that like how do you estimate the scale of data you would need and whether it's the right problem to solve or switch the problem so that we need less data to just iterate faster.

01:07:31 Yeah, I think that's really one of the core ideas between if you can teleop first you it's not obvious that all many businesses work without full autonomy. Um even know a lot of the robo taxi businesses don't need full autonomy right it's so a lot of these companies they take they see autonomy as a scaling factor right so you tell up generally if you can tell up you will be able to learn least important parts at some point and it just becomes a question you learn that by training models and kind of trying to you know

01:08:04 plot your own scaling curves and and so on I don't know that I know anything up front uh but that's kind of the idea of the strategy you don't guess and avoid because it's equally likely that some task the task that you thought you needed to solve isn't really an important task anyway, right? >> Thanks. >> All right. Thank you, Nico. [applause] >> Uh we're from General Instinct.

01:08:30 My name is Bill and then Guam is going to present later. I come from a kind of a technical background working on VLMs at the beginning. worked at seammens on their foundation model to train to predict time series and then guine worked mostly on robotics RL what our company does is we build infrastructure for you to run physical AI models really fast so you have for LLMs you have VLM and SG lang for physical AI models like word action models and VAS you would have us general instinct yeah this is a meme from Jim fans

01:09:02 talks that vlas are dead and then we're going to have world action models from now on So, most of us in the room know what VALAs are already. I'm not going to try to explain it. Basically, you have a VLM that's trying to predict an action through an action head. What a word action model is, however, is uh you have a central diffusion transformer that's trying to predict what the future looks like and future kinematics at the same time.

01:09:26 So you have current observation from a robot's camera in the form of video streams and then you're trying to imagine future frames as a condition to try to predict future kinematics. One example is Nvidia stream zero. You have uh basically the robot trying to predict future actions and then you have the flow matching that allows the robot to act on those action chunks in future frames.

01:09:55 And Dream Zero did really well. So these are some benchmarks that you have on Dream Zero compared to some of the state-of-the-art models, some models in Pi, some models also from Nvidia. But one problem that we noticed is that although action models perform really well because you're you're still trying to use a diffusion model to try to predict frames, it's really heavy.

01:10:16 So even after all these optimizations that you can do on it, it still takes two GB 200s to run the same model. And then each one costs around 70k. So economically for robotics as an industry, this is not scalable. Yeah. Another meme, VAS are not dead because they're smaller. So since we know uh wordex model like dream zero is super slow and we will talk about how can we optimize it.

01:10:45 So this is the architecture on the left is a training pipeline and then on the right is the inference pipeline. So for the training pipeline you basically uh you treat current observation as a condition for the flow matching and you add noises to the future latent and then you will train the model to learn the future latent and then predict the future velocity field and then send it to the OD then you can drift back to the future latent with clean state.

01:11:10 That's the same thing for the inference as well. for the inference you're doing this auto reggressively for maybe 50 steps or some some of the people they do 100 steps to ensure the uh accuracy of the models and a problem for this will be uh on the left you have the video prediction on the right you have the IDM which is the infer dynamics model so basically for each chunk production to produce one chunk for 16 frames you need to run the DIT which is a diffusion transformer for 32 times because of the cfg.

01:11:46 Uh the cfg is you need to run the condition for the flow matching and also run another unconditioned flow matching. Then you can take the derivative of the gradient that you can you can do for the gradient descent for the flow matching. If we go back to the architecture like this people were talking about oh why not just do not run the diffusion models.

01:12:05 So we don't need to predict the future frames that works and there's a research paper called image run. Basically, they are not predicting the future video chunks. Instead, they're predicting the future end state, which is the uh future end state of the uh single frame. For instance, I'm predicting the future video for maybe 16 frames. Rather than predicting the whole video, we can just predict a t plus n, which is the end state of the frame.

01:12:31 And there's another research paper called fast one. Fast one is more extreme in some sense. uh they think all the word representation they are already learned in the hidden state of the DIT. So you don't even need a decoder to decode all the videos. You can just use the hidden state as a condition to train your action head by doing this. You don't even need a decoder in the training and also in the inference pipeline.

01:12:56 If we if we take an analogy of those two different word action models for the generative word action model which is you need to decode and then p predict the future frames you're pretty much like a VR of Google maps but for the latent word action model it's like you look at the navigation of your Google maps and think about what the model the policy is heading to and what kind of action you're going to produce in the future.

01:13:22 In conclusion, all the problem began all the problem comes down to the question about how to keep the rich word representation. Some of the people like Liquin they think about uh because they were doing Japa they think about maybe we can have two different encoders and then one encoder is encoded the current observation and then the other one is decoding the a future observation and by learning by doing loss on the current observation lat future observation latent then you can teach the model to learn how to predict

01:13:56 the future and that's one way of doing this and they're doing this use MSU loss and the other people they're using they treat the future as a distribution of of possibilities it might be you take this possibility of this action you might take uh other other possibility of taking other action so you treat them as a distribution and then you estimate those kind of action distribution you think flow matching we talk about those different optimization angles we might have and for us since we are doing the infra thing so We

01:14:29 did all those optimization on our infra. Uh we did dissolation on the VA part which is a VA encoder decoder. We also did dissolation on the DIT part. So the DIT became smaller. We also divided the because previously they were using the same transformer same DIT. We divide them to two different DITS. Rather than decode all the future frames we can just use the cross attention from the video transformer to the action transformers.

01:14:59 So the trans action transformer learns the hidden state which which is a word representation from the video transformer and you don't have to decode future frames anymore. We also did distillation on the uh auto reggressive flow matching sampling. Uh previously it might take 50 steps or like 100 steps to do the flow matching decoding but we made it to down to one or two steps which is immediately 50 times of speed up and we also did some changes on the modality side because we know future representation or the word

01:15:36 representation uh you can learn through pixel level or you can learn through through latent level. Is it possible we can find a more suitable modality to represent or to retain the word representation? So one way of doing this it might be mask and then the other way might be flow. We tested both both of them and you can see the heat map is the visualization of the model and it guides the word model about which action you're going to take in the future 0.5 seconds by using our infra the word action model can runs 500 it

01:16:18 should be 500 milliseconds per chunk which is 16 actions on JSON sore and we also wrote a full blog about how we did this on our website. This is the QR code if you will to learn more about it. And also this is the LinkedIn of the founders of us. That's it. Thanks guys. [applause] >> Uh going back to that slide that you had with the training and inference of the world action models.

01:16:50 >> In the training stage, >> you're using flow matching and teacher forcing. Um, that makes sense. But in inference, you're going to start from a fully noisy space and then you're going to come to the future time step. How do you ensure that at inference time the model actually collapses to the right thing and it just doesn't degrade to noise? >> Uh, that's a very good question.

01:17:11 So for the training part, you begin with the clean future latent and then you add noises to the clean future latence gradually. So eventually the future clean latency will become pure noise at the end for the training part. So by doing this the inference learns how to reverse it back. So when you give into a pure noise the inference learns how to do this gradually and then reverse back to the future clean latent.

01:17:40 >> So that's basically how it works. >> So like curriculum learning you're starting with like very little noise initially and as you train for longer longer you have more noise >> and that's why it's super slow because people doing this maybe for 50 steps. So we think it's too slow. So we just found a way to distill it to like two steps or like three steps.

01:17:55 So it'll be way quicker and without performance drops. >> Yeah, that's super interesting. It'll be all. So when it comes to the world models, uh is the performance improvement due to the action heads looking at more details when we are forcing them to predict the whole frame? And second one is when there are multiple agents that are involved in the scene uh does the model actually develop some kind of a theory of mind and predict other agents actions in order to be able to predict the world.

01:18:24 >> Oh that's a very good question. So I want to go back to the architecture of VA and word action models. So uh I got lots of question about what is difference between VA and the word action model. Why does word model need to predict the future rather than just predict the action itself? And a way of how to answer this is for VA especially there are just based on the current observation to predict the current action.

01:18:53 So they don't have the explicit learning of the future kinematics. The only reason we need to introduce videos especially for future videos for word action model is we want to teach the model to learn the future kinematics. For example, I'm holding a a bottle of water and then I drop the bottle of water. So from pixel level, if we give the future videos to the model, the model learns how the kinematics will change the pixel level.

01:19:22 And then we know this is kind of a superficition of teaching the model to learn the future dynamics and we believe the future dynamics help with the action generation as well because it's physics. So by doing this we teach a model how to learn the correlation between the physics from from the pixel level to the action you produced. to test that hypothesis.

01:19:40 Um, could you train a transformer stack body to do next frame prediction and then more and more and then rip off the head and just do the action. Why would it do the same thing? >> I think like we divide them to two different transformers, >> two different stages. >> Yeah. Um, of course there's a more efficient way to do so and it's called mixture of transformers.

01:20:08 And if you check the image one, they actually do the same thing. They did the cross attention from the image editing backbone to the action expert because uh just like uh image RAM and also fast one they realize the word representation you don't have to explicitly decode the frames. You can just keep them in the hidden state and then do a cross attention from the last layer of the backbone to the action head.

01:20:34 So short answer for this is of course you can do so. And then we realized this is a more efficient way to of to maintain the word representation meanwhile produce the best action based on the word representation. >> And then is it is there no test time planning that's done with wham where you will invoke let's say 10 samples and I'll get 10 different end states and 10 different actions.

01:20:59 Um, and then I'll pick the best end state of the ones that were sampled and then emit that action. Is that not done? >> Yeah, I think it's pretty much how flow matching works, right? Because for flow matching, you're basically train you trade action as a possibility of distributions and then you always sample the a best trajectory and we use a teacher forcing to teach the model to sample the distribution of the actions.

01:21:28 So I think for flow matching they're already doing the same thing. >> And then last question for me like why choose the business model of being a a VLM equivalent for whams versus just actually you know do like Nico says become a robotics application company and actually go end to end. >> Uh I think Bill can't answer this question. [laughter] Yeah, I think the reason we went with this route is because we really believe in having this understanding of the world for your models.

01:22:02 But I think right now everyone's focused on maintaining the research so that it can be as generalizable as possible. But eventually every model like that needs to go on the edge and needs to go in real time. So I think right now not a lot of people are focused on building the infrastructure that allows those models to perform really well on your robots first.

01:22:20 We want to be the first company to do that. >> Hi, I have two questions. Um the first one is more of a clarifying question of like um so for a wham like do you do is the prediction auto reggressive in like previous frames or is it like one previous like t minus one and then to t without any other like t minus 2 minus 3 and so on. >> I think what you talking about is the chunk size because you can change the parameter as well.

01:22:44 So for the training you can do 16 frames which people they all do 16 frames which is you predict in the future 16 frames and people they also do uh 32 or like even higher frames but if you increase the chunk size which is more frames you produce then it will be harder for the model to learn the future states because it's longer. >> Okay. Uh and my second question is are you uh familiar with any word that uses like um encodings for the differences between frames?

01:23:13 So like temporal difference encoding I think like a recent work by Yan Lun um as well as like a tech blog by I think induction labs where they like train a image imagination model um that predicts like lat encodings for um the differences between uh adjacent frames in a video. >> Yes. And uh if you check here we mentioned about asymmetrical dnoising because we found a way that you can actually measure the energy of the ki cache.

01:23:40 So rather than you can produce all the future frames, why not just produce those maintain the highest uh details of the the action you are doing right now and uh that's basically what we added to the infra as well. So if you we found a way to measure the k to measure the different energy of the kf cache and based on the energy then we decide if the model going to predict different resolution of the frames or just purely doing this on the latence base. >> Thank you. [applause]

💡 Answer

Robotics is not solved because physical modeling, sensory-motor feedback, sim-to-real transfer, long-horizon memory, and hardware drift remain difficult; it could advance soon through better memory, selective reasoning, simulation-trained dexterity, and focused application companies.

🧠 AI Summary

Robotics remains unsolved because systems still struggle with physical-world modeling, the sim-to-real gap, deformable objects, limited sensory feedback, and embodiment drift. Long-horizon tasks require memory, while useful embodied reasoning must be selective, action-predictive, and embodiment-specific. Simulation-trained reinforcement learning can provide broad dexterity and zero-shot tool use, but pose tracking and real-world complexity remain bottlenecks. The most practical path to commercial robotics is to solve narrow customer problems end to end with teleoperation and off-the-shelf hardware first, then improve autonomy through iterative data collection, evaluation, and deployment.

🔑 Key Points

  • The sim-to-real gap persists because simulated world models do not reliably respect real-world physics.
  • Robots lack the distributed tactile sensing humans use to detect force, friction, moisture, temperature, and vibration.
  • Embodiment drift makes teleoperation data stale as actuators corrode, collect dust, or lose battery performance.
  • Long-horizon robotics tasks need short-term visual memory and compressed long-term memory to track progress, time, and mistakes.
  • Self-supervised bootstrapping can discover reasoning traces that are concise, non-trivial, and predictive of actions.
  • Simulation-trained goal-conditioned reinforcement learning can generalize to unseen tools and tasks without task-specific retraining.
  • Robotics application companies should own the full customer workflow instead of beginning with general-purpose foundation models.
  • Physical robotics data is multimodal, multirate, episodic, and deeply nested, making conventional table-based data systems inefficient.

✅ Actionable items

  • Decompose robot policies into a high-level policy and a low-level action policy, using compressed long-term memory for the former and dense short-term visual memory for the latter.
  • Train long-term memory representations with supervised annotations, while exploring reinforcement learning to discover which information should be retained.
  • Use selective reasoning traces scored for concision, non-triviality, and action predictiveness, then retrain the embodied policy on the filtered data.
  • Train dextrous manipulation policies in massively parallel simulation with random objects and random goal poses.
  • Use a human video only to specify a desired object trajectory, extracting goal poses with Foundation Pose and SAM while keeping the robot policy frozen.
  • Add random perturbation forces during simulation training so policies learn to recover from dropped objects.
  • Start a robotics application by identifying one customer problem, validating it with teleoperation and off-the-shelf hardware, and iterating before increasing autonomy.
  • Deploy a replica of the customer environment for repeated testing and manually define success criteria before automating evaluation.
  • Train and evaluate continuously rather than collecting all data first; update data-collection instructions when evaluation reveals problems.
  • Track rollout metadata and failure classifications, and split complex tasks into composable subtasks when faster adaptation is needed.
  • Distill diffusion and flow-matching models from 50 or 100 sampling steps down to one or two steps to reduce inference latency.

💡 Business ideas

Automate paper-plane factories with a robot54:21

A hypothetical narrow robotics application that folds paper planes and could later pick up paper and pack boxes for shipping.

For
Paper-plane factories
Solves
Automating paper folding and related handling and packing operations.
Validate by
Start with teleoperation and off-the-shelf hardware, then measure required production volume, failure handling, onboarding speed, operator training, and hardware reliability.
  • Producing 1,000 perfect planes per day
  • Sorting out bad planes
  • Picking up paper and packing boxes

🏗️ Business models

Robotics application company52:36

An end-to-end physical-world business that solves a narrow customer problem using robotics operations, teleoperation, off-the-shelf hardware, and progressively improving autonomy.

  1. Choose a single customer problem that someone will pay to solve.
  2. Deploy a teleoperated solution using off-the-shelf hardware.
  3. Validate the full business workflow in the customer's environment.
  4. Collect and evaluate data while operating the service.
  5. Fine-tune models and gradually scale autonomy.
  6. Expand into adjacent operational tasks.
  • Data center management
  • Construction
  • Warehouses
  • Manufacturing
  • Food
  • Paper-plane factory automation
Physical AI infrastructure51:21

Infrastructure software for collecting, storing, querying, visualizing, training on, and serving physical-world robotics data and models.

  1. Record and store multimodal episodic robotics data.
  2. Inspect and query data.
  3. Evaluate rollouts and classify failures.
  4. Use the data for model training.
  5. Scale storage and production workflows.
  • Rerun open-source SDK
  • Rerun Hub
Physical AI model infrastructure01:08:30

Infrastructure that optimizes and serves vision-action and world action models in real time on robotics hardware.

  1. Distill model components.
  2. Reduce or remove future-frame decoding.
  3. Separate video and action transformers where useful.
  4. Reduce flow-matching sampling steps.
  5. Deploy models for edge and real-time robot inference.
  • General Instinct infrastructure

📣 Marketing

Sales

  • Demonstrate that the system solves a valuable operational problem before pursuing broad autonomy.

Branding

  • Build defensibility through operational expertise, evaluation systems, deployment, support, and full-stack ownership rather than modeling alone.

Distribution

  • Sell and deploy quickly using teleoperation and off-the-shelf hardware.
  • Provide robotics services in data centers, warehouses, small-scale manufacturing, and food operations.

Customer acquisition

  • Begin with a single customer problem that has clear willingness to pay.
  • Deploy directly into customer environments and learn the end-to-end business requirements.

🧭 Frameworks

Multiscale embodied memory11:59
  1. Separate the robot policy into high-level and low-level components.
  2. Feed dense short-term visual memory to the low-level policy.
  3. Feed compressed language memory of the previous minutes to the high-level policy.
  4. Use recurrent memory updates to track task progress and past mistakes.
R&B encore24:24
  1. Generate candidate embodiment-specific reasoning traces.
  2. Score traces for concision, non-triviality, and action predictiveness.
  3. Resample the highest-quality traces into a synthetic reasoning dataset.
  4. Retrain the embodied reasoning vision-action model.
SimToolReal36:27
  1. Generate primitive objects and random goal poses in simulation.
  2. Train a goal-conditioned policy with massively parallel reinforcement learning.
  3. Extract goal-pose trajectories from a human video.
  4. Track those poses with the frozen policy at 60 Hz.
  5. Deploy the resulting behavior on novel real-world tools.
Play to Perfect41:30
  1. Learn a shared dextrous prior through task-agnostic play in free space.
  2. Fine-tune the prior on a sparse-reward contact-rich assembly task.
  3. Deploy the policy zero shot in the real world.

🧰 Tools & AI usage

  • MAM — Multiscale embodied memory system for long-horizon robot policies.08:05
  • Foundation Pose — Extract object pose sequences from human video demonstrations.37:46
  • SAM — Assist with extracting goal poses from human video demonstrations.37:46
  • Rerun SDK — Log, visualize, query, load, transform, and analyze physical robotics data.51:21
  • Rerun Hub — Provide a data catalog and large-scale cloud backend for physical AI data.52:10
  • LSTM — Process the current robot state, object pose, grasp bounding box, and goal pose to produce joint position targets.38:18

AI is used for

  • Compressing robot memory — Represent recent visual context densely for low-level dexterity and represent longer history as compressed language for high-level planning.11:59
  • Generating and validating embodied reasoning — Create synthetic, action-aligned, embodiment-specific reasoning annotations for training vision-action models.24:24
  • Goal-conditioned dextrous manipulation — Control a hand and arm to move objects through desired pose sequences across unseen tools and tasks.36:27
  • World action modeling — Predict future visual states and future kinematics to improve action generation through learned physical dynamics.01:09:01

📊 Numbers mentioned

Costs

  • World action models were described as requiring two GB 200s to run the same model.
  • Simulation can generate experience about 1,000 times faster than real time.
  • Simulation can collect decades of interaction data in a few days.

Growth

  • SimToolReal runs at 60 Hz.
  • SimToolReal controls a 22 degree of freedom hand and a seven degree of freedom arm.
  • SimToolReal was evaluated across 12 unseen tools and target behaviors.
  • Roughly 60% of SimToolReal failures were attributed to pose tracking.
  • World action model flow-matching sampling was reduced from 50 or 100 steps to one or two steps.
  • The optimized world action model runs at 500 milliseconds per 16-action chunk on an unspecified processor.

Pricing

  • Each Nvidia Dream Zero model instance costs around 70k.

💬 Quotes

This time it's different.

Captures the recurring optimism surrounding claims that robotics will be solved soon.00:54

Selective reasoning is way more important than exhaustive reasoning.

Summarizes the central finding of the embodied reasoning presentation.28:22

The physical world is brutal.

Explains why rapid real-world deployment is necessary for discovering robotics failure modes.54:57

Robotics is this like death by a thousand cuts kind of thing.

Summarizes the many small hardware, data, software, and operational problems that make robotics difficult.01:06:38

👤 People & companies

Marcel

Stanford PhD student who presented multiscale embodied memory work completed during an internship at Physical Intelligence.

07:59
Milan Gennai

PhD student associated with Marco Pavone and Clark Barrett, with research experience at AWS and Whimo; presented embodied reasoning work.

07:05
Tyler Lum

Presenter of SimToolReal and Play to Perfect, focused on simulation-trained dextrous manipulation.

43:42
Nico

Co-founder and CEO of Rerun, discussing robotics application companies and physical-data infrastructure.

51:21
Bill

General Instinct team member who presented infrastructure for fast physical AI model execution.

01:08:30
Guan Ming

General Instinct team member who works primarily on robotics reinforcement learning.

01:08:30
Chelsea Finn

Researcher whose lab hosted Marcel's PhD work and whose lab was referenced in the event introduction.

06:56
Marco Pavone

Researcher associated with Milan Gennai's PhD work.

07:05
Clark Barrett

Researcher associated with Milan Gennai's PhD work.

07:05
Carl

Co-first author of the multiscale embodied memory work.

16:21
Y Combinator

Organizer of YC Paper Club and YC Robotics Club.

00:08
Physical Intelligence

Robotics company where Marcel completed an internship and developed the multiscale embodied memory system.

08:05
Rerun

Company building an open-source SDK and data infrastructure for physical AI, including a data catalog and cloud backend.

51:21
General Instinct

Company building infrastructure for fast execution of physical AI models such as world action models and vision-action models.

01:08:30
Nvidia

Company associated with Alpomeo and Dream Zero, which were referenced in discussions of vision-action and world action models.

20:50
DeepMind

Organization referenced as part of a speaker's previous research experience.

07:47
AWS

Organization where Milan Gennai conducted research.

20:45
Whimo

Organization where Milan Gennai was described as currently working.

07:05