Why this is the best time to build hardware: AI tools, energy, memory and optics are creating a hardware renaissance — but the hardest constraints are now manufacturing, memory, power and energy, not design.
Rebuild the entire virtualization and management stack — security profiles, performance management, policies, migration — but designed so the primary 'user' of the infrastructure is an AI agent rather than a human. Every fundamental element of the VM abstraction (security, performance, abstraction, migration — 'vMotion for agents') gets a new embodiment, while humans set the policies, constitutions and guard rails and view dashboards.
Searchable transcript of Former Intel CEO: Why This is the Best Time to Build Hardware — a16z (53:15). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by a16z. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 In a AI digital age, energy capacity equals economic capacity. Why build a new data center and buy the million GPUs if I can't power them? You're going to see more and more defaults happening on many of those data center projects cuz the energy won't be there. >> Whenever you have the technology to make something easy, that means the bottleneck moves somewhere else.
00:21 >> Nothing's a chip anymore. It's a rack. It took me 3 months to design it, but it's 9 months until I can actually start to use it. Exactly. How many major new memories have we had over the last 30 years? >> Zero. >> Zero. >> Memory innovation for the first time in 30 years is nigh upon us. I declared the death of copper about 25 years ago. Eventually, I'll be right for all of us hardware guys.
00:40 This is like renaissance in front of us. AI inference accelerator chips. And I'm sure I don't even know them all. Why are you guys funding so many of those? >> You funded your fair share, too. >> Historically, there have not been a 100 competing processor vendors in any industry ever. This is a temporary thing. It'll converge back to a view. >> I see it as a >> I'm here with uh our esteemed guest and dear friend and former boss, Mr.
01:05 Pat Gellinger. Welcome Pat. Pat is currently the general partner at uh at Playground Global, but uh as uh you're very very well known in the industry for leading Intel, for being the CTO of Intel, of course, leading VMware and many other things. >> So, welcome. >> Hey, thank you, Rayu. Great to be with you and Guido, right? And to me, this feels just a little bit like old home, right?
01:33 You know, it's like, you know, you know, it was like uh we superimposed a year of change, right? But uh you both look the same. I feel uh energetic. So let's dive in. >> There has never been a day when you stopped being energetic. So there no news there. Uh but yeah. No, let's start actually from uh from your Intel career, right? I mean recently Andre Horovitz created this Horowitz and recreent academy >> which was to find talented people between 16 and 22 and then put them into a a modern educational setting so they'd
02:10 be well prepared for either doing things on their own or joining the companies of today. You had a similar experience like you went to a regular trade school and then jumped into Intel. >> Yeah. Yeah. When we were what 18 or 19 is >> Yeah. Uh, you know, it was really a sort of magical period. You know, I'm 16 years old. I accidentally win a scholarship.
02:30 I go to tech school, right? Skipped my last year and a half of high school and Intel comes recruiting when I'm 18 years old. And, uh, you know, I never been on an airplane. Uh, had already fallen in love with computers at tech school. And, uh, the interviewer, right, Ron Smith was his name. He writes on his page. And I was number 12 that he interviewed that day.
02:49 And if you've interviewed 12 people in a row, you can't tell male from female, you know, by the end of it, right? You know, I'm number 12. He says, "Smart, aggressive, arrogant. He'll fit right in." So, I got invited to go, you know, come to uh Intel and, you know, just became uh, you know, really uh glorious. uh you know starting as a technician moved into the design team at the end of the 286 you know engineer number four on the 386 you know uh architect and design manager for the 486 all of that while doing my
03:23 master my bachelor's my masters and PhD work so it was like the the the best career that you could possibly have you're learning by you know day and you're putting it to practice at night and >> that's right that's right you're creating new completely new era of chips for I bet most of what you learned and what you built there was a big gap because you're breaking new ground.
03:47 >> Yeah. In a lot of ways and uh you know really was a pretty incredible time in the industry and I remember you know one of the professor classes uh you know he uh postulated a new carry look ahead adder design right in one of my classes at Stanford. So I said well I'm working on the 386 carry look ahead adder so let me go try it. So I come back two days later to the professor and I said it doesn't work right and I started argument with my you know professor at Stanford and it was like you know it was like what you
04:15 know the next two class periods were just me and the professor arguing over why his design didn't work. >> He had a book that was being published what that design is. So he was he was very annoying. It was very time sensitive for him >> and you know we started builtin self test right you know and I remember Ed McCcluskey at Stanford was my professor and he was sort of the father of that and we you know had wonderful arguments around why most of his ideas were bad they weren't practical for real chips but I was putting
04:42 you know builtin self test to work on the 36 you know John Hennessy uh you know was my thesis advisor and obviously he did okay I say you know I helped his career he became president and he's done okay since then Yeah. Yeah. Yeah. >> Later on we'll talk about the CIS versus risk debate, but I want to talk about the 486. Sorry. Go ahead. >> I had the same experience when I was teaching at Stanford.
05:03 You know, you you come with something that you think is the truth and then the student comes that works in industry is like, well, actually not quite. So, you know, the next lecture you come back with like, okay, uh, I looked this up. You're right. You know, this is this is outdated. >> Yeah. And I tell you, you know, that's part of the set of students.
05:20 Yeah. that interaction and I think that's part of what makes Silicon Valley so unique. Yeah. >> Right. You know is you have you know these so many companies and you know entrepreneurship it's like in the water here right you know starting things up and new ideas and challenging the professors and you can't be a professor at a place like Stanford and not expect to get challenged and enjoy it.
05:41 >> There's a good foot traffic in and out of that ivory tower. >> Yeah I think it's good. >> Yeah. Now 486 was your big >> Yeah. >> accomplishment if you will. Obviously lots of people worked on it. Um and you guys broke some new ground in chip design there, right? >> Mhm. >> And um increased what functions the chip was supposed to do and so on and so forth.
06:05 >> Yeah. It was a pretty magic period uh as well because it was before what you would think of as the EDA industry. >> Oh wow. Yes, that's right. and uh you know the the end of the a little bit of the 386 but the 486 was really the first chip to implement what what you would think of as modern EDA techniques. Yeah, we did a highle design, right? Uh description and RTL description, but there was no vera log.
06:33 >> So, we invented HDL, right? The Intel hardware description language. So, you know, I wrote my own language. Well, you had to build a compiler for that language, right? You know, so we created the compiler for the language. There was no automatic place in route, you know. So, we had to invent that. And we worked with Alberto Sanji Yavanji Vincentelli at Berkeley and some of his students for the first you know placement the first routing the first automated timing you know management and >> so they created EDA to some
07:00 degree. to a great degree that uh you know and that was really one of the hallmarks that not only was the 486 you know this compatible you know pipeline microprocessor but we ushered in you know many of the foundations of what became the modern EDA uh industry and you know it was really you know pretty magical period of the industry. >> Yeah. Yeah. So if you look at today >> Yeah.
07:25 Will people look back as Blackwell and saying you know this was of the last generation that was was done before AI took over the design process for for micros are we in a similar transition right now >> I think there are certainly aspects of that and uh you know I think you know when you look at you know things like jalapeno today right you sort of say you know that you know that's that's sort of a first principles use of AI in the chip design process right where you throw out a lot of those things now you know un You
07:55 know there are pieces of design that are still really hard right uh in that sense and a lot of the analog right aspect certis is probably the best example of that you know are still not aiable right you just need lots of silicon data right to get those so you either get so conservative in your analog design you know that you're able to I'll say AI it uh or right you got to do those the hard way uh you know for but to now many of the logic functions can really be done with uh AI tools and techniques in pretty incredible
08:27 ways, right? Your transistor budgets are big enough, the design tools are getting smart enough, you know, a few experts guiding the tools and how to apply it. And I really think it will be somewhat like the 46 that way, right? Where you sort of look back and say, "Yep, that was the beginning of a new era of chip design." >> Whenever you have the technology to make something easy, that means the bottleneck moves somewhere else.
08:52 >> Yeah. Can we guess at this point where the bottleneck will be in the future? >> There's some things which it seems like AI is incredibly good at, right? Like creating the software layers, you know, writing kernels, um, you know, a lot of the >> I think a lot of the sort of being able to specify the the the objectives at a higher level which then gets sort of translated into very log right and so what is the new frontier?
09:14 What is the thing that that will be difficult going forward? Well, you know, when you look at today, you know, so let's say, you know, Guido and Pat, we're off to do a great new chip together. We're building, >> right? You know, right? You know, and it's the one that's going to leap ahead because we have understanding of 8646, >> you know, we have understanding of the AI workloads.
09:34 We have understanding then of how to compose that into the right, multiply, accumulate, you know, register structures, >> the perfect visibility model workloads of tomorrow as well, >> you know, and all of that kind of stuff. And you know let's say you know we turn our AI you know tools loose we are great within three months we have an awesome design right you know you know for that we've been a little bit conservative on all the analog components of it >> you know now we still have like a couple of really major
10:02 bottlenecks today you know one is right we still have nine months of silicon processing time >> right you know someone like that right so I can design the thing in three months but I can't actually get it into real silicon at scale for 9 months. Okay, that sucks, right? So, we have to really see that bottleneck improve if we're going to have this innovation because we know that >> about the validation verification stages or even beyond that.
10:26 Well, you know, first I can't get anything out of fab. Yeah. You know, in less than three months and that's if I have, you know, like supercharged design flows, right? So, I said, you know, you know, how do I compress that, right? Uh, you know, for it and then I got to get into advanced packages. 3D packages are getting complex. So, that's like another three months, right?
10:44 You know, and then I have to put it into a rack scale solution because nothing's a chip anymore. It's a rack, you know. So we're 9 months, you know, it took me 3 months to design it, but it's nine months until I can actually start to use it, you know. So how do we start to compress, you know, those aspects of design? And I think we need new forms of lithography, you know, to enable that.
11:03 We need, you know, design flows that don't require, you know, $50 million of mass cost until I can get things into prototyping. So you know, to me, how do you compress that to a month or two, right? Because I can now do the design in three months. How can I have that done in a month or two? Because if it takes me a year and a half until I actually get scale and software on it, >> okay, my understanding of the AI workloads is no longer applicable to the chip that I designed.
11:31 >> Right. You know, so you know, all of those >> you see that play out live right now with AI. Yeah. >> Right. You said, you know, it's almost, you know, we can look at it historically. It was like, you know, Graph Core wasn't a bad design, but the world moved on, >> right? and you know uh in that. So all of those aspects somewhat obviate the fact that AI has made the design piece easy.
11:52 All of those manufacturing scale issues to me become a huge bottleneck. Right. Another bottleneck of course is memory. >> Yep. >> Right. You know memory bandwidth >> we noticed. Yeah. >> Right. And you know for that as I've described HBM is a hideous memory. It's just the best one that we >> ask you about that. Yes. >> You know it's a terrible memory.
12:09 You know I get terrible bit density. I'm limited to shoreline bandwidth the power right uh issues thermal issues DRAM doesn't like to be hot >> concentrating everything in one spot >> you know it's terrible to give >> that AI is a memory compute workload yes >> right so right you know for that so I you know I need to you know compress that development uh and embodiment into real scale you know I need to fix memory >> right uh associated with So to me that's another critical bottleneck for us to attack.
12:44 Uh and uh power right you know power in you know thermal uh and most of the power simulation aspects today are bad >> right you know you don't have good 3D modeling tools for power power dissipation hotspots uh etc. you know my voltage in right you know I'm essentially guard banding almost my full power envelope correct >> right so essentially I'm losing you know 40% of my power and guard banding so I got to get much better at power modeling you know and then my power dissipation techniques are bad right so you know >>
13:18 delivery model has to change too >> yeah so all of these things you know to me those are the you know the next big bottlenecks right you know how do I you know improve each one of those and now right you know if I've done that Okay, now we're on to an exciting era for what's next. So, you know, thank you for those great AI tools, but you know, I got big challenges even when those are fully deployed.
13:39 >> And so, we've seen sort of a rebirth in the number of chips and chip architectures and so on and so forth. >> Yeah. >> Thanks to these workloads. >> Yeah. It's like there's a 100 AI inference accelerator chips and I'm sure I don't even know them all. >> Yeah, exactly. Yeah. So, so >> why are you guys funding so many of those anyway? >> You funded your fair share too.
14:06 >> Yeah. So, um, so are we, you know, they're all to their credit making different design assumptions either in the components that you talked about or >> in the >> in the part of the para curve that they're targeting. So do you think there's room for all of these or are we in a world where partly because of what Guido talked about AI tooling um we now have the ability to build more different kind of chips.
14:33 >> Yeah. >> And the time is enormous whichever way you want to even more provocatively historically there have not been a 100 competing processor vendors in any industry ever right. So >> is this a more heterogeneous future? Will we see is this a temporary thing? It'll converge back to a view. >> Yeah. And you know, I see it as a more temporary thing and it will converge is is what I expect.
14:54 And there's probably three different reasons that I see that to be the case. You know, one of them is, you know, this emerging heterogeneity that you see in the compute uh environment where okay, now I have specialized prefill versus decode. Well, you know, and now people are saying, well, we really need a specialized tier for midfill, not just prefill, right?
15:17 You know, so now I have early prefill and now I have midfill, which is, you know, differential. >> Maybe speculate and verify if we listen to OpenAI here, right? >> Yeah. You know, and you know, all of a sudden my compute fleet becomes more and more granular across the workload. And I think anytime that you've seen that in history is not sustainable.
15:36 And then all of a sudden right when you know as you look at that and now people are saying well as I go to reasoning models I want something that looks more like a CPU again. So you know all of a sudden I didn't want that level you know that didn't become the predominant portion of my compute capacity. I needed more of it over here, right? and model and model workflows and you know so on you know so I really don't like I'll say at scale specialization right because I think the workloads are going to continue to
16:08 moderate you know and migrate so significantly first you know second is I think today's models are about to go through some evolutionary breakthroughs as well right you know I think the LLM is sort of reaching the limits and now as people look to how do I really model 3D right you know where you a flattish LLM doesn't work well when you go to molecules and chemicals and you know imaging you know so I think we'll see limits of that that'll cause different shifts in the algorithmic domain I also think that uh some of the
16:38 most interesting workloads become where you start you know I'll say bringing HPC like things back into AI like things where all of a sudden you know 64-bit provision uh precision counts again right so I see >> what would be an example of that >> well you know imagine that my AI model is now inducing the five most interesting domains for a chemical analysis, right?
17:02 Okay, those chemical algorithms are going to run with high precision, right? You know, when you know which is going to be the dominant piece of the workload, right? Getting to the five algorithms I need to run or running the five algorithms, right, on different chemical or biological systems. So you know I do think that uh you know or as we optimize our AI systems they will lead us back to some of the things that have been more traditional high performance computing.
17:27 So I see that aspect of workload uh as well. So you know all of that said is you know this extreme specialization to me in the compute architect I I sort of don't like it because I don't think it's going to be what the workloads will look like 2 3 four years from now. So that's one reason I don't see all of these AI chips as they get more and more specialized for portions of the compute workload to necessarily be right you know second is right there's like a hundred of them now >> right you know there's multiple
17:56 optical ones there's probabilistic ones and so on and you know they're not all going to win and they're not all going to win because at the end of the day you have to get scale >> on these things and scale requires you know you have to win you have to get capital you have to get workloads onto it so I think it sort of defies logic that you're going to have a hundred of these things.
18:15 So I see them you know coming back uh as a results of you know simple you know capital market share etc. you know which are the winning teams you know winning designs winning architecture so I see the workload driving that you know I see the natural you know industry consolidations you know that will occur and then I also expect that the winners will pick some winners >> right you know where you know an open AI an Nvidia anthropic will say I like this one because it isn't just hardware it's how do the hardware and
18:49 software code evolve Yeah, >> right. You know, for it and I do expect that there's going to be h, you know, here's this 10 I could pick for this. I like that one. And it does require investment in the software and the workload evolution to take advantage of those platforms. So for those three reasons, yeah, we're going to see, you know, a narrowing of the field uh in that sense >> and the deployment complexity as well.
19:14 Sorry, go ahead. >> I mean, makes total sense and I think I agree in most point. Let me still try to stir a little bit controversy here just to to stirring controversy to you guys have changed it right >> so so obviously 100 companies won't access consolidate right and and completely agree with your argument that at the end of the day the large consumers will pick winners to some some degree right today I think >> to be a first tier chip startup to some degree whether you you're that or not is defined by how much
19:43 fraction you have with with the with the really big Nice. That said, um, and there's of also an argument to say like, look, in any kind of transition period, you probably have a bit more of an explosion of different approaches over time that that'll shake out. >> Now, all that said, I think there's an argument you can make that basically says, look, >> historically, we standardized on, you know, one or two architectures because, you know, different instruction sets were hard.
20:06 we had to build multiple compilers you know like the this basically um this the the software complexity dictated that you can only have a certain small number of hardware platforms >> that part seems to be changing with AI right today you know writing an optimized kernel for I mean I think was a great example where they basically said look humans can't really program this thing anymore but an agent swarm overnight can do it that's all we need right so where basically you can say like if you can build a chip that has a
20:32 certain capability it's now much much easier to build the software layers on top that fully leverages that could a result of that be that we have a little bit more hetrogenity just because the you know it's it's easier to program that hetrogenity going forward >> and and I I do think it changes some of you know you you know the you know where the bottleneck lands right changes in that regard you know the first bottleneck is the one I'd point to that we've already talked about still takes a year and a half to build
20:56 these things >> right you know and then the second bottleneck is not only does it take a year and a half to get these things into racks the scales and so on like that now to make them useful full. Okay. Well, could I borrow 50 billion, right? I mean, you're not talking about small capital expenditures. You're talking about very large capital expenditures and data centers, you know, power, you know, commitments and so on like that.
21:19 Well, I got to build this thing at scale, right, associated with it. So, and that's going to defy a lot of heterogeneity uh as a result. Now, you know, clearly big players and we saw this with what, you know, Nvidia did with Grock, okay? you know, we're gonna sort of own some of that heterogeneity, right? And you know, you know, two years from now, you're not going to tell what Grock is inside of Nvidia, right?
21:38 It's just going to be like another, right? You know, someone like that. So, you'll see the architectures, you know, sort of hide the heterogeneity, right? Uh inside of it, which to me is just fine, right? You know I mean that's what layering that's what abstraction has always done associated with it you know but I think those forces are powerful ones to I'll say expose a lot of the heterogeneity no right I think it will get abstracted underneath you know some of these but at the end of the day it's still hard to build
22:06 these things at scale >> earlier you talked about HPM being hideous memory >> by far the best one we have today. >> Exactly. It's like the thing about democracy. >> Um, we can't recall when there was the last memory innovation. It might have been obtained. I think you probably had something to do with it. >> Yeah, that one died. So, >> that one died.
22:27 >> We killed it. >> Yeah, we killed it. We killed it. >> So, do you think memory innovation is around the corner or what needs to happen for obviously we understand the scale problem because it's even more >> Yeah. >> larger there. >> Yeah. But if you keep the scale aside even from a technology point of view. >> Yeah. And let's you know uh you know look at that history.
22:48 I've probably personally been associated with at least five different new memory architectures. Optane just being one of them right you know that did not see the light of day right uh as well. Uh I'm probably familiar with close to a hundred that have happened in the industry over the last 30 years. And exactly how many major new memories have you had over the last 30 years?
23:13 >> Zero. >> Zero. >> Right. You know, which you know, okay. DRAM, SR, RAM, flash. Okay. What else is there? D RAM, SRAM, flash. Yeah. You know, so I I do think >> flash stacking console a little bit. >> Barely. Yeah. Right. You know, but anyway. Right. >> So, uh, you know, in that it it you know, it has been right you know, a a disappointing >> that we haven't created, you know, the next physics.
23:35 Yeah. that enable you know lowcost high performance resilient manufacturable memory right you know for it so I do think and and now and and over the last 30 40 years you know memory was a terrible industry >> right because you would make money one year and then you lose money for the next four right so let's go do R&D in an industry that makes money one out of five years right you know it's just always you know it's just been so terribly difficult because it's been such driven by such violent commoditization cycles.
24:11 >> That has changed a little bit though recently >> over the last 5 years. >> Happy days, man. >> So So all three big memory vendors are on the in the top 20 most valuable companies on the planet. >> Yeah, that's a crazy change, isn't it? >> Oh, it's just so right. you know who who would have thought right uh associated with it and anybody who looks at the AI workloads and it's always you know it's much harder to you know you know it's easy to predict the past but it's a little bit harder to predict the future right
24:38 and you know in that you look at it and say AI is a memory role >> exactly >> right you know inside of it so I think there is the combination of capital and technology needs that will justify memory innovation now there's two domains things that I'm excited about uh you know for it you know one is you know this idea you know I do think that the shoreline bandwidth limitations of HBM is a fundamental issue >> right so with that the idea you know and obviously you know we have a company Dmatrix that's doing this but I
25:09 think others will be you know whether it's you know Cabbrris you know and their techniques >> the future is stacked yeah >> somehow you got to bring memory and >> in more ways than one >> you know you got to bring memory and compute together >> right you know in a fundamental until uh new structures. Um I do think that uh given the workload now has such capital being deployed against it you know in the AI space that I do think we'll actually be able to break through some of those physics challenges as well uh as well.
25:41 So I do believe and you know I've just substited a new memory company as well that's still a stealth but uh you know you know I'm quite excited about some of the innovations that will occur in new memories and you're going to see new materials uh you know for memory people you know a lot of people looking at things like ferro electrics uh you know finding non-capacitive you know uh highdensity memories finding memories that can you know be high performance high density and stackable right but you know thing you know
26:13 people are looking at uh lowcost uh flash you know for that I'm not particularly enthusiastic about that idea myself I think you know speed and performance of the underlying cell are somewhat problematic with a flash or an MRAM structure you know but I do think there's memory innovations are coming for the first time in 30 years and I do think that's going to be an exciting space and you know we'll have a couple of companies there I'm sure hopefully we'll do a few of them together What do you think?
26:42 Should we do Okay. You know, but we'll we'll find some things to do here. But I think memory innovation for the first time in 30 years is nigh upon us. >> Yeah. Yeah. >> It's it's amazing what a couple of$10 billion of market cap. I don't even know. You say a trillion. Okay. Sorry. >> You know, I think the memory industry has now increased by market cap by $2 and a half trillion dollars.
27:03 That's great. You know, over the last four years, >> I think that's about right. That is Yeah, >> we could have just invested there and not done any forential stuff. >> Yeah, exactly. Use your money. Um, how long what's your view? Sorry, go ahead. >> Yes. To to expand a little bit on the stacking part if you have to predict h how quickly will our chips get get taller?
27:27 Is this going to be like, you know, we're going to go to two and then sit there for a while or will we have, you know, the 16 layer sandwich, you know, happily mixing logic, memory, >> pulled down this HPM is going down at Nvidia, but how how do you see this evolve? Yeah, you know, I think there's um um you know, the problem with stacking, right, is now essentially the value of the stack, right, becomes greater as you go to higher dimensionality of the stack, which means the yield of any individual layer in the
27:59 stacking process has to keep coming, you know, essentially it can't go up linearly. Yep. Right. you know, it has to go up exponentially and improving the manufacturing and defect characteristics in an exponential way. You know, when you get to eight or 16 high stacks gets to be so so hard when you run the math on that uh associated with it. You know, your manufacturer flows have to be so close to perfect.
28:24 >> Although you could see designs that take into account a certain model cost like we've seen it for like a long time for two dimensions. >> Yeah, of course. But remember that was the nature of DRAM. they always have bad block replacement and you know spare lines and so on like that. So I'm not in any way, right, you know, minimizing that, but assuming you're doing that, you still need the yields to be so good.
28:44 You know, if I have a cracked die, I don't care how much resilience you've built into it. You can't, you know, it's a bad die, right? So, you know, those kind, you know, uh, so as a result of that, I'm not, you know, a crazy 16 32 stack, you know, kind of guy. I think there's going to be, you know, nice three, four, five stacks, you know, that just end up sort of being the sweet spot.
29:05 >> The mid-rises. Yeah. Yeah. Right. Uh you know for and I and I expect for you know part of its manufacturing part of its complexity uh associated with it you know where's those yield sweet spots associated with. So that's sort of where I think most things land. But even there right you know so let's say I now have you know I now have a great compute tile right you know how big does that tile need to be?
29:28 I mean everything's become reticle limited now right? you know hm can I go to more chiplets that have better yield characteristics you better thermal you know and then you're going to say I want a memory stack on that you know is it a two stack memory a four stack memory you know I don't see it becoming an eight stack memory you know so to me two or four sort of right you know but then power delivery needs to get integrated into that so that becomes sort of another element and I want that in the Z dimension right
29:54 associated with it because you know I want my power rails coming into this as well >> right and Then I need, you know, fairly intelligent RDL layers to distribute signals across that platform. So to some degree, you know, my canonical thing is already about eight layer thick when you add up all of those pieces associated with it. And I think that's about all you can do, right?
30:16 At least as we understand manufacturing today. And then of course you we expect optics to come in. So you're going to have more 35 material structures, you know, that are going to be closer to this complex because I want, you know, my IO getting, you know, essentially, you know, deeply integrated into that construct as well, you know. So my siliconentric pieces need to get complemented by some of these other 35 materials and how do I get optical connectivity built into this?
30:44 How do I get power rails, you know, built into that? You know, that is a, you know, I'll say that is an engineering feat. and a manufacturing nightmare, right, to bring all of those pieces. But that's where the physics is taking us uh over time. So because of that, you know, throwing a 16 high stack and to me, you know, it's like that's crazy. You know, a two or four high memory stack, you know, I think that's probably where the speed sweet spot will be.
31:09 >> And you didn't talk about the thermal dissipation >> and new materials needed for that. >> Yeah. And obviously there's going to be a lot of, you know, you know, first thing, right? You just can't have some of the, you know, the physics hotspots that you have today. So you have to be more intelligent about how you use the the heat and spread the heat.
31:24 But you know then you're going to be saying okay how do I both get power in and dissipate power out uh of those environments. And by the way if I can you know de you know get much more efficient power in at the voltages I want with much better you know power responsiveness and my voltage regulation you know then I don't create as much thermal chimneys.
31:45 >> Yeah. But I do have to think okay how do I cool these things and you know what role do new materials diamonds and other things have been postulated as ways to you know do a better job of the thermal cooling how do I then have better you know liquid cooling techniques and others you know sitting right on it you know with good thermal characteristics so yeah you know uh you know we we engineers are becoming plumbers right >> exactly yeah we got to worry about MEC the chip scale >> let me ask about an operate
32:14 phenomenon which is some people have said hey let's get the memory off of the package and just go all optical between the compute and the memory and that'll give you more degrees of freedom and more ways of building things what's your view on that as opposed to stacking that's other extreme >> yeah you know I think the more OEO right uh interfaces that you have you know it's you know you know optical to electrical to optical to you know the more of those that you demand in your core compute complex you enmy is just
32:48 hard right uh in that sense and you're always going to have losses associated with it you know people are you know working hard to come up with more clever techniques to decrease you know both the cost and the power right you know how how much of the power am I wasting as I go through my you know my DAX and uh you know conversions so I'm not a fan of those you know because I think it gives you access to large pools of memory right but I I burn a lot of power getting there.
33:16 >> Yeah. >> Right. And remember my you know the essentially you know the phentojoules per compute is now actually pretty amazingly good. Right. The phentojlesw per bit of communication is comparatively a thousand times worse right and if I'm spending that much to go get to large pools of so I don't like those approaches you know myself. I think they somewhat defy the underlying physics of the solution.
33:40 I mean and then there's sort of the other approach to say like look can we actually integrate compute and memory directly right shift some of the compute into the memory it's >> some of the pim approaches >> yeah exactly right so what >> what are your thoughts around that >> I'm not a fan of those either right um because I think workloads you know because now I have to take my workload and I have to like push these little pieces of compute into the memory right you know so that you know I don't need as much bandwidth
34:09 of memory coming into my real compute Right? And that puts constraints on workloads. And again, you know, I could be wrong on this, particularly some of the design flexibility and the intelligence that we can now bring, you know, through the the design aspects. But most PIM ideas, you know, have been around for 2530 years. And I think 2530 years from now, they'll still be hanging around, right?
34:29 In that sense, you know, give me great memory bandwidth, you know, that I can have, you know, high density, uh, high performance, and get it close to the real compute engines. I expect that's going to be a more generalizable solution across many and evolving workloads over time. >> Well, so we're still stuck with denser and denser chips, more compute, more logic, more power, >> and much better IO.
34:54 >> Now, I am a I am a huge fan, right, of opt, you know, moving to optics for all IO function. Yeah. Right. You know, I declared the death of copper about 25 years ago. Eventually, I'll be right. >> You're still right. one day one day soon uh I'll be right but I think is that a you know to me all IO should go to optical right you know but that core compute memory complex no right >> so when do you the first step for it going um all optical is to scale up even more so than scale out >> what's your views on that >> yeah
35:30 you know and I think we're now >> what would be the like the application or the workload driver if you will Well, I think the workload driver is the one in front of us. >> Okay. >> Right. I mean, you know, today, you know, our scaleup environments, we're putting tons of copper, right? You know, and essentially copper is becoming a wave guide that's getting shorter and shorter, right?
35:48 And all of a sudden, you know, to make my copper work at uh, you know, 5 m is more expensive than making optical work, right, at 100 meters, right? You know, so you know, I think the physics is driving us there today, >> you know. And you know so if that's the case Pat why haven't we moved there already? Yeah it's a complex supply chain that you know has never been tested at scale and I think now you know I think everybody is looking at you know some form of NPO CPO etc you know in the 28 29 time frame because you know
36:21 I can't scale up to large radics you know compute uh clusters for my scale up environment without making that move. So, you know, clearly Nvidia has uh indicated that. I think everybody else in the industry expects, you know, that's the conversion point. You know, in reality, I should have never built an NVL72. You know, it's an engineering marvel and it's a manufacturing nightmare, right?
36:44 You know, it took the industry, you know, one extra year and a half to digest that beast, right? And it really is amazing that, you know, we're able to get it to scale, you know, so on. But you know that really proved that had we had a more mature optical supply chain, we'd be much more advanced today. We'd be able to build better scale up environments today, you know, better, you know, power, better cost, better energy.
37:08 But we didn't have the maturity of the optical supply chains to allow that and those things are not trivial to get resolved. There's a lot of work to happen, but I think 28 29 is the year >> capacity. Good. for CSS. >> We we're probably gaining a lot of experience how in package optics work in networking, right? And and I think many of that will translate nicely to to >> GPU.
37:31 I do think but all of a sudden, you know, it's like, okay, these things have to work, you know, at 10 million GPU scale. Wow. Right. You know, it has to fit in packages and, you know, there's just a lot of things there, you know, and many forms of optical don't like bad thermals. Well, guess what? Bad thermals is the nature of the game for it, you know.
37:49 So we have thermal problems to deal with, you know, package integration problems. We didn't have enough, you know, laser capacity in the industry. So a bunch of things, but I don't see anything there fundamentals. It's just a lot of hard work at this point to pull all of those pieces together. >> In terms of the network architecture, will this also mean a switch towards more circuit switch networks just because, you know, you you can't scale these these very fast optical networks indefinitely?
38:14 Well, you know, uh, you know, I like, uh, you know, uh, Nick Mune, you know, another Stanford professor here, good friend of mine, you know, as well, you know, I mean, you know him a little bit also, uh, Guido, you know, >> was he adviser? >> He was not the advisor. >> Yeah. Yeah. So, but, you know, to me, he describes it well, you know, we built, you know, uh, networks for to be able to handle any packet going anywhere with no knowledge of where it might go.
38:41 When you think about AI, you know, it's almost exactly the opposite. Yes. Right. You know, we're doing flows that are large and predictable. >> So, in that sense, just the very description says, hm, that looks more like a switch than like a packet. >> So, I do think this confluence of moving to optical, right? And, you know, having a workload that is far more predictable, manageable, you know, large and flow, you know, to me, you know, I do see that.
39:12 Yeah, we will go all the way to bringing optical into the compute complex and we will go all the way to OCS or some form of optical uh switching you know with more predictable flows you know I mean that is the right you know architecture and then you sort of say is that scale up or scale out and to some degree >> say it sort of dissolves the boundary there >> yeah right if my scale up is big enough then yeah you know just saying okay what's the radics of my scale up environment and when do I want to hook it up to other
39:40 clusters that are really big and to me that ends up being a pretty beautiful answer uh you know to the future but I'm looking at somebody who's way more expert on this than myself. So >> no I I I look I always found scale up versus scale up to some degree an artificial distinction. You understand where it comes from right is we're kind of using different protocols one is front end one is backend network and so on.
40:01 So right now practically speaking it is a it is a big difference. >> Yeah. >> If you take a step back and look at it from a system architecture perspective it achieves the same goal you know with slightly different flavors slightly different protocols that over time I think will converge. >> Yeah. >> I mean especially for the AI training workload. >> Yeah.
40:21 Although I mean training and inferences actually look fairly similar. >> The inference corp is less of >> Yeah. Now I'll go on a you know right. Does your brain stop learning? >> Yeah. >> Right. you know you know when you know by the way as another characteristic of workload evolution I do think the you know the separation of training and inferencing is going to become you know less so going forward not more so right that's a little bit of the argument against you know you know more and more heterogeneous architectures
40:44 is I think continuous learning you know I think there'll be algorithmic domains that will say yes continuous learning you know becomes the nature of the algorithm where yeah I do want my inference environment to be updating my model weights because you know that's giving me more and more refinement you know evolution across you know uh workloads moving into environments that are less formalized training uh as well we'll have smaller data sets in those environments so you know I'm not sure that's a >> but the number of
41:15 nodes you need in order to be useful is a smaller diameter for inference than than for training that's what remains >> today that's true >> yeah yeah >> I mean I I still think there might be a role for hetrogenity coming from a different angle which is Different types of AI workloads require very different underlying like some are more comput limited some are more more memory bandwidth if you have a small diffusion workload or so right the optimal card for that looks very different from you know a top-of-the-line you
41:43 know 15 trillion 15 sorry 15 trillion weight LLM right that needs to operate at massive scale right so it's I I think that that might driveity if anything >> yeah I I think there's some room for that you know but I'd also say okay now I'm talking about putting a million GPUs to work in my data center, right? And you say, well, you know, for that particular diffusion workload, you know, I should really put something optimized this way, right?
42:08 So, I'm now going to start, you know, putting, you know, a thousand of those in, right? And for that workload, I need, you know, 2,000 of these in and, you know, for some, you know, and all of a sudden, you know, you don't get some of the, you know, the flexibility of scale, right? that you would say and you know how much does it cost me to manage and upgrade and you know connect and you know fault domain you know all of those types of things >> the agents will do up all the upgrades so there's no longer a problem >>
42:35 yeah yeah yeah thank you everything gets easy so good out there >> yeah so fleet management is done um so we talked about what compute memory networking power is the next one >> yeah yeah yeah >> and you touched upon it with vertical delivery inside the chip on the chip. >> But I mean there's 800 volt DC, there's power delivery to the data center. >> Um there's an entire ecosystem all the way to power generation and and the sociopolitical dimensions of that and so on and so forth.
43:14 >> Yeah. Yeah. >> So where does innovation come in and where does like its pure execution? Well, you the you know and if we start at you know uh first principles here you know our our nation has done a terrible job with this energy capacity right you know and essentially we went for 10 to 15 years where essentially I was taking coal offline at the rate I was adding renewables and essentially the nation was flatlined in terms of overall energy capacity.
43:42 Not good because in a AI digital age energy capacity equals economic capacity right so essentially my economic capacity as a nation flat for 15 years if you buy that thesis right which I think is very you know provable over the last 5 years I you know we've seen this enormous influx into just give me more capacity uh but you know I think uh you know even with that we've gone to maybe increasing our national energy capac capacity 4% per year >> something like that >> going from you know essentially 0 to one right wow 4x
44:20 or hm 4% right so to me you know that you know it's a bad situation so we just need more energy capacity >> uh for it so number one and obviously we have some companies in that space you know we talked about nuclear operating as one of those with Alva um and uh you know I think nuclear right you know as a base load is you know fabulous thing that we want to go build more of.
44:43 You know, unfortunately, all of our renewables have, you know, deep dependencies in China, which is very unfortunate uh you know, for us. You know, we're, you know, uh, you know, gas turbine lead times are only eight years at this point, right? So, you know, you know, I mean, we really have a difficult environment to scale up rapidly, but one, we just need more capacity.
45:05 Yeah. And fundamentally energy capacity somewhat dampens how euph fork we can get on AI, right? Because you know why build a new data center and buy, you know, the million GPUs if I can't power them, right? And I think you're going to see more and more defaults happening on many of those data center projects because the energy won't be there. >> So you think it'll be a significant headwind?
45:28 >> I think it will be a headwind and I think you're already starting to see some of the first indicators of that. the Oracle, you know, uh, comet, you know, was just I and I think that's the first of what you're going to see many. Now, it's not like we're going to slow down as a result, but I think people will start saying, do I really have the energy, you know, to underwrite the capital commitments that I'm making, you know, for, you know, cement for data center racks and, uh, GPU purchases.
45:50 So I do think it's you know going to be a you know an upper bound >> but we need more energy capacity and we need to get more innovation into that and we need more of those supply chains in the US you know nuclear one of my favorites you know when was the last nuclear reactor that came online in the US right you know it's 20 years ago >> right you know we just stopped you know a very capable uh industry so that's one we need more uh you know we need to get much more efficient at delivering you know the power networks
46:21 and as you you know, moving 800 volt DC data centers. Absolutely. You know, right, you know, and being able to eliminate so many conversion steps along the way because of uh of uh uh historical uh you know, reasons and you know, standardization reasons. Okay, we got to fix this. You know, an 800 volt DC, the revenge of Edison is upon us. >> Yes, I was just going to say, >> yeah, I like to say so, you know, I do think that ends up being we need better power conversion.
46:50 You know, we have a couple of companies, one of those vertical, you know, vertical GAN, you know, I think will be, you know, a killer technology, you know, to be able to do 800 to 48 or 800 to 12 or even 800 to 5 in one conversion step, right? Which just gives you a lot more uh energy, you know, we touched on >> solid state transformers, you know. >> Yeah.
47:10 You know, all of those are, you know, solid state transformers, switches, etc., you know, rebuilding, you know, the power delivery network. And I think that will, you know, have a lot of, you know, innovation, you know, it's sort of like, okay, power's cool again, right? You know, you know, for it. And then you're going to need all of those in the cooling systems as well, right?
47:27 You know, and so there'll be a bunch of innovations and that, you know, so you know, you know, turbines and refrigerants and, you know, all of a sudden they get to be, you know, exciting new technologies again. So, you know, for all of us hardware guys, this, you know, this, you know, is like renaissance in front of us. It's just going to be >> after 25 years.
47:44 I had to go back and read about car's car >> carinal efficiencies. Oh yeah. Is it like good stuff? How does that first law of thermodynamics work again? >> Something after all here, >> you know. So, so it's going to be you know exciting time that way you know but you know you know we also say you know we need to change the physics right because cos hasn't improved you know the essential power per terror you know power per teraflop for the last five uh generations of Nvidia chips you know flatline not okay right and
48:15 that's where some of the things you know we have a superconducting company snowcap right which affords the opportunity to fundamentally change the physics you know a thousand times better power performance. So, I do think some of those innovations are upon us in the near future as well. >> I've been a terrible time manager here for the past hour. So, so many topics yet to cover.
48:37 We didn't cover, but we got to cover at least one before we leave. >> Okay. >> Because now >> you're in charge. >> Yes. So, I thought um now VMs are back. >> How do you think about that as a former CEO of VMware? >> Well, you know, I I uh I feel excluded from this conversation obviously if we go on >> part of it was I was going to talk about networking as well so but I don't think we have the time >> you know hey you know I I think the you know every innovation has always led to the next abstraction right and I think
49:13 many of these ideas and you right you know I always like to say you know you know what's what's new in our world today you know data structures algorithms you know an abstraction you know you know we bring it back to those you know foundationals and hey you know I you know the virtual machine abstraction you know whether it's done from the infrastructure level or from the application level you know I think it's still a foundational abstraction model you know that deserves to always have a place uh in the compute
49:37 hierarchy so you know so what do you think one of us go back and run VMware again >> yeah that's what I tell I think I told it's time to go back and run VMware again >> yeah yeah and uh you know and I do think some of these abstractions you know but you also think about it and like one of the things we did at VMware where we you know essentially you know managed every aspect of the computing hierarchy.
50:01 >> Yeah. >> Right. You know and you manage the workload, you manage the network, you manage the storage you know system and how do you abstract those? And I think in the AI context, right? We think about you know these agent swarms. Well, who's going to manage the agents? Who's going to create the security profiles around all of the agents? Who's going to be manage the performance of the agents?
50:21 You know, essentially every fundamental element of virtualization and management needs to get recreated in this next computing hierarchy and there's going to be lots of wonderful companies that you'll sort of break through for how those things will be, you know, done. >> Okay. Do you want to end it on networking? >> Actually, I'd love to stay with the the VMs for a second.
50:42 And the most interesting thing for me is that we're now building VMs so that they're usable for agents as opposed to usable for humans. What we've seen in sort of other companies is that that changes a lot the form factor that changes a lot how you market them, right? You want you want the agent to make the pick for for for your VM offering. >> Um it it changes a lot how much complexity you can have.
51:01 It changes how how quickly the startup should be. Humans are a lot more patient than than agents. >> Um >> so it's if you have to rebuild VMware today, how would it look different? >> Yeah. And and I do think >> like the VMware for agents. >> Yeah. And I and I do think there's, you know, does there need to be a VMware for agents for humans, right? You know, >> I see what you mean.
51:24 Yes. >> Right. You know, in the sense that, you know, I need humans. >> Let's stick to the agents. They're way more interesting than right now. >> Yeah. But, but, you know, I think it's an important point because I do think that there needs to be, you know, and I'll use, you know, Dario's constitution, right, or guard rails, right? Or being able to set policies.
51:43 And I think then you you really say okay every aspect of the virtual machine as we know we'll love it today it is servicing agents >> right as opposed to servicing hardware and humans right so I think you know that becomes your fundamental design constraint on this side you know how do I make agents great >> right how can I make them secure how can I make them performant you know you know uh you know how do I abstract them how can I migrate them all of those things you know have the new embodiment you V motion for
52:13 agents right it's like every one of these will have the new embodiment you know for them but then I also need to think about how can humans be setting the policies describing the constitutions you know getting the you know dashboards uh etc and I think that's sort of you know where the interesting thing is because you need to do both of those right not just one of those but the fundamental you know I'll say the hard design stuff like we did for you know virtual machines was associate abstracting the hardware in a
52:41 performant way for applications or operating systems here. It's, you know, abstracting the hardware and the operations for agents. It's a fundamental thing that you're serving. It >> was great discussing this and thank you so much, Pat, for coming in and we'll do it again. >> Well, two of my favorite people. Anytime. >> All right. Thanks. >> That was awesome. Thank you. >> Thank you.