Searchable transcript of New AI models, token minimization and IBM’s new sub-1nm chip — IBM Technology (51:02). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:01 This transistor. The technology we talk about at the Angstrom provides 50% better performance, or it will save 70% more power when it does the computing. This is a massive improvement that our industry hasn't seen for a long time. All that and more on today's Mixture of Experts. I'm Tim Hwang and welcome to Mixture of Experts. Each week, MoE brings together a group of the researchers, builders and thinkers working in artificial intelligence to walk you through the week's news.
00:37 On this week's episode, we've got Abraham Daniels, Senior Technical Product Manager for Granite. Martin Keen, Master Inventor. And Gabe Goodhart, Chief Architect, AI Open Innovation. We're going to talk a little bit about A24's partnership with Google DeepMind this week. We'll talk a little bit about the phenomenon of token mining. And we're going to talk a little bit about these new models that have hit the scene.
00:56 Sakana Fugu and GLM 5.2. But first we've got a special segment with Sascha Brodsky interviewing Huiming Bu VP of Silicon Technology Research and Development on the sub-1nm chip. Hi, Huiming. Thanks so much for joining. Hi. Thank you for having me. So before we get into the technology, help us understand why this announcement matters for someone who doesn't follow semiconductors every day.
01:26 Why is breaking the one nanometer barrier such a significant moment? So let's start with what semiconductor essentially are the chips. I would say charge every single thing we are doing in our modern life when it comes to devices. So if you look at it at that angle and realize our industry has been relying on semiconductor for transistors scaling for the last 60 plus years.
01:53 What does it mean? A MOSFET actually was invented in 1959. And since its introduction we have been making the transistors smaller and smaller, more powerful, and consumes less energy. So far, we have been successful for most of the time, but in the recent years we have been hitting some of the very, very, very difficult challenges. Takes billions of dollars to invest, provided solutions.
02:30 And we are at a point now looking back at the 60 plus years we have been doing transistor scaling in the X direction and a Y direction, essentially, in a two-dimensional world, which we call it defined by lithography scaling. Today, this sub one nanometer announcement is for the first time in the history of the semiconductor industry. We are stack the device in the vertical direction, Z direction, which is a direction our industry has not explored in the past 60 plus years.
03:07 And now we are at the end of the shrinking in the two dimensional. The third dimension of Z provides tremendous opportunity for the future of scaling. So interesting. Now, for years, we've heard that Moore's Law was slowing down, that the industry was running into the physical limits of silicon. What challenges were you trying to solve with this technology?
03:31 When I joined the company in 2005, we faced a major challenge in transistor. So what happened back then is that when we tried to make the transistor smaller and one very, very important material in the transistor is a dielectric. We call it a gate dielectric. And is that a point of time it was silicon dioxide. It started to leak like crazy when you make it thinner.
04:00 And we provided the solution. The solution came with a new materials. We introduced high-k metal gate to replace the silicon dioxide and a polysilicon. At that point of time was a gating material for a traditional transistor and it was suppressed. The gate leakage and the scaling. We call the Moore's Law continued. Another example is only a few years later.
04:28 What happened is there is for the first 45 ish years of the transistor scaling, we have been using a planar device. As I mentioned in 2D that is called a planar device, and when you make a device smaller, the short channel effect shows up, which means there is another leakage starts to show up. Electrons start to flow in the channel when the device is even turned off.
04:57 So at that point of time, our engineers introduced FinFET. I can show you a cross section, actually, about of how a FinFET looks like. This is a gate. This one is a fin similar to the fin of a fish. That's why it's got a FinFET the gate wrapped around this thing on three sides. Left top and the right side. So it's called a tri gate device. And this device has much better gate control.
05:26 Hence the suppressor leakage in the channel. These are a couple of examples of the key challenges. One are the material side and one on the device architecture side we faced. I would say at this point, more than ten years ago. IBM is calling this the world's first sub nanometer chip technology. What is sub one nanometer actually mean today? Is that literally about the transistor size, or is it shorthand for a new generation of manufacturing today?
05:59 The technology node name is no longer directly correlated to any physical dimensions in the real device. It actually becomes a node name quite a few years ago. And it's you can consider it's just a name to show the technology keeps evolving. So after two nanometer today, best available technology in manufacturing is two nanometer after two nanometer is 1.4 nanometer after 1.4 nanometers is one nanometer after one nanometer.
06:34 There's 7 Angstrom technology. But one thing you probably have realized is if you divide the number, the new number versus the old one. Generally speaking, the ratio is 0.7 that is called a scaling ratio. That is a real number. So we have kept scaling ratio of 0.7 because the point 0.7 times 0.7 is essentially 0.49, and 0.49 means your area of scaling, or in other words, your density keeps increasing.
07:03 That is the scaling factor. What are the phrases that stands out in this announcement is nano stack. What actually is nano stack, and how is it different from the chip architectures that are being used today? Yes. Great question. Here is the cross-section of the nano stack. I'm going to let you just look at it for a few seconds. I explain what's actually in front of you.
07:30 So let's look at this direction. If you look at it, you're going to realize there are two nanosheets here. This is one nanosheet. Here is another one. They are not on the same plane. One is on top of the other, but not exactly right on top of each other. This is a transistor design. We call it vertically stacked and also staggered. So think about two transistors.
08:01 You stack them, but you also stagger them both in vertical direction. That's why we're saying nano stack is a device enabling the transistor stacking or scaling in the Z direction, in the vertical direction. The beauty of this device is the following. One is these two devices or two wafers. Top wafer from bottom wafer, top device, bottom device. They are actually not monolithically defined by litho etch process.
08:33 It is actually using a thin dielectric to bond them up. Also, because of the thin dielectric bonding process, these two devices can be optimized independently, which means you can use the best material you can think about for the bottom device. Best material you can think about for the top device, and as they can be optimized for performance independently.
09:02 The other unique nature if you look at this device is because they are stacked and staggered. Both the front side of this top device and the back side of this top device. Front side of the bottom device. Backside of the bottom device can be contacted with your signal line, and a power line directly. That itself provides tremendous density scaling benefit with this nano stack structure.
09:33 This chip packs nearly 100 billion transistors into something about the size of a fingernail. At that scale. What are the biggest engineering challenges? Are you fighting heat, quantum effects, manufacturing precision, or all of the above? So let's do it one by one. I explained each of these. The thickness of each of the sheets is about a five nanometer.
09:56 You can consider there is some process margin with it, whether it's a five and a half or four and a half. So let's talk about a quantum confinement effect using the four and a half nanometer sheet. Do we observe quantum confinement? In fact, in this device, the most pronounced observation when you observe quantum confinement effect is threshold voltage variation.
10:19 So far, we have not observed the quantum confinement effect in threshold voltage variation for the thickness down to four nanometer. So we are still operating in a safe zone away from the quantum confinement. The second one you asked is heat effect. Yes, heat is a very, very important factor in this device. The way to think about it is these devices that are vertically stacked, you have less heat dissipation path.
10:51 So you do have to have innovative way to provide thermal conduction, just like how you do the electro conduction to do the same thing, to conduct thermal better in this device. And we have engineering work to provide that solution as well. AI is driving unprecedented demand for computing power. How much of this breakthrough was motivated by what you're seeing from generative AI and increasingly demanding AI workloads?
11:19 Yes, that is the fundamental reason we are enabling this device. I spend quite some time to talk about why you would observe density benefit, but let me talk about the power performance of it with this structure compared to two nanometers we announced in 2021. By the way, as I mentioned today, the best, most powerful available chip is manufactured with two nanometer.
11:45 Consider that is the best world's best chip compared to that one. This transistor. The technology we talk about at the Angstrom provides 50% better performance, or it will save 70% more power when it does the computing. This is a massive improvement that our industry hasn't seen for a long time. That is why we are so excited when it comes to high performance computing, especially for AI computing.
12:15 Another thing I want to mention maybe share with this audience is the other benefit of this transistor architecture provides when it comes to AI computing is the very significant extra size scaling the nano stack provides 40% extra area scaling compared to the two nanometer. That is a significant step forward in scaling that our designers haven't seen for more than a decade.
12:48 And all the AI computing requires memory. And SRAM is embedded memory solution with your computing unit. This level of the SRAM density scaling enabled by nano stack will truly help AI computing to be a lot more efficient moving forward. So finally, if we're sitting here ten years from now, looking back on today's announcement, what do you hope people will say this breakthrough made possible?
13:17 I would say this nano stack architecture innovation is in the first time after 60 plus years in semiconductor industry has truly made our computing logical device go from a 2D to 3D in its design for the first time. That has opened a completely new dimension for the future of scaling, which at least we can see 10 to 15 years of technology roadmap ahead of us.
13:52 Thanks so much for joining me. This has been a wonderful opportunity and I appreciate you being here. Thank you for having me. First, I really want to talk about these two new models that have hit the scene and have been at least eating my social media feeds, completely over the last week or two. One of them is a new model, a coding model from Z.ai, a Chinese lab called GLM 5.2.
14:21 And the other one is a model by a Japanese lab called Sakana called Fugu. It's sort of an agentic model. And we'll talk about both of them. You know, I think the first one I wanted to talk about was Sakana, though, because we had never, ever talked about them. They were not on my radar at all. This is the first time I had heard about them. And the numbers for the Fugu model are really shocking, right?
14:44 They have these charts where it's like, oh, you know about Mythos, that model that everybody's been talking about for a really long time? Well, we're we're already there. And actually above in some respects, I guess, Martin, maybe I'll bring you in. What what's going on here? I thought Anthropic and OpenAI were, you know, so far ahead, so on the cutting edge that no one was ever going to catch up.
15:03 And as yet, I guess, I don't know, maybe I just haven't been watching the Japanese AI scene, but the idea that a model could come out and have these kinds of metrics, it's a little bit shocking, isn't it? Remember when Claude Mythos was like the best of the best? That was like last week? Yeah. So now we have something else from. Yeah, from apparently out of left field.
15:21 So yeah, it's absolutely crazy. And you know, it does particularly well in coding tasks. So it's doing very well on LiveCodeBench benchmarks, things like that. Which, you know, already, we could say the the big frontier model labs, they are absolutely laser focused on this. It's not like they've picked a metric that nobody was really looking at and then suddenly maxing on that.
15:44 So, yeah, this is this is really impressive to see and just, just kind of the whole Sakana AI platform is really interesting because this is a multi orchestration model. So you send your request in and it can go to one of multiple models. And in fact the models can change at any point. So if a model goes away or a better model comes along it can kind of slot in there.
16:05 And it is interesting to think about going forward. Are we going to see more multi orchestration kind of harnesses where we just go to multiple platforms multiple models and and that changes constantly. Or are we going to continue to see this focus where we you know. We're all getting super excited about this. This new model that's come out and we're suddenly sending all of our workloads to that new model for a couple of weeks until the next one comes out.
16:31 And I'm really interested to see where that goes, because this, this whole idea of multi-agent orchestration and a platform means that if there's one model that's really good at a particular task, and it's another model that's really good at another, that they can diverge, and it will be interesting to see where that goes. Gabe, do you, do you buy the numbers here?
16:51 I mean, I don't know, I look at these stats and I don't know, they're just sort of so unbelievable in a certain respect, you know, my immediate instinct is, oh, is this metrics cherry picking. Like, should we, should we buy kind of what they're putting forwards or I guess like Martin I mean, you know, Martin sounds like you you buy it. This really is the real deal.
17:09 I know Gabriel on a similar, similar page. I think the numbers are great. It's really, really cool to see success with this architecture. But I think we're we're doing it a disservice by calling it a model. Because fundamentally, that's not the innovation here. Like, yes, there is a single novel model here that is the router. But everything else is all the existing frontier labs.
17:31 It's just passed through. So what they are doing is not actually creating a. Net new model that can achieve these benchmarks. Like you were alluding to. Martin. They are figuring out how to get the best results out of the models that already exist, and stitch those together in a meaningful way to solve problems. And many folks on this panel and in other discussions have been saying for a long time that there's a lot of greenfield in this orchestration space.
17:56 But I think in some ways, the fact that every, every endpoint looks like a list of JSON objects with a role and a content field, and, you know, you could put whatever you want behind that API. And in this case, what they are putting behind that API is a carefully routed, orchestrated multi-model ecosystem. So calling that a model endpoint is not very accurate.
18:21 And the other piece of this that is interesting is that one of the claims they made, is that this gives a, an enterprise resilience to model fluctuations and model changes. And that's true. But on the flip side, it also means that your quality could change wildly depending on which models it actually gets routed to behind the scenes and how well the routing performs for a given task.
18:49 So there's you're essentially trading the fact that you'll always get some answer for the fact that you really can't predict what answer you're going to get. You get even more non-deterministic because now you're not just dealing with the non determinism of the specific model, but you're dealing with the additional non-determinism of which model your request is even going to land on, or which collection of models are going to service the subcomponents of your request.
19:14 So it's it's a wide open field. So to your question about do I believe the benchmarks? Yes, I believe that at their perfectly routed settings, they are actually able to achieve better than any of the individual models can achieve on their own? Do I believe that that's going to be the floor of quality? No, no I don't. So I think it's probably just got a much wider aperture for quality than the other ones, because it's got a much wider set of quality models behind it.
19:42 Yeah. It's like those robotic demos where the robot does it like perfectly and you're like, this is the trial one of 100, and the other ones are the robot, like failing to be able to open a door just like falling over and up. But usually it's not trial one. Usually it's trial like 74 out of 100 just right. But as yet, I mean, so all these points are granted.
20:04 Abraham. Are we? I actually wonder whether or not we worry about the wrong thing with AI. So obviously the big stress around Fable 5 about Mythos has been, say, the security aspects of all this, right? That the models are now so good that it poses this huge cybersecurity risk. And I think granted. Right. What we're seeing here is an assemblage of models, not a model in of itself, but it kind of says that maybe like the the is the cat out of the bag already.
20:27 Right? Like if we feel like Mythos is a model that's so dangerous that we need to do Project Glassing or whatever, isn't the fact that you can kind of stitch together a bunch of models and get to the same result mean that we're trying to control, like, the wrong thing, or in the very least we can. We can already get the kind of scary things we're worried about with sort of not commodity parts, but certainly like off the shelf parts.
20:43 What do you think about that? I mean, I've never thought of it that way, to be honest. It is interesting that you bring that up the way I kind of saw this from Sakana's point of view was it kind of pitches frontier capability without the risk of export controls, to be honest, or government intervention without any sort of, control over the actual user of the model.
21:05 So I kind of saw this as like maybe not a direct shot. But like, you know, worse at times, you know, US frontier models are getting restricted. Enabling orchestration gives you a little bit of a protection around what you can and cannot use. Given you could build the same capability and aggregate kind of to your point, I was really going to echo a lot of what Gabe said in this regard.
21:31 Like, I don't see this as a fundamental jump in terms of a new model. I think we're gearing towards orchestration is the product as as opposed to the actual model. Is the product really with Sakana? And I think it does a really good job of kind of providing a very simple user experience, like the idea that it's all behind an API is great because the user doesn't have to care about the orchestration or know what models actually getting hit with respect to a particular, you know, query.
21:51 So I thought that was a really good direction. And I think that's really where things are going to start going in terms of being able to build out frontier capabilities as opposed to frontier models and really focusing on that aspect of how do you actually get the most out of a model for a system of models or orchestration of models for the user? And just to jump on that, you know, something that I know you focus a lot on.
22:14 Abraham is, essentially this virtual model idea and what you can do behind a virtual model endpoint to get the most out of a collection of models or even an individual model being orchestrated carefully. So, the Malaya project out of IBM is a great example of this, and it's not aiming necessarily at pushing that frontier. So one thing I'm really excited about is taking this same concept of a virtual model endpoint and bringing it to the small scale.
22:44 So maybe it's not going to push those capabilities, but if it could push Sonnet capabilities and run on your smartphone or run on your commodity laptop, that might actually be the real innovation where we can take we can get acceptable quality in a much like order of magnitude smaller footprint than we would based on a single model just based on the weights.
23:05 So I'm really excited for this overall trend of orchestrated virtual models. Chasing Mythos and Fable is a lofty goal, and it definitely makes headlines and it's really cool, but I think the real impact is going to be at horizontal spread of smaller models with great capabilities. Let's make me take a few minutes to talk a little bit about Z.ai and the new GLM 5.2.
23:30 Gabe, you just commented, maybe we'll stay on you with these coding models. I'm not doing coding day in, day out. I don't know if you've played with the model yet, but curious about kind of like just your your taste test. You know, I think we can look at the metrics, but really is someone who's doing this day in, day out, I feel like you've got probably a better feel for like, what's the is there a good you know, TFR for the model.
23:48 Is there a good kind of like what's the what's the taste of this one. Do you like it. Yeah. Yeah. So so you know, I guess in fairness I haven't tasted this one. I've, I've read I've read about other people's tasting experiences, but I haven't tasted it myself. And a lot of that. Oh, it's you know. Yeah. Exactly. Exactly. Yeah. Yeah, I've got the flavor profile, but no, I think, you know, staying on that point for a minute, one of the things that's interesting is that with these open weights models pushing size
24:20 frontiers that rival the frontier labs, what we're really starting to hit also is usability gates. Right. So I do the vast majority of my coding in a professional capacity because when I'm not doing that, I'm chasing children around and there's not a lot of time for writing code while I'm chasing children. And given that I'm basically gated behind what is acceptable to do in my professional capacity, so I'm not able to actually pull this model locally because it's enormous.
24:46 And, you know, I also am not allowed to send any of my professional work to GLM. So, you know, that that's just an interesting aside that, yes, this is really cool because it's an open weights model. But the size is still a physical limitation for tinkerers like myself that want to just sort of unbox it and play with it. That said, you know, it still is the story we've been telling over and over again over the past couple of years is that we get major leaps in quality out of the proprietary labs and then open models,
25:12 often out of China, catch up in a meaningful way. And the interesting thing that I think is that it's not just one Chinese lab that's doing this. This wasn't another round of DeepSeek or another round of Qwen. This is, you know, another another lab that, yes, has been making great models for a while but hasn't been the one grabbing the headlines. And now they've kind of surged ahead by a nose and we'll see who wins the race type of thing.
25:35 But you know, I think it's awesome to see that these capabilities keep staying on the frontier. And I also think it's really cool that the innovations at the model architecture level are continuing to be published in the open. So that was one piece that I looked into with this is how they, you know, which which architectural widgets are they borrowing from which other models?
26:00 It's highly inspired by the latest DeepSeek architecture. But it's it's able to actually bring speed and quality at an extremely long, context length, to an extremely large model. And that's that's awesome. So, I love seeing those quality, you know, pieces pushed and again, true to myself, I like things that I can hold in my hands. I hope this same architecture gets released in a much smaller footprint, say, approximately 100 billion parameters that fits on a DGX Spark so that I can play with it myself.
26:37 Martin and I actually want to pick up on the point that Gabe is making. You know, it's kind of interesting. It feels like out of the kind of Chinese market for models it seems like there's a new, new player, you know, every few months. Right? We're like, oh, wow. Is Z.ai, their new, it's not just like another DeepSeek thing versus I think it feels like in the US market, we've gotten very much into the mode of like, okay, OpenAI is doing one now Anthropic is doing one, OpenAI is doing one, but there's not so much like
27:00 we're blindsided by a new lab that comes out with a model that's like, you know, kind of capturing the headlines. And I guess kind of the question is like, you know, do you feel like the US market is like almost insufficiently competitive, like, you know, you only have these two labs now that are doing like a lot of the model work versus where in the US or in China, you feel like there's always a new lab kind of popping up?
27:23 Well, I was not even aware of GLM models at all until this 5.2 came out. So to me, this is just something out of left field. But of course it's not. I mean, this 5.2 tells you something. This is not their first not their first rodeo. So yeah, I mean, it is an interesting point in that there are so many, so much innovation coming out of a, you know, a country that has been traditionally restricted with its compute capacity as well for being able to build these models.
27:56 But, it is also interesting the direction that this has gone, because it seemed like the direction initially we were going is you have big frontier models from the big labs, you know, the big three that we're always talking about. And then you have these open weights models, which are maybe nine months behind everything, but they're also a bit smaller.
28:18 I can run them maybe on my laptop or a dedicated GPU. Well, that is not really the case here. This this model. This this GLM 5.2 is about the same size as Claude Sonnet 4.6. It's enormous. So these are not models that you're just going to kind of take and run in your basement anymore. These are models that need to be deployed somewhere to go to Gabe's point.
28:39 Like, if I want to access this model now, firstly it needs to be somewhere that I can get it because I can't install it myself. And now we have all sorts of governance issues as to, well, am I allowed to access it and who is hosting it for me, and what sort of workloads can I send to it and can I trust who's doing the hosting, all that sort of stuff as well?
28:57 But also it is great to see that this research is now it's published, it's available for everybody to see. It just speeds up the whole development market for all of this. And I've been been following along with, with some of the people who have been able to use this model. And it does seem that in coding tasks in particular, it's doing a great job with a few little caveats.
29:19 Like apparently it had a very large Mandarin training data set. So it will just like, blurt out random Mandarin phrases, apparently, even when you're talking to it in English. So you never know when that's coming along. But yeah, just kind of fascinating to see how these things are developing. Well, I'm going to move this on to our next topic today.
29:52 Really interesting headline. We've been kind of having a subplot of this, you know, as we get into the summer here at MoE, which is, I think, the kind of relationship between the, you know, emerging AI industry and, you know, filmmaking, art, entertainment. And there was a big partnership that was announced between Google DeepMind and A24, which is the studio behind movies like The Back Rooms and Marty Supreme and kind of known as in some ways, like a kind of prestige sort of film studio.
30:21 And, you know, the interesting thing is basically what it sounds like is DeepMind is going to team up with A24, and they're going to build new AI tools for filmmakers to use. And so this is a little bit different from, oh, hey, they're going to just start producing AI movies. It's a little bit thinking a little bit about the back end kind of toolkit.
30:39 And Abraham, maybe a question for you. You know, I don't think it's unfair to say that, like the relationship between entertainment and the AI industry has been pretty fraught. I would say, you know, a lot of people do not like what the AI companies have been doing. And one way of doing this partnership is maybe a kind of like, you know, olive branch peace deal to say, well, look, look, we're going to work with you to develop tools to do filmmaking.
31:01 We're not going to just try to do, you know, AI generated film or whatever. I guess question for you, Abraham is like, do you think this, like, changes the narrative? Like, is this the beginning of what a beautiful relationship might look like? I think between the AI industry and sort of filmmaking and, and the arts or, or is this just not going to be this is just a blip in what is already a negative relationship.
31:23 It might be getting more negative over time. I want I think it's a great olive branch too. I think it's the beginnings of a better relationship, maybe not a beautiful relationship. So, the the way it was framed out, I think was just really strategic in terms of building tools with artists and the idea of trying at least as best as possible to preserve creative control.
31:44 And I know that the first prototype of this was kind of the storyboarding tool, which is really precursor to the movie. So if anything, it adds a lot of kind of value to expediting that process. I read earlier, too, that, you know, Martin Scorsese, arguably one of the most, you know, recognized directors, also partnered with an AI firm to build out some AI storyboarding tool.
32:06 So I think you're starting to see a little bit more appetite from the industry in terms of AI tools. Truthfully, I kind of see this like, you know, in CGI and 3D animation, 3D kind of came onto the scene and, you know, a lot of individuals on the animation side of things or on these, you know, set design were kind of ousted for with with these respected technologies.
32:31 And now it's kind of moving upstream to actresses, actors and actresses. I think it's inevitable, but I do I kind of I do applaud A24 in terms of getting in front of this and trying to kind of frame this more as an additive as opposed to a replacement. Truthfully, let's just see how it goes. Moving forward, I don't really have a position in terms of how this is going to land, but, you know, with more and more artists recognizing that this is going to be an aspect of it and potentially protecting your copyright or
32:59 IP, you know, whether your image, you know, I know Taylor Swift has done that in certain cases. So I think it's just going to be something that we continue to see in terms of trying to protect your brand or really your ability to make money in this industry. But yeah, to your initial question, I do think this is the beginning of trying to mend or at least build a bridge between the two sides.
33:21 Gabe, why do you think A24 is doing this? You know, they're they're obviously a very well-regarded studio on the scene. You know, this using this technology, partnering up with big, bad technology, you know, is maybe a little bit fraught for them. What's their angle? What do you think they get out of something like this? I guess I have a cynical view and an optimistic view on this.
33:41 So the cynical view is that a studio like A24 has the credibility to spend, and they can see where the ball is going to go one way or the other. It's somebody is going to do this. So they'd like to be the ones doing it right. They'd like to set the set, the playing field for everybody else. So that's kind of my cynical take that it's more of a getting out in front of the narrative play.
34:05 But my optimistic take is that, you know, this is actually akin to, I'm going to make the analogy to, again, something I know and love, which is coding agents. I think, you know, you think about the early attempts at AI in the film industry. They're very akin to vibe coding. They're basically single shot. This thing like, I want a video that looks like blah.
34:26 Give me one of those cool at the end. And now, just like in coding, we're recognizing that vibe coding is wonderful for prototyping and for demonstration of the possible and pretty terrible for actual professional usage. I imagine that this is going to look a whole lot like that within the context of the film industry. So again, just like Abe said, I think the framing it as tools rather than models is intentional and important because, as we think about, you know, coding agents in the hands of developers, we think
34:57 about the, the nuanced usage patterns, how you map these tools to the existing, workflows of developers. I hope that DeepMind and A24 will do the same thing in the context of movie production. And so storyboarding is a great example that's akin to project planning and software, and that's a place that you can absolutely get boosted without having to necessarily replace the, the actual quality and architectural, integrity of a piece of software.
35:31 I hope the same thing gets gets held true in the movie industry so that we have high quality stories coming out. They're just coming out better with the help of the new tools. I get that, Martin. Do you want to I mean, so they haven't talked a little bit more about, I mean, to a great point, like haven't talked too much about, like the precise tools they're building.
35:48 But I think it might be fun to do a little bit of, like, pie in the sky. I mean, I don't know if you have kind of dreams about what these tools might look like, but it just seems like there's is a lot to be done outside of the completely kind of naive, like, I need a movie about blue and like the movie comes out. There are a lot of kind of really interesting tools here, but curious if there's anything that you have in mind really like, oh, be really cool if they build XYZ or I wonder if this might be what the future of
36:10 it looks like. If you want to kind of dream a little bit about where this all might go. See, Tim, I absolutely have some dreams about this. To me, this this deal is absolutely wild because DeepMind are going to be creating tools that will help A24 create their movies, right? So how much are A24 paying Google DeepMind for them to develop the tools? Oh, they're not paying anything.
36:36 They're getting paid $75 million. Google are paying the studio for the studio to use Google's tools. That's absolutely wild. And of course, why are they doing that? Well, because they need to understand the use cases of what is it they're doing in terms of the storyboarding and the production and the animations and the, you know, all of it, whereas these labs already know how to do coding.
37:04 I mean, that's what they do. So they don't really need to be doing these big deals with somebody to learn how to code. But they do need to learn how a movie is made. And and that is so valuable to valuable to them that they're willing to spend tens of millions of dollars doing it. So you mentioned dreams, Tim. My dream on the side. I, I like to brew like, fancy pretentious coffee.
37:30 And I have all these, these coffee machines that brew, the coffee. And some of them have started to have AI features in there. So this just in case, like Demis Hassabis is listening to this, I just want to put out an open call if they would like to learn how to brew fancy coffee and how that DeepMind could help with that, I would be more than willing to offer my coffee experience for much less than $75 million, and teach the model how to do that.
37:55 Well, this is great. And a lot more to come. And I think, Martin, I buy your point that like, we will see these sort of interesting collaborations. I mean, entertainment, of course, is the most prominent one, but I wouldn't be surprised if you find like, oh, well, we're going to team up with a prominent coffee maker to build tools for this, right? We're going to team up with a prominent, you know, every domain, right?
38:13 I think like is amenable to this. My friend, you know, actually reminded me that back in the, I think 80 or 90, there was actually a rice cooker that was sold with one feature being that it contained a neural net, and it was like a really, really old school perceptron that was used to just measure kind of like moisture in the rice cooker. And it was using it to kind of improve the technology.
38:35 And so, you know, I think we're going to see all sorts of partnerships like this emerge in the future. All right. Well, let's move to the final topic of the day. This is kind of more of a culture piece. Needless to say, for a period of time, I think corporate America was very gung ho about people using AI more. And, like so many things, you know, what you measure is what you get.
39:00 And so a lot of these companies were measuring their success at adopting AI based on how many tokens were consumed, because, of course, the more tokens that are consumed, the more people in a company are using AI. The problem with that, of course, is that tokens are expensive, really expensive, and sometimes colossally expensive. And so there's been a number of pieces.
39:19 One of them that we caught our eye was a piece by Eli Tan at The New York Times called Tech Workers Maxed Out Their AI Use. Now They're Trying to Minimize It and specifically talking about the phenomenon of token mining. Now, we've gone a little bit too far in terms of token maxing. The idea is how do you get done what you need to do with the least number of tokens?
39:39 And Abraham, I guess question for you is, you know, like, how widespread is this phenomenon? Do you feel like in general, enterprises are starting to be like, whoa, whoa, whoa, we've just spent way too much on tokens. You know, everybody's got to be on this, like, much more fixed diet of tokens now. And is that going to be the norm of the future? Are people going to just start tightening their belts when it comes to this?
39:59 Yeah. I think it's you've seen headlines, you know, whether it's Uber or Microsoft or other enterprises either restricting or limiting, you know, API or access to cloud models just because, you know, within 3 or 4 months they've already swallowed up their entire yearly budget for, you know, for the usage. So I think you're definitely going to see it.
40:18 And to your point, there is, you know, the fundamental shift from token maxing to token mining. I think the narrative's shifting away from, you know, more tokens means more productivity and more so token economics, like, how do you get the best result for the unit economics of a single token? And then kind of that will I think that kind of plays into our earlier, you know, conversation about orchestration of models to define, you know, how do we get the best output independent of whatever model was used using some type
40:47 of, you know, orchestration layer? So, so in short, yeah, I think this is kind of the direction forward. Inferencing costs exponentially more than to organizations want above the person model training. So there has to be a better way of actually, you know, building a business case around this and the user experience. I just think it's it's it's something that we, you know, we're already starting to see.
41:13 You know, it seems like the problem here is that we have a really hard time measuring like what the business value per token is like. We've been maxing because there's an assumption that the value was greater than the cost of every token. Now we're mining because we actually believe the value is less than. But at the end of the day, the problem is like we just don't know how many like business points you get from a given token.
41:34 I don't know. Is that the right way of thinking about this? I have, yeah. It's. Yes. In some dimensions and no in one critical dimension. So I have I had a whole bunch of reactions to this, and my first reaction is, 100% token maxing is the wrong thing. And this token mining manifesto asserts that measuring outcomes is the most important thing, but it doesn't really explain how to do that.
42:02 Just like you said, we can't really measure business value in any meaningful way. Every business measures its value in different ways. There are so many degrees of freedom between a given line of code, a given token, and actual monetary value, or societal value, or just maintainability value or whatever, whatever value you want to call the output of that.
42:26 So it's really hard to measure. Another interesting aspect of this, I read an adjacent article by Steve Yegge called The Flat Curve Society that was focusing, its initial premise was talking about how, given Mythos class models, we're going to sort of see a flattening of capabilities because we're reaching the point where models won't be generally available to the public.
42:49 And that's an interesting premise, but the follow on from that was that the real work then becomes the educational journey of the majority. So right now we've been on this overall experience where everyone is pushing toward like we've got spiky capabilities. We've got like the people really running ahead in AI that can get 10x 100x done. And then we've got a very large majority of people that are sort of working their way up that ladder of capability and understanding.
43:14 And one of the things that Steve cited in his article was that up to the point of spending about 5 million tokens a day, token maxing actually is a decent predictor of overall output. But after that, and I think I think actually he said up to 15 million after that, it becomes a really bad predictor of the overall quality of output. And that's where you get gamification of the metric.
43:44 And so the idea of token mining, I think, becomes really important as the general populace reaches that level playing field of overall AI literacy, that everyone is now able to effectively use these tools, and now it's about effectively using them efficiently. So the last point on this one that really stuck out to me is that mining is obviously a very nice antonym to maxing, right?
44:11 But I actually don't think it's quite the right metaphor for what needs to be done here. I think the real metaphor is efficiency, and there's a critical element that is missing from this token mining manifesto, and that is that not all tokens are created equal. So the tokens that I can run on my laptop are, for all intents and purposes, free. I can use as many of those as I want.
44:35 I can max the heck out of those ways. Exactly. The tokens that I spend on a very inexpensive model are less burdensome to my overall budget and to the environment, and all of the negatives that we have associated with the cost of a token, both monetary and environmental. So back to that initial point about Fugu orchestration is going to play a key element here.
45:04 And I think one of the big untapped areas here is this idea of local offloading that happens to align with my personal interests, because I love models that I can run locally. But truthfully, in my daily flow, I spend almost half the tokens I spend are on local models. Now, granted, I have some pretty nice hardware here in my house, so that's not accessible to everyone.
45:23 But in the range of 30 to 100 billion parameter models, you can get a whole lot of work done. You can't get your most challenging frontier problems done. That's true. But you can get a huge amount of internet research, basic code implementation, exploration of your repository for, you know, summarization and information, like, all of that stuff is not your frontier of capability, and you don't need an expensive frontier model for that.
45:59 So, that's where I think the story really needs to drive is that efficiency of output. And that has both a use less and a use more, aspect to that overall equation. Yeah. There's so many variables here. It's like almost like what tokens, where the tokens are. It's a big, I think gawky math equation. It's awfully hard to encapsulate that in a pithy phrase like token max min.
46:24 Exactly. Well, max, under certain conditions, which you specify. Exactly, exactly. Yeah, exactly. I think one aspect I mean, Gabe, you mentioned that kind of interesting like curve where it's like, okay, up to this many tokens, it's a good predictor after this many tokens, it's a bad predictor. And it strikes me that that's actually a really interesting number because it almost reflects to your point how efficiently people can use this technology.
46:52 So you imagine a world where, like everybody's using AI for the first time and they're just trying to like, use it to figure out the problem. Like that predictor threshold is like way higher, right? Because the average person takes like just a lot longer to use the technology to solve the problem. But then over time, you actually want that number to slowly kind of go down, because your hope is that people get better at using the technology they can, I don't know, they prompt better and so they get the answer more
47:12 quickly. And so your token consumption like decreases over time. I guess, Martin, just to turn this into a question a little bit, it's almost a little bit like, you know, there's almost going to be like token deflation is what you want to see as people kind of get better and better because it's less about the tokens. It seems like it's actually more about the, the person, right, and how they use the technology.
47:34 Yeah. What this reminds me of is the laser printer at my first office. When I got my first job, they had this really fancy laser printer. It could print like 50 pages a minute or something, and we had full access to this thing, and we were encouraged any time that we had like a PDF or manual or something that we wanted printed, just send it to this fancy printer and it would like, spew this thing out in seconds.
48:00 They had folders and you could put them in the ring binder. Yeah. And then at everybody's cubicle, they actually had shelves set aside so you could stack them up. You could basically make your own personal library. So that was print maxing. That was print maxing because who wouldn't want the biggest fancy looking library in their cubicle to show how busy and important and productive they were, right?
48:23 Well, I left that company and I moved to a startup and this was a tiny startup. There were six of us in one room above the harbormaster's office at a Marina, and I get there my first week and I come across this PDF with this this Java manual. I'm like, oh, that might be useful at some point. So I send the the job to the little laser printer we had in the office there.
48:46 And within like two minutes, the founder is coming over to my desk and he's like, dude, we have a policy with printing here. You can print as much as you want, but you've got to read every word. And I'm just sat there looking across at the printer as it's spewing out page after page of this Java manual. And I'm like, I am at most going to read three pages of this.
49:07 I just don't know which three pages. So I printed the whole thing. Print mining, right? Like, the fewer the better. So I think that idea of just measuring productivity by I printed this whole bunch of manuals. Look how productive I am, or I've hardly printed anything. Look how productive I am is really the wrong metric. It's a very easy metric to measure, right?
49:32 I mean, that's that's the thing that tokens are a number. We can measure them. We can compare them against other people. But as as Gabe and Tim, as you've both pointed out, you know, a printer page is a printer page, but a token is well, it could be a token from a model that's running locally in my laptop. It could be a pretty efficient token from a, from a sort of small 100 billion parameter model.
49:57 Or it could be something from a frontier model that we are paying $0.50 for every million tokens or something. So it could get real expensive real fast. So yeah, just just the raw number is is not a good not a good measurement. But, you know, I still miss that laser printer. That thing was so fast. Yeah. I think it's like, you're revealing the absurdity of it because, like, no one's like, bro, you got to be electricity maxing, you know?
50:21 Or, like, you have to be, you know, like, why are you not water maxing? You know, I think, again, this is probably like the process of becoming more mature about how we manage and measure, these technologies. A terrific discussion, like always. Abraham, Martin, Gabe, thanks for joining us on the show today, and thanks to Huiming and Sascha for the special segment.
50:41 That's all the time that we have. And thanks for joining all you listeners. If you enjoyed what you heard, you can get us on Apple Podcasts, Spotify and podcast platforms everywhere. And we'll see you all next week on Mixture of Experts.
Orchestrate models behind a single endpoint to provide stronger capabilities on smartphones or commodity laptops than a single local model might provide.
Develop filmmaking tools that assist storyboarding, production, and related workflows while preserving creative control.
A platform routes requests across multiple existing models and can change the models behind the endpoint over time.
Google pays a film studio to use its tools and provide understanding of filmmaking use cases.
This chip packs nearly 100 billion transistors into something about the size of a fingernail.
The real work then becomes the educational journey of the majority.
VP of Silicon Technology Research and Development interviewed about the sub-1-nanometer chip technology.
01:03Director described as having partnered with an AI firm to build an AI storyboarding tool.
03:15Artist mentioned in the context of protecting image, copyright, or intellectual property.
03:18Person addressed in a hypothetical invitation to help develop AI for brewing coffee.
03:36New York Times writer associated with an article about minimizing AI use and token mining.
03:56Company developing the nano-stack sub-1-nanometer chip technology and Granite-related AI work.
00:59Japanese AI lab associated with the Fugu agentic model and multi-model orchestration platform.
14:24AI model lab whose architecture influenced GLM 5.2 and whose models were mentioned as examples of Chinese open-model progress.
24:30Enterprise mentioned as restricting or limiting access to APIs or cloud models because of usage costs.
39:59Enterprise mentioned as restricting or limiting access to APIs or cloud models because of usage costs.
39:59Publication associated with an article about minimizing AI use and token mining.
03:56