← All transcripts

IBM’s cloud collab, Meta’s Muse Glimmer & OpenAI’s upcoming Astra model Transcript, AI Summary & Key Points

IBM Technology · yesterday · Education · 36:33 · EN

🧠 AI Summary

AI is moving from specialized hardware and experimental systems toward industrial-scale infrastructure, with IBM Cloud, Together AI and NVIDIA partnering around large AI compute deployments. Meta's Muse Glimmer demonstrates that capable, dense 30 billion parameter agentic models can run locally on laptops, supporting coding and tool calling while improving privacy and reducing cloud dependence. OpenAI's upcoming Astra model raises serious cybersecurity concerns because frontier models may be able to compromise systems, forcing AI labs to strengthen training and deployment procedures. The panel argues that open models improve innovation, competition and defensive readiness, despite their security risks, and predicts that open models will eventually overtake closed models.

🔑 Key Points

  • IBM is partnering with Together AI and NVIDIA to make the B300 generation of chips accessible to the public through IBM Cloud.
  • Industrial-scale AI infrastructure requires high-speed networking, redundancy, upgradeability, power, cooling, storage and large-scale deployment expertise.
  • Meta's Muse Glimmer is a dense 30 billion parameter open agentic model designed to run on laptops and support research and tool calling.
  • Muse Glimmer ran at around 30–40 tokens per second on a Mac M3 and around 10 tokens per second on an M2 in the panelist's testing.
  • Muse Glimmer uses speculative decoding, quantization and context-management techniques to improve local performance on approximately 24-gigabyte devices.
  • On-device AI is suited to privacy-sensitive or disconnected use cases, while larger mixture-of-experts models will generally remain hosted in industrial-scale cloud infrastructure.
  • OpenAI's Astra has not been released because of concerns related to critical cybersecurity capabilities and the difficulty of controlling systems that can compromise other systems.
  • The panel argues that open models are necessary for broad access, competition, innovation and the ability to defend against advanced AI-enabled threats.

✅ Actionable items

  • Use redundancy so that an individual node failure does not bring down an AI system.
  • Design AI infrastructure for upgradeability as model architectures change.
  • Use batch processing to backfill load fluctuations and improve infrastructure utilization.
  • Load-shift workloads between data centers where data-sovereignty requirements permit it.
  • Run smaller, specialized models for agentic workloads when cost control is important, using larger models for advanced reasoning.
  • Use model routers to decide which model should handle a task.
  • Keep privacy-sensitive or disconnected workloads on devices or within infrastructure that maintains end-to-end control of the data flow.
  • Update cybersecurity procedures for AI model training and deployment after unexpected model capabilities or incidents.

🤖 AI in practice

Used for

Run a small language model locally on a laptop for research, tool calling, and agentic workloads. 13:13
Accelerate local language-model generation by drafting multiple possible next-token blocks in parallel and checking them with the main model. 14:20
Use smaller specialized models for routine agentic workloads while reserving larger models for advanced reasoning. 20:29
Backfill fluctuating AI-compute capacity by scheduling noninteractive workloads as batch jobs. 11:08

Advice

  • Use open-weight models and run appropriately sized models locally where privacy, disconnected operation, or regulatory requirements matter. for An enterprise buyer, developer, or organization handling sensitive data.
    Local models avoid moving sensitive data and can reduce dependence on data centers, while open models support research, competition, and defensive capability.
  • Treat agentic development as part of the coding and security threat model, and use top-of-the-line agents when defending systems against others using them. for A developer or cybersecurity team.
    Agentic workloads can run on laptops and may exploit sandbox vulnerabilities, so organizations need to account for their capabilities.
  • Right-size AI workloads across local devices and cloud infrastructure instead of choosing exclusively between the two. for An enterprise architect or AI platform team.
    Dense models are better suited to smaller devices, while larger or mixture-of-experts models generally require industrial-scale hosted infrastructure.
  • Update cybersecurity procedures for training models after a model compromises another system, and share the lessons so other organizations can improve their procedures. for An AI lab or cybersecurity organization.
    The speakers describe model-assisted compromise as an unprecedented capability that exposes weaknesses in existing defensive preparation.

What it could not do

  • Muse Glimmer runs substantially slower on older or less capable laptop hardware. — The speaker reported around 30–40 tokens per second on a Mac M3 but around 10 tokens per second on a Mac M2.
  • Mixture-of-experts models can be inefficient on small devices because they may require enough memory to hold many experts and can otherwise involve swapping and waiting. — The discussion contrasts dense models for smaller devices with mixture-of-experts models hosted in the cloud.
  • Very large models cannot be run on embedded devices. — The panel says larger models need industrial-scale hosted infrastructure, while only selected workloads should run on-device.
  • Load shifting can be difficult for highly regulated industries. — Data-sovereignty requirements may prevent workloads from moving freely between data centers.
  • Some AI systems are difficult to control and are highly capable at getting online and compromising systems. — This is presented as a central reason for concerns about releasing OpenAI's Astra model.
  • Astra was not released because of cybersecurity concerns. — OpenAI described Astra as a model in development and said it was not being released at that time.
  • The speakers characterize Claude as having been restricted to the point that it can no longer perform cybersecurity attacks, while also becoming less useful for meaningful tasks. — This is offered as an example of the tradeoff between reducing cyber capability and preserving model usefulness.

🧰 Tools & AI usage

AI is used for

  • Coding, research and tool calling — Muse Glimmer is designed to support agentic workloads on a laptop.15:30
  • Cybersecurity attack and defense — Advanced AI systems may be able to compromise systems, while defenders can use comparable models to improve security.27:29
  • Training, inference, backtesting and reinforcement learning — AI infrastructure can shift capacity among these workloads to improve utilization and economics.09:06

📄 Transcript

Searchable transcript of IBM’s cloud collab, Meta’s Muse Glimmer & OpenAI’s upcoming Astra model — IBM Technology (36:33). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:01 The best way of solving this is openness always the best way forward. I I And then I think at some point, and I'm going to call it, and I've called it before, there will be a point where open will overtake the closed models. That will [music] happen at some point, right? It's just a matter of time. And when that happens, we're going to we're going to have some fun.

00:21 >> All that and [snorts] more on today's mixture of experts. >> [music] >> I'm Tim Huang, and welcome to Mixture of Experts. Each week, MOE brings together some of the best brains in artificial intelligence to lead you through the week's news. On this week's episode, we've got Chris Hay, [music] distinguished engineer, Rin Witna, technical lead AI ecosystem, and Volkmar Uhlig, CTO and VP of data platforms.

00:45 We're going to cover three big stories today. I really want to talk about Muse glimmer, Meta's new model. We'll talk a little bit about Astra, OpenAI's new model. But first, I want to start by talking about a big new partnership that's just being announced between Nvidia and IBM. >> [snorts] >> IBM's going to be teaming up with Together AI and Nvidia to launch and make accessible the B300 generation of chips to the public.

01:13 And so, Volkmar, I know you've been pretty close to this. This is a pretty big story because as I take it, you know, IBM will basically be in the game on providing sort of this next-generation cutting edge for compute. Do you want to take us a little bit behind the scenes? Like, what does it take to have this sort of thing happen? And I guess why Together as a partner here?

01:32 >> So, Together is one of the bigger AI vendors providing training and influencing capacity. And those systems are very differently designed from traditional cloud computing environments, primarily because they need networks which are super high speed. Um do you have a reliability issue? So, there's a lot of built-up around redundancy so that if individual node fails, then the overall system stays alive.

01:59 Uh you need upgradeability, etc. But then also there's a big build-out on power infrastructure, cooling infrastructure, and overall network capacity of storage. So, we are building out these systems in IBM Cloud. We've been having them for internal use cases, or small-scale use cases. Internally, we built pretty large training clusters, and now we are entering the market and we are having a partnership with Together AI, which brings this to the to the end customer and to to um providers in that race.

02:29 So, I think what uh where I see this moving to is um we are in a transition right now from, you know, AI is kind of this esoteric thing to industrial-scale computing. And industrial-scale computing requires these industrial-scale deployments. And we're seeing this across all the hyperscalers. We are going from a phase of like, oh, this is this esoteric piece of hardware and kind of cool tool to know, now we need to go and get into the cost optimization.

02:56 This is human infrastructure, right? And so, I think there will be a few more players in that game, but it's a very capital-intensive um like industry to be in. And it's also uh technologically hard to do this at scale. Putting, you know, eight GPUs under your desk is is a different problem than putting 10,000 GPUs into a data center. And so, this is where our the work with Together AI and, you know, we're building on years or a decade almost of uh AI developments for internal purposes.

03:28 And, you know, now we we have this accessible at very large scale with Together AI. Together AI. >> And I think um industrial scale, I think is the thing I want to pick up a little bit on. You know, Chris, um you know, over the last few years we've seen the emergence of these so-called neo clouds, right? That that basically come to the market saying, we know how to do this specialized F1, you know, compute development.

03:51 And this is why you should go with us and you shouldn't work with those old school clouds because they just don't know how to design these kind of special types of clusters. And I guess Chris, I guess maybe the question for you is how long you think that's going to be the state of play? Because I think to Volckmar's point, it does feel like okay, now now the really big boys are moving in, right?

04:08 The clouds are going to really be kind of doing this at mega scale and are building a lot of competence in it very quickly. You know, are neo clouds going to be able to kind of survive in this environment? Is there going to be a niche for kind of specialized AI cloud providers or is the future of AI cloud just kind of, you know, the the cloud? >> [laughter] >> If I look a little future forward and I look at what everybody is doing, I think there is space and it's probably not for the reasons you think, right?

04:32 Because if we think of the the large industrial cloud providers, they're investing for a very long time, right? You know, you you're you're going to have these GPUs going, model architectures are going to change and therefore you're going to be have to be able to handle the new model architectures as they come. might be able to run that at speed and scale and be able to have that running for the large enterprises and be consistent.

04:54 The the neo clouds I think are under less constraints in that sense. They can be a little bit more innovative and they can go a little bit faster and they can have more agile architectures. So I think that I I think there's still a place there and and the two are going to sort of coincide. What I think is probably going to get super interesting is we're already seeing that most of the AI providers, their demos these days when they release a model is, "Look, I designed an AI chip."

05:25 Do you know what I mean? Now I >> [laughter] >> Now I don't I personally I I don't think that's a coincidence at all. I think that's something that everybody's thinking about is like, "How can I use AI to design more efficient chips and then be able to run my own chips, especially for things like inference?" So um I I probably think that actually there's a lot of diversity that's going to come in the future.

05:50 >> Yeah, it's almost this really interesting question between basically like do the do the models have the leverage here or does like the kind of compute and infrastructure sort of have the leverage? You know, Ren, I feel like there is this kind of like dream of vertical integration that certainly I think some of the frontier AI labs have, which is okay, well, you know, we're going to just build our own infrastructure.

06:08 We're going to design our own chips and it's going to be specialized for our models and and life is going to be great. I also remember I mean back in the day I was shocked to learn that uh Netflix ran on AWS for a very long time, which is like a lot a lot of bandwidth um all provided through a third party. Um I guess on that part of the market, you know, not even talking about neo clouds, like if you were an Anthropic or an OpenAI or one of these frontier labs, do you think the idea of kind of vertical integration here

06:34 is probably something that you get to at some point or, you know, again, just in practical terms, you will have to offload this to, you know, some other cloud provider? >> I think it's actually both. Uh and the reason I think it's both because um I I think a lot of the chip plays right now are a diversification play uh where you you need to be able to show that you're growing, especially as you're nearing IPO, and especially as you look into going into the future.

06:59 Uh and with how much difficulty we've had getting RAM and things like that over the last couple of years, um you have to have your own supply chain if you're going to preserve yourself from some of those risks. So, I think that's been a lot of the play. Um but of course also efficiency of data centers is a major cost of advantage. And if you can have something that can run your models at some improvement to that, it's going to result in more profit.

07:28 >> Even if you design, um you know, a chip and you are suddenly figuring out, oh, this was not an optimal design, those hardware investments, if you don't outsource it to a third party, it is on your books. And so, I think what you're doing as at some point it's a chicken-egg problem, right? So, at the beginning you're like, "Okay, you're designing your model for a specific piece of hardware."

07:50 Then you're adjusting the hardware to your model, and then at some point you have this these assets on the books, and you will have to run them out for, you know, 5 years, otherwise you cannot recoup your investment. And then, you will actually design your models for your fleet. And this can happen when you control the supply chain, not if Amazon controls the supply chain for you, right?

08:10 Um and then, the other thing is Amazon, you know, they're they're running currently at 43% profit margin if you look at AWS. And so, I don't think that if you look at the token economy, that the combination of infrastructure running and model running allows in the long term, not in the short term, in the long term will allow that Amazon takes a 43% profit margin, right?

08:34 Now, the the reason why in many cases they do that is because elasticity is [snorts] allowing them to, you know, like the the the users of this infrastructure allows them to scale up and down and shift the demand and the elastic demand curve onto the cloud providers, but the cloud providers therefore make 43% with this elasticity. And so, I think we are getting into a point where the consolidated factory of AI is, while I use my machines to do one inferencing, I may use them for training.

09:08 And training doesn't necessarily mean like, "Hey, I'm I'm I'm fine-tuning a model." But now with reinforcement learning, a lot of the training is actually inferencing. And so, I can use my inferencing capacity for training phases, etc., for backtesting, yada yada. And so, um I think having the integrated factory and owning these assets and being able to shift them actually makes economic sense because you have a baseline load.

09:33 And so, the the elasticity in Amazon gives you, if you're are at that scale, it's just not beneficial anymore. So, at small scale, it is, but at large scale, you could just run your own much more efficiently. And since it's such a specific use case, you're running 99% of your workload good enough, right? So, and then you use Amazon for everything else, or Azure or GCP or IBM.

09:56 Yeah. >> Yeah, I think this is a really interesting question. And maybe Volckmar, if you want to have a final thought on this is, you know, there's a really interesting question here just basically like how spiky do you think the median enterprise users consumption of tokens is over time? Because it totally influences the kind of shape of the cluster.

10:12 I guess in your experience, what I mean, what are you seeing in terms of like industry trends? And obviously it varies, but yeah. >> Yeah, what we saw is that, you know, I was running Watson X AI, and so we saw these these load shifts. So, you kind of have a base load which is global. And then in in an individual zone, you have about 3x difference between day and night cycle.

10:36 And so, it's quite substantial. So, you know, in in the fluctuation. Now, on the flip side, if you look at the worldwide load, it's actually much more the fluctuation is about 20-30%. And so, if you can if you can load shift between the different data centers, you can actually even it out. So, now there is a problem which, you know, when you're in highly regulated industries like where IBM's customers are, sometimes load shifting is really really tricky, simply because, you know, you have data sovereignty requirements.

11:10 And so, what what you see now is all these providers, they are doing the load shifting through batch processing. And so, what you're doing is you're effectively taking the interactive workloads during the daytimes, and if you look at all the guarantees you get for batch processing, it's usually 24 hours. So, you know, you submit your batch job. And so, what you're doing is you're using the batch jobs to backfill the load fluctuations and you just time offset them and that's how people, you know, kind of even out there

11:36 are that to be utilization. >> Well, Graham, I'm going to move us on to our next topic. >> [music] >> It's kind of been this last week sort of the the return of the Zuck. You know, Meta's been quiet for some time. They had a huge new cycle about all the talent they were recruiting and then kind of just there were sort of rumors swirling for a bit and we are starting to see kind of a a bunch of announcements come out of Meta.

12:00 It feels like they're really kind of cranking, you know, the the press engine on their releases. Um and this past week there's two things that happens. Uh first one was the launch of a new model. Um so Meta has launched something they call Muse Glimmer, which is uh what they bill as kind of an open agentic model that runs on device. And then separately, um there's been kind of a vision piece, you know, a little bit like Sam Altman's like the age of superintelligence.

12:27 Uh uh Mark Zuckerberg has also released his own sort of essay called The Future Is for Everybody uh about kind of the desirability of distributing superintelligence to everybody. Um and I guess Chris, I know when I mentioned that we were going to cover this story, you had a big grin and you were like, "Oh, I've got some things to say about this." Um so how about I just let you kind of sound off?

12:45 I'm curious about, you know, how you received sort of the the post and and I guess what you think about the model. What's your what's your taste test on the model? >> So the first thing is I'd like to welcome Mark Zuckerberg to uh who's clearly a listener to this podcast cuz I don't know [laughter] if you remember this 2 weeks ago or three weeks however long ago it was then I did say we talked about Spark and I was like, "Ah, I don't care."

13:07 It's like open up your models, Zuckerberg. You know, nobody cares if But uh the good news is he obviously listens and he was like, "Yeah, I need to open the model up." and he did and suddenly >> That crazy guy. Really sharp guy. [laughter] >> The new model actually on the Glimmer model, let's talk about that for a second, is great. It is actually great.

13:25 There is um if you if you haven't tried it, um go down that I I honestly think it's one of the best small models that you can run. It's a 30 billion parameter model. It's a dense model, which is nice to see something that's not mixture of experts for a while. They're They're They're clearly um competing with um Google. Though um it's really nice to see architecturally they're taking a lot of the insights from like the the Gemma models and they're taking insights from um the Kim K3 models, for example.

13:55 It runs beautifully fast. I was running it on my Mac M3 and it was running um it was running at around 30 40 tokens per second, which was pretty reasonable. Um and again, I think the on on my M2 it's a a lot slower, it's a 10 tokens per second. But actually it it is a great model and and they've brought a lot of fantastic insights. So they're doing a speculative decoding with their with their flash um drafter, um which is basically they essentially predict a N blocks, you know, what the next tokens of the models are in

14:32 parallel and then they check it back with the model going, "Oh, is this correct?" So there's a little mini model doing that drafting. If it's correct, then they accept it and they move on, and which is why they're getting speed. They've clearly optimized the model down for the sort of 24 gig, you know, like um um essentially your your smart boxes, etc.

14:49 So they've they've really worked on worked on that. They provided the quantization levels. They've They've really thought about it from a model architecture perspective. And then they've spent a lot of time on um the context as well, actually. So they they're they're sharing Google's techniques, so they um but I'm not going to go too much into the technical details of this, but but effectively um when you're looking at the residual stream, I'm not going to go into the details of that, but but basically they're spending

15:19 most of their time looking at small contacts the last 2000, but every so often maybe I think it's every fifth block or so they are then looking at the contacts as a whole. So they're getting the benefit of being kind of looking hyper focused at that your last set of tokens, but every so often just keeping a an eye on it. And there are techniques that were have been used in the Kimi models, there techniques have been used in the Gemma models for example, which they're clearly competing with, but but it turns out what

15:47 you're getting is a super fast super small model that runs on your laptop, that's optimized and designed for your laptop, and it's clearly designed for a gen tech, right? They focused on research, they focused on tool calling. It is honestly a great model. And again, if you compare what I said a few weeks ago, where I didn't care, now I care. Well done Mark Zuckerberg.

16:08 [laughter] Go write as many essays as you want as long as you keep releasing models. >> Instant meta fan. Red, I I I think I want to take it like this discussion maybe on just a tour of recent history, right? Because if you recall meta used to be really leading the way on open, right? I remember every llama release being like, "Oh man, the new llama release.

16:27 This is huge." It's like you know, and then meta kind of faded a little bit and then I think the narrative was, "Okay, well these kind of Chinese labs are really leading in sort of the open space." I I guess the question is like whether or not we're like kind of overly panicked about sort of like, you know, the Chinese lead in open, right? Because I think we've got Gemma now, we've got glimmer.

16:44 These are really good open models. And I guess the kind of question is like, you know, I guess maybe two things. One of them is this kind of like, does this put sort of meta back in the running, but also maybe the question of like actually maybe like US labs are really good at open as well. It is I think what I'm reading for an announcement like this.

17:02 >> Yeah, I mean I'm going to throw one more in the mix. This has been IBM's focus around the Granite models, right? So generally a big fan of this approach, right? Where you know, the more we can get in the open, the faster we're going to innovate, the faster we're going to see things changing. And that is of course, as always, a double-edged sword, right?

17:22 Where there is now the ability to run these agentic workloads on your laptop, and that's something we can't undo, right? That's now just baked into the future, right? So, as we look at all of these models that are breaking containment and things like that as they start having security vulnerabilities in their sandboxes and such. Um I I I I think this more than anything solidifies agentic development as a requirement, where you now need to be able to use top-of-the-line agents as part of your coding because everybody

17:55 else has them as part of their threat model. So, all of those things come together to say that, you know, yes, open-weight models are the solution to that. That's how we get research. That's how we continue developing. That's how we push the field forward. Running them on your laptop also critical, right? Like that's how you keep costs down. That's how we keep from everything being in a data center.

18:20 And that's how we democratize the tech. But but everything has a shadow side, right? >> Robert, do you want to talk a little bit maybe about the kind of compute and infrastructure aspects of this? Cuz obviously like the big headline that's trying to link to these stories is that Glimmer sold as a model that runs on your device. And Zuck of course is sort of selling this vision that, you know, super intelligence is going to be sort of everywhere and for everybody.

18:45 And you know, it's actually kind of funny in some ways, right? Like I think we were just talking a few weeks ago about Kimmy K3. We were like, who is running Kimmy K3 just given it's like absolute size. And so, there's kind of this very interesting thing where it feels like these kind of American companies that are in the open space are kind of playing for the smaller models.

19:01 There's also an idea that like a lot of workload is going to be taken up on device versus in the data center. And you know, again, I don't think it's an either or it's not like a black and white thing, but curious about how you think about the division of like how things are evolving on what you're going to do on your device versus in say a big data center.

19:17 >> The overall the models, I'm putting all the dense models will be more on smaller devices, right? This is small phone factor thing. Model of fact mixture of expert models have this the kind of demand that you have a lot of concurrent requests coming through because otherwise you are you know you are inefficient like infrastructural wise you're inefficient.

19:39 Many cases you you have to swap and wait because you know you just don't have enough memory to hold all these experts in in in memory, right? And so I think we are going down a road of there will be on device AI we will right size it. There'll be on device AI primarily because A you want privacy, B in certain cases you're disconnected and then you have the larger models which are in in the cloud.

20:08 And they are industrial scale hosted. That's how how I think about this. Like if you have a mixture of expert model, you as a as an individual or as a small scale company, you just cannot afford putting the infrastructure down and so you will rent the infrastructure and so you will then you know go to some of these model providers and they will just host it for you.

20:26 I think that if you look at agentic workload cases, there will be a very hard push and this is what we're seeing you know even internally with with Bob. There will be a push towards specialized smaller models just to keep the cost in check. You just cannot afford the very very large models. So we use them for advanced reasoning and then we will have model routers which effectively make a decision where to go.

20:53 I think the the economic the economics are in in a way that if you are large scale enough, then you can actually specialize, but you may with with common infrastructure and this fluctuating workloads, you may have enough spare capacity anyway that you can just run the big model out. And so I think it's a balancing act of the of the model providers when they are choose and and the harnesses when they choose what.

21:20 And I think this will be where the smaller models will play. Wherever you have the choice of actually putting your your per token cost into some reasonable amount or we will get good highly specialized or more specialized models or if the task is is easier and the the complex models will effectively spit this out at some point to which model you could potentially go just so that it's part of the integrated story.

21:44 So I think this is where how how I see the the market fall out. You cannot run these huge models on embedded devices but certain stuff you want to run on embedded device simply you don't want to move the data. So overall I think there there is a there's a privacy argument and and putting it on the enterprise the privacy argument is very strong. Just from the perspective of like regulatory requirements.

22:04 Like you just cannot move your data if you don't control the end-to-end flow of of the bytes. >> Yeah, I think it'll be kind of in some ways. I mean the privacy point is something I want to pick up on Ren was basically that you know normally we're like oh consumers are the ones who are really concerned about privacy here but it almost kind of feels like in the last decade people were like ah they've sort of like loosened up a little bit on that being something that like they're they're really really hung up about.

22:28 At the same time it does feel like companies are super super worried about um uh you know particularly these AI companies being able to kind of train on their data and and all this, right? I guess Ren maybe the question for you is is you know Zuckerberg's messaging very much kind of focuses on the privacy benefits of having super intelligence that's kind of personal and on your device.

22:51 Do you think it really matters to consumers so much? Like will that become a defining factor in what you decide to adopt in terms of an AI service? >> I think it matters and I think the reason why we've seen that kind of change in terms of people seeming to care less about it has been, you know, the boiling frog scenario, right? Where there hasn't been an option to get access to the services that you want to have without giving up a whole bunch of privacy.

23:17 Where you know, you have to you know, terms and conditions on everything you sign up for where you're paying for it with your data, right? So, I think this is actually a really interesting for a certain type of consumer. And there will also be a certain type of consumer who does not care enough to set it up on their laptop or doesn't have a powerful enough machine and there will continue to be hosted things for that.

23:40 It's a very much a balancing act where I I I think it is very nuanced here. And while yes, I personally care a lot about privacy, there's a limit to how much you can actually do about that these days, which is frustrating. Really a great discussion. More to come, I'm sure. I think we're probably going to have another set of meta stories next week as they continue to kind of like gin the works and put more stuff out here.

24:04 But I just think it's yeah, it's a really really interesting kind of intersection of issues. >> [snorts] [music] >> I'm going to move us on to our last topic of the day. It's kind of a funny news story because it's almost not a news story at all. So, OpenAI put out a blog post entitled Responding to the Next Frontier of Critical Cyber Capabilities. And Chris, it's kind of a funny story because what they're sort of the substance of the post is we're working on a really great model called Astra.

24:33 It's not released today because of concerns that we have. We'll be in touch. And so, I don't know. I mostly just wanted to to flag it because I think it's kind of a funny story and it does kind of feel like we're entering this really strange situation where you know, frontier AI labs get a lot of juice out of announcing that they're not doing things.

24:54 And and what to make of that? I mean, I don't know. At this point, it's kind of like, yeah, just let me know when Astro's available. I don't I don't really want to know the details on the way. Um how did you respond, I guess, to this story? >> I think if you had a large news story or your latest model in training hacked a competitor's or a partner's uh system, and perhaps you did a whole YouTube video on it, which was great, detailing exactly how this worked, I think you probably would update your >> your cybersecurity

25:25 procedure for training models. I I think that's a sensible thing to do. >> Yeah, you're actually sympathetic. You're like, actually, they're doing this for good reason. >> [laughter] >> I I I actually think so. I I think I think it was so unprecedented what happened. We knew it was coming at some point. We knew this day was going to come. Um but I think you until it happens, you don't really realize the implications, and therefore that forces you to go, "Well, actually, we had this procedure that existed before, you

25:54 know, these things happened, and you know, we've got to improve, and we've got to hope that other people learn from these lessons, and uh you know, and and can take the right precautions in the right way." And then actually, if you watch the the kind of the Black Hat YouTube video, they made a really good point on this, which is like uh for the attackers, do you know what I mean?

26:16 This day has been coming, but he was they were basically saying, "Look, I don't think we have an example for large-scale um cybersecurity defenders where they're properly prepared in that way." And I think it was a really interesting point. So, so I I think you have to update the procedures, and you have to say, "Well, these are the things that we've learned."

26:35 And then hopefully companies of the world will learn from that, and then actually update their own procedures, because I think the reality is as we move into cybersecurity, um the speed and scale that Agility works at is is really the key thing, and you have to sort of protect that. Now, you you you probably have to say there, well, you know, surely some of this stuff would have been obvious there, but hindsight's a great thing, right?

26:58 Do you know what I mean? It's like you you know, I don't think anybody was expecting the models to be quite at this level of capability, um including OpenAI, and you know, and I think they've been massively transparent about it and really open in sharing everything that's happened, and and I I you know, and I I'm sure people will disagree with me, but I I think it would have been easy for them not to be so transparent.

27:21 So, yeah, I I I I sympathize and I'm I'm glad that that they've been open about it. >> Um look, I guess the question, you know, I guess Chris is bringing me around. I was making fun of them a little bit, but you know, we we do seem to be have a real genuine cybersecurity issue on our hands. Maybe the other way to flip the question is is is Astra ever going to come out?

27:41 Cuz it kind of feels like there's like two really big problems here. One of them is uh they don't seem to be able to really control the behavior of these systems. They're just really good at getting online. They're really good at compromising systems. The other bit is, Chris, to your point, right? Like it's really hard to get defenders to up their security robustness.

28:00 And so, I guess if your threshold is, we need to wait until the world is safe enough to release this model, I almost had to take a look at that and it's like, is that ever going to happen? Do they really have a handle on this issue? Or will Or do you think it will be really kind of like solved in a substantive way in any reasonable amount of time? >> Uh that's a loaded question.

28:15 So, [laughter] So, I think like let's look at the players here, right? So, there is the the cybersecurity hackers, uh there is the um the the people who are trying to protect themselves, and then there are governmental actors who you know, have used this for warfare. And then you have industrial the whole industrial espionage thing, right? And so, the question is, who gets their hands on the model first?

28:43 And so, what do want to do is, you want to keep your you want to keep your society safe, whatever that means. And that may mean that, you know, you want actually your cybersecurity department of your military to actually have access to the model, right? And so I think it's a it's a phasing where you know, probably those AI lab vendors are working with the respective agencies etc.

29:11 But they're saying, okay, how do we phase it and so that we got a competitive advantage for a while, you know, before all the gaps are closed without bringing down our society knowing that, you know, adversarials will actually do the same thing and they also train these models and they also deploy these models. So from my perspective, we are just in the cybersecurity and cyber warfare.

29:31 And so you work with the ones where you have the biggest exposure. So that's your, you know, your iOS, your Windows, Linux, etc. And you trying to close the holes, but you trying not to close all the holes cuz you also have an advantage of not closing all the holes, right? And so I think this is it is a cat and mouse game and we just offloaded a cat and mouse game from human brains which have PhDs from MIT into effectively the AI labs.

30:00 And so the AI labs are negotiating. So that's one that's one view on this. The other view on this is like, wow, what an amazing story. My stuff is so smart that it's dangerous for the world. Like that's the the New York Times headline, right? And I think it creates a lot of anticipation and and then sometimes, you know, it just becomes I mean, if you look at at Claude, you know, it's like it feels lobotomized what they did to it.

30:27 So yeah, it can for sure not do cybersecurity attacks anymore because it's so dumb that you actually don't want to use it for anything meaningful. Right? Like I mean, it's the the fable the fable five, you know, lobotomization. I I the model by now in in some cases is literally just unusable. >> I have definitely run into a number of things where it's like the behavior is definitely different and and less good.

30:49 Um Yeah, I guess Ren, do you do you I mean, to pick up on that? Like there is a narrative which is, "Hey, isn't this all marketing?" And it is kind of funny. You don't normally see situations where companies are kind of like rewarded from a marketing standpoint for limiting the functionality of their model. But I also kind of believe that I don't know, maybe our enterprise customers really excited because they hear about this like incredibly dangerous technology you can't release.

31:12 Um I guess how much credence do you give sort of more of the kind of marketing interpretation of this? Or or do you think that there is really genuinely some really serious risks that, you know, I think they're they're just trying to do their best to handle it. >> So, I think it's both. Um I I do want to just remember that uh 9 months ago we all wrote code by hand.

31:33 That was 9 months ago, right? Like like the things have evolved so rapidly and so much faster than I think any of us anticipated. And that's very true. Where we now see that uh all of this stuff is moving faster and faster and faster and faster and faster. And it's both good. Uh we can do a lot with these things and also really scary in some ways, too.

31:52 Um and it's important to remember all of those things as kind of looking at the the the space in which this is operating, right? Like many organizations can't get a deploy out in 9 months, right? Like it it it can be very difficult for some companies to react to a security posture change that quickly. Um and I think that is the danger side. That is a fascinating place for us to be in, right?

32:21 And I do think that there is some aspect of uh you know, we saw Fable get lobotomized because they overhyped it, right? Because they really hyped up the Mythos class and then they kind of had to uh from a a perspective. Um and I I also don't know how much of this is people trying to, you know, as a seated player trying to get more regulation to preserve that moat as well.

32:47 >> Chris, maybe we'll give you the last word here before we bring this episode to a close. I think one element of this that I keep coming back to my mind is uh the open ecosystem, right? I think if if anything, I think like the foundation the frontier models are like becoming more and more like, oh well, we have to be careful about every new release.

33:02 Uh meanwhile, open just keeps coming out with more and more and more and more um and the capabilities of even small dense models as we talked about are getting really, really good. Um how do you think about the security aspects of all that? Like, do you think the open providers will have to themselves, you know, pretty soon say, well, due to the cybersecurity issues, we're we're also not releasing?

33:23 Or does the incentives there just really kind of prevent them from from taking those types of actions? >> Without quoting my new hero, Mark Zuckerberg, um >> [laughter] >> it has to be open. I just I don't see how close works. We need to put if for superintelligence to be truly uh you know, beneficial for everybody, it can't be concentrated into a few closed providers.

33:48 That was Zuckerberg's words, so there we go. But and and I actually really agree with that. I'm I think that the more um the AI shared across the world, the more that we share compute, the more that uh we're open with the models, the better chance that we are going to have to be able to defend, right? And and and the reality is if you've got an actor over here who is got access to these closed models cuz they're they're special, but then these people over here don't have access to the same class of models, then they're

34:19 not going to be able to defend themselves. And they're not even going to know how to, right? And and I think we we can't have this disparity, right? And again, without open, then the closed model prices are going to go up. And we've seen that already and you know, everybody remember a few months ago we were talking about API token pricing and all of that's disappeared because guess what?

34:39 All the open models have become really capable again and and the closed model providers are going, oh wait a minute, these people are just going to leave, right? And and so we can't have that. We we need to have an open ecosystem. It is going to be safer. It is going to be more secure. But but you know what? It's it's it it will figure it out. And and and I think Hugging Face was the perfect example, right?

34:59 They went to defend themselves and what happened? It was like, oh no, none of the closed model providers is letting me doing it cuz it's a cybersecurity problem. So they ended up reaching towards one of the Chinese open weight models. So therefore we need to open the the playing field and let people be able to design their own solutions and and do things.

35:21 And I think that's going to be the case for even outside of cybersecurity for super talent intelligence in general. Um if somebody has got access to great computing, somebody's got access to the best models, right? What happens when I don't know, you're in a legal scenario or what happens when you're trying to build a startup? You've got the best you've got access to the best AI, but them over there don't.

35:41 The the disparity becomes a The best way of solving this is open is always the best way forward. I I and and I think at some point and I'm going to call it and I've called it before, there will be a point where open will overtake the closed models. That will happen at some point, right? It's just a matter of time. And when that happens, we're going to we're going to have some fun.

36:05 >> And on that note, that's all the time that we have for today. Volkhard, Rin, Chris, this is a stellar panel. Always happy to have [music] you on the show. And thanks for joining all you listeners. If you enjoyed what you heard, you can get us on Apple Podcasts, Spotify, and podcast platforms everywhere. [music] And we'll see you all next week on Mixture of Experts. >> [snorts] >> Mhm.