← All transcripts

Pacing the AI frontier, IBM Granite 4.2 & Meta’s Muse assistant Transcript, AI Summary & Key Points

IBM Technology · 13 days ago · Education · 38:47 · EN

Watch on YouTube

AI Summary

Frontier AI development is undergoing debate over pacing, audits, capability checkpoints, agent limits, sandboxing, and government controls. AI agents can coordinate at large scale, reduce the cost of finding software vulnerabilities, and create risks that defenders may struggle to investigate quickly. IBM Granite 4.2 provides Apache 2.0 dense models in 3B, 8B, and 30B sizes with native step-by-step reasoning for enterprise workflows, tool calling, planning, coding, and other agentic uses. IBM also discusses Granite Speech 5.0 O, mid-training, synthetic code, and future Granite releases. Meta’s Muse runs in a secure virtual machine with its own browser, but handing it access to personal data, accounts, banking, and purchases raises security, privacy, and unintended-use concerns. Personal agents may become more widely accepted as consumers shift toward mobile and AI-mediated shopping.

Key Points

  • Pacing — Dario Amodei’s essay "We Must Pace the Frontier" proposes slowing frontier AI development through measures including embedded auditors with employee-level access and capability checkpoints tied to alignment certifications.
  • Pacing — AI agents can spawn or coordinate thousands or tens of thousands of agents and sub-agents, potentially lowering the cost of discovering hundreds of day-zero software vulnerabilities.
  • Pacing — Potential controls include guardrails around model, agent, MCP-server, and tool calls; limits on concurrent agents or language-model calls; observability; sandboxing; and kill switches.
  • Pacing — AI safety concerns involve both model behavior and the environments in which models are trained, tested, and deployed, including internet access, containment, and the complexity of investigating incidents.
  • Pacing — Abraham Daniels questions broad unilateral government regulation and emphasizes that frontier-model controls may need to focus on a small number of labs with exceptional compute and capital, while also considering competition with China and restrictions on open-source models.
  • Granite 4.2 — IBM released dense Apache 2.0 models in 3B, 8B, and 30B sizes, with native step-by-step reasoning for planning, tool calling, automated coding, and enterprise agentic workflows.
  • Granite Speech 5.0 O — IBM’s speech-recognition model has approximately 450 or 470 million parameters and was described as transcribing three hours of audio in one second on a laptop; its stated speed is 1,200 times real time on a single H200.
  • Granite training — Mid-training uses structured data between pre-training and post-training to guide capabilities, domains, output structure, and agentic tool calling while supporting faster and cheaper serving.

AI in practice

Used for

Agents

  • OpenAI agents — Carry out agent experiments that involved interacting with external systems. 2 held 02:18
  • Muse — Act as a personal assistant that can access the internet, book trips, make purchases, and interact with a person's everyday digital life. 1 held 26:10

Tools & resources

1 item

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Pacing the AI frontier, IBM Granite 4.2 & Meta’s Muse assistant — IBM Technology (38:47). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:01 just was actually calling for these guard rails to be implemented and the slowdown to be implemented is the fact that the likes of Dario himself are calling for a slowdown even though Andropic is the one that's actually driving the frontier. All that and more on today's mixture of experts. I'm Tim Hang and this is Mixture of Experts. Each week, Moe brings together a panel of smart cookies working at the frontiers of artificial intelligence to lead you through the week's news.

00:31 On this week's episode, we've got Abraham Daniels, who's a principal product manager for AI Foundations, and Mihi Kvetti, who is CTO of Watson X orchestrates. Uh, as always, I'm joined by my co-host, David Zach, who's staff writer for IBM Think. Uh, we got a bunch of big stories. Uh, feels like the news is just overwhelming in AI over the past few weeks.

00:51 Um, but there's three that we're going to pick up on on today's episode of We're going to talk first a little bit about uh pacing, which has been very much in the news. We'll talk a little bit about the new granite release and also the release of Muse, which is um Meta's first real personal assistant release uh that we have been long awaiting uh here ate.

01:14 Let's start by talking about pacing. I think it's uh all the rage. Um uh people are talking about existential risk in places that I would never have imagined. Uh would be places they'd be talking about it and uh we want to have that as the top conversation today. >> Yeah, I I saw our colleague uh Nick Douglas uh on LinkedIn um called this apocalypse week.

01:34 It does feel like all of a sudden everyone's talking about um the existential threat of AI and um the call to you know arguably slow down or to pace development of AI. I think before we get into the the nuts and bolts of what is meant by pacing and turning it over to the panel, I I just think it's worth maybe rewinding a little to discuss how we got here and sort of the why now because I think a lot of us who follow AI know that um AI practitioners have been concerned about the risks of AI for a very long time.

02:03 So what happened all of a sudden that everyone's talking about it right now? And a few things happened this summer and in the past couple weeks in particular, but rewind back to May through June, I believe it is. And um Open AI was doing some testing on some of its agents and as view listeners of the podcast know um these agents sort of broke out of their sandboxes, hacked this platform called Hugging Face.

02:27 And in recent weeks, more details have come out about that that that hack and have been published by auditors. And we have, you know, learned about sort of deceptive behaviors and frightening swarmlike behaviors of these agents. So even though this happened months ago, that the details are trickling out in recent weeks, I think, and that's really uh raised the temperatures for a lot of people.

02:50 Um, another thing that has happened is the the rise of um, so-called recursive self-improvement or sort of this observed phenomenon within the AI labs where the people researching and developing the next uh, wave of AI models say that the AI is itself kind of doing most of the heavy lifting of doing the research and the analysis and the devel the the the improving of itself.

03:16 And there are concerns about this as sort of like an infinite or accelerating feedback loop. On top of this, a a researcher at Anthropic resigned and issued a a dramatic series of tweets that really caught a lot of attention. That researcher's name is Jacob uh Coxin and he had tweeted some very stark things about existential risk and AI. Um he assured everyone that um it was not a marketing stunt that uh that people are genuinely really concerned about uh long-term and even short-term risks uh that AI technology is

03:50 posing. that blew up, got 170 million views, and then we get to this theme of pacing, right? So then, um, Dario Amode, the CEO of uh, Anthropic, publishes, um, I I think of these as encyclicals now, like the Pope sort of issues these sort of encyclicals every 3 to six months. and and the I think the AI CEOs they they have uh uh taken to they're all essayists now and they they publish every on a cadence of every couple months these multi,000word essays.

04:21 So Daario publishes this multi,000word essay about pacing called we must pace the frontier and that is what we want to talk about today. Um so what is meant by pacing? Dario has some ideas. Um I just want to focus more on the technical and organizational components than maybe the complex geopolitical things that we don't need to get into. And um you know a few of the ideas that he has suggested would be the embedding of auditors within firms with employee level access.

04:50 Uh he also has some more technical ideas that I want to get into. But um broadly maybe starting with Mihi, you know, if you were able to read um either Daario's 4,000word essay or summaries of it, what's what stood out to you uh in this we must pace the frontier essay and this call for pacing broadly? I think one thing that stood out to me is just who's actually calling for these guard rails to be implemented and the slowdown to be implemented is the fact that the likes of Dario himself are calling for a slowdown even

05:25 though Andropic is the one that's actually driving the frontier. They're responsible for building and generating very powerful bubbles. And I think it's just an interesting perspective to look at why they're calling for this slowdown. So they're seeing some risks that obviously we don't know yet. We don't have the same information they do. And some of these risks are starting to become more and more I would say relevant across the other players.

05:55 So one thing that I've been personally calling out for is to implement guardrails for large language models for agents talking to other agents and for MCP servers talking or being talked to by agents. So for example I've implemented something called context forge which is an MCP gateway with guardwells that can be applied before and after an agent call a model call a tool call for similar risks.

06:18 But what what we've seen with the recent open AI attacks with the recent issues that even entropic and meta have highlighted is that these models have the ability to call thousands tens of thousands of agents and sub agents to really scale an attack which we didn't think was previously possible. So the economic implications of that are that discovering things like day zero vulnerabilities in software for example uh can cost very little money and you can find potentially hundreds of these day zero vulnerabilities in

06:55 software that can be used to attack the likes of you know nuclear power plants or you know the grid energy grid or large companies and cause major economic damage. So I think it's it's mostly an economic thing. Only highly skilled hackers and nation states could af could afford the research required to create these very large exploits and they would only use them for you know viable targets.

07:26 Now the cost of developing such a software vulnerability has gone down considerably because of AI where almost anyone with limited skill can find these you know day zero vulnerabilities in software and can use them even on you right you're now a viable target prior to that nobody would burn their day zero exploit on a random target now the economics are in favor of the attackers so I think this is this is a situation that needs to be addressed test in some form or another cuz defenders don't have access to the same

08:02 tools either. We've seen it with the open AI attack on hugging face where guards were actually preventing the defenders from being building and patching the vulnerabilities within their platform. So we don't know if guardrails are going to be sufficient. Maybe we're talking about things like a kill switch. We're talking about how these models behave.

08:20 If we're talking about limitations on number of concurrent agents which could run or concurrent LM calls which could run, it's potentially a scary situation. >> I I confess the scale of that that exact number of uh that question of sort of how many agents should be allowed to coordinate and and uh spawn and instantiate at any given moment um has been what has been concerning me.

08:45 I think one thing that's been concerning me is just how much time it's taking merely to investigate this event that happened a few months ago, right? And one thing I'm interested in is the the ratio of um how slow a handful of meatbased brains at Meter and other auditors need, you know, how much how how much glucose you need to burn to comb through these transcripts of um this these silicon based brains that just uh just work so so rapidly and in such concert with one another.

09:22 Um, I didn't see a lot about that in Dario's essay, but uh uh but um Abraham, I'm I'm curious to know what what are you thinking about um pacing uh whether it's to do with this uh theme of whether there should be limits on how many agents should be allowed to coordinate or maybe also some of the other themes that AMD talks about. Um one that he offered was this idea of capability checkpoints.

09:50 So he writes, "One possible scheme might be a series of checkpoints. If models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z." So by alignment properties, um I I assume we mean anything that uh um correlates to or helps improve uh you know AI alignment with human values, right? So um he lists a few properties but um uh one might be the ability to interpret or just sort of even read the the model's mind and understand it.

10:22 Um another property uh might be um just well he talks about if if a model can escape like security sandboxes maybe we really need to have a process of certifying and analyzing and auditing a training its training environment before we ship the thing. So how does this all resonate with you as someone who designs, trains and and deploys models? >> Yeah.

10:47 So a lot to that question. I may take a step back and think of this from the perspective of like the the I guess the the narrative of the the essay where you know pacing frontiers in my mind pacing frontier models is a frontier lab problem like the the core value that enterprises actually extract happens below the frontier level. You know smaller local more auditable models.

11:07 So I think pacing the frontier is really kind of targeting a very select number of labs and if you want to outline the labs that have actually had these security concerns um you're really thinking about two of them. So I think there is a conversation to be had about you know um instead of having these unilateral conversations about what to enforce across a broad spectrum of labs some of which might not be necessarily well equipped to handle some of these enforcements.

11:38 um as well as having a governmentmandated policy that truly only impacts you know a select number of labs given the scale of compute and capital that they have. Um I think this needs to be a conversation about how do we you know from the perspective of pacing frontier more specifically from the you know protecting um you know systems in terms of how models connect to them.

12:04 I think this is a little bit like I I'm not necessarily of the mind that this this this perfectly makes sense to be honest. Um there's also the balance between you know you know the the end of the essay commented on we also need to limit any sort of slowdown in order to maintain our um lead over China. So I think there is a little bit of a chicken and egg here.

12:26 There are some comments where, you know, Dario is making a case to to to constrain open- source, you know, models being deployed given a lot of these government controls that may or may not be, you know, enacted. Um, but yeah, I I think in short, like I definitely see the need for some type of controls. I don't necessarily fully agree with like a unilateral governmentbacked regulation across this in terms of what the essay outlined.

12:52 >> Um, just a thought on that. Um Abraham, I wonder how much of this is really an issue with the model itself or an issue with the harness the experiment or with the guardrails or sandbox environment that is put in place whenever one of these experiments is conducted. So if you look at the open AI uh incident, it's not necessarily that the model you know escaped and become became sentient and decided to hack X Y and Z.

13:14 Um the experiment was designed in such a way and the guardrails and the sandbox and environment were designed in such a way uh to make this possible right and is it sufficient for these steps to be taken during model training or is it more of a question that the deployment of these models when being used in either an experiment or in production needs to be done in such a way to ensure safety to ensure for containment to ensure that the sandbox is designed in such a way in which you know models don't have access to the

13:51 internet unless they're supposed to have access to the internet and that the frameworks themselves have things like observability and a kill switch and a mechanism to detect when they are doing things that are on section. So do you believe this is really an issue of how the model is trained or partially a component on how models are used in these experiments?

14:11 Yeah, I think that's right. And I think a lot of this ends up being that we don't really have a good discipline about how to build the world around these agents in some sense. Um, and it's a multiplying problem as the number of agents expands, right? I think David, you mentioned a moment ago part of the worry here is just takes so long to investigate these incidents.

14:27 You know, there's no m magic about that because like what's being spun up is very very complex and so it can take a lot of time to do the forensics. Um, and so I I think that's kind of the really interesting thing sort of that's at stake here is I think you know in some ways the researchers just don't know really. Um, and um, you know, the same the same problems that we had when we talked about Navier Stokes last week uh, are in some ways the same problems that we have here, right?

14:53 Is the proof good? How do we know that the proof is good? Well, it seems to have solved the problem. You know, I think we're in that world, but now it kind of applies to everything that these agents are going to do. And um you know part of the question which Abraham's raising which I think is really good is it's it's partially about like well what is a level of complexity past which we have no idea.

15:12 Um and what's the zone of ambiguity where hey I'm not looking at everything that my agent is shipping when I tell it to do something for me but I have enough trust in it that I'm okay with that. Right. Um and so that balance I think is the really really interesting thing that we're going to have to keep an eye on. I'm going to move us on to our our next topic.

15:32 Um, you know, longtime listeners to MOE will know that, uh, every time Abraham is on, we're going to talk Granite. And, uh, it's actually a really good moment to be checking in on Granite because, uh, Abraham, I realized just a few weeks ago there was a big release of Granite 4.2. I guess curious if you can tell us a little bit about, you know, what's new in the big release.

15:48 What are you guys focusing on? What should people be be looking after? Um, for those who have been following the Granite Project for a while. >> So, we released our Granite 4.2 language models um earlier uh earlier this month. So they come in three sizes, all dense models. Um so 3B, 8B and 30B, you know, and the scope of the size is really to kind of fit the use case that you're targeting.

16:08 All Apache 2.0. So that means you can pull them down from hiking face and use them as is train on top of them. So what's net new with these models is we've actually natively integrated thinking. So step-by-step reasoning. The reason behind that is you know we want to be able to provide more agentic enterprise um you know capable models for any particular workflows that require planning you know the right tool calling um anything that really does you know end to end agentic whether that's automated coding or tool

16:37 calling use cases along with the language model we've actually released a granite speech 5.0 O um turboctc all that really means is its own no LLM backbone attached to it so it's extremely fast um it's only 450 or 470 million parameters um again this is all available on hugging face to be used it's a ASR model so um audio speech recognition and just to kind of give you a scope in terms of the speed we're talking you know 1 1200 RTF on a single H200 which means you know nothing to no one but in the context of what that

17:13 you know what that actually allows you to do. So, say let's say you had a three-hour audio, you can transcribe that in a second on a laptop. So, um you know, this is just another kind of feather in the cap in terms of the local story. Speech models have been phenomenal and they really topped the leaderboards in terms of you know, ASR. So, we're really proud of what we've done there.

17:32 Um and we're not slowing down. You know, we're in the middle of driving for Granite 5.0 release, which you know is >> Yeah. No, no. Yeah, we're not we're not we're definitely not based, but we've got a lot of cool things in the hopper right now. So, we're really excited. >> We are granite accelerationists at IBM. You say it's available uh at Hugging Face.

17:52 Is that on a human controlled server or a botnet controlled, you know, pirate server that AI agents hack? >> It's whatever you want to Yeah. So, it's available on Hugging Face as well as a number of our partners. So, um yeah, yeah, pull it down, give it a use, you know, give us some feedback. you know, we're always listening to anything on Reddit or any other social site.

18:10 So, looking forward to kind of seeing how people uh how people interact with the model. >> I I do have a followup question or two about um I gather that uh there were a couple innovations in training that um really strengthened the reasoning capabilities of 4.2. Um there's a few we could talk about the training on synthetic code. The one I'm interested in, if you'll indulge me, is this so-called mid training.

18:34 So uh I am not a technical person but I recently just did a lot of independent study on sort of how models are trained and obviously the traditional way um is that you have this long period of so-called pre-training where basically the models read the internet and learn language and then you have um post-training and fine-tuning which is often uh reinforcement learning and sort of often humans upvoting and downvoting responses but I gather that um IBM is something of a pioneer and advocate for this mid intermediate

19:10 process called mid-training and that this has been showed shown to improve reasoning capabilities. Um Abraham could you tell us a little bit about what mid-training is and and why IBM has decided to make this investment in this sort of intermediate training step. >> So I I I I want to stay away from like this is a specifically IBM based training mechanism.

19:29 Um so mid training can mean a couple different things in terms of just like a precursor to to post training. So warming up with you know a select amount of data um specifically in the SFT side of things to either one gear you for a more um smooth training when it comes to post training in the RL environments. It also allows you to do some you know uh output um kind of structuring what's through speculative decoding layer.

19:54 So it allows you to do faster and cheaper serving. That's kind of the reasons why you know the mid training kind of come to fruition. Um I I don't know how much value it is going to the you know the the bits and bites of it but I guess the at the treetop it's really one to kind of give us a better understanding of or give us a better output with respect to uh you know structure and then just from the model's perspective it just makes it faster and cheaper to serve any sort of props or or questions that pass through.

20:24 >> Got it. And just so I understand uh is it so if pre-training as I said is sort of reading the internet and this kind of highly automated high volume process and if post- training is often slower and more expensive and involves RL or reinforcement learning this upvoting and down voting by by humans mid-training is something in between or it's something closer to the the former or closer to the latter.

20:46 >> I would say it's closer to the the latter. Um it's it's definitely not unstructured data, you know, early onset just trying to get a sense of the the the structure of the model. It very much still is structured data u organized in order to move the model or at least nudge the model in the direction of what you're actually hoping for it to learn. Whether that's a a capability, you know, either, you know, or domain or even just a particular agentic, you know, tool calling capability.

21:13 Um it's much more on that post-training uh side of the fence from a from a what the data structures look like. >> Well, you indulge me with my nerdy interest. What what innovation in the training process um has most excited you whether it's the synthetic code use of synthetic code or anything else? >> So the the 10,00 tokens of synthetic code which you're speaking about that's code alchemy.

21:34 So this is a uh open-source um data set that IBM research has released. I believe it was in August. um the it's really a a small part of a broader effort um in open elco which is a IBM you know researchdriven initiative to drive a diverse set of RL environments as well as SAT trajectories across a host of work streams that are not only important to you know developers but also have a very large enterprise footprint whether that's uh you know terminal uh interactions you know co-work interactions so working with

22:10 documents or web browsers um So I think co code alchemy was really the initial push to the open source community in terms of what we're we're cooking up. Um we are hoping to you know build out the capabilities or the work streams along code alchemy to be able to release the open source. So I think what's super excited to me is that's like you know 4.2 in this kind of open alchemy environment is really the precursor to what we're building with 5.2 in the open alchemy and I think we're just going to continue to build on

22:38 top of this to you know to deliver highquality u you know models that are not only fast but you know work within the particular environments that you know you operate in whether that's local whether that's on the cloud or whether that's you know on a device uh at home >> I'm I'm always advocating for those at the back of the class or maybe the less technical viewers for those who are less technical um synthetic code and uh maybe more broadly synthetic data why does it turn out to work right that you generate this data

23:06 and it doesn't wind up being like a fax of a fax of a fax and and degrading like why is it what how is it that we've been able to strengthen models using um synthetic code that is not sort of human produced. >> Yeah. So I I the the basis of it is that like synthetic code very much mirrors the structure the scale the diversity of code out in the wild.

23:25 The only difference with synthetic code is you just have a lot more control over the breadth. So the amount of the data that you can have. Um so what's great about it is if you have particular you know um you know domains or uh languages if you will that are either one um you know not as prevalent in terms of the data that's available whether that's because it's a lot of it is proprietary so organizations hold on to it or the data that's available in the open source market just doesn't have the you know the scope and

23:56 scale that you need to or the the data points that you need in order to sufficiently train a model to um be able to confidently from a user's perspective, you know, output good code. So what synthetic data allows you to do is to generate vast amounts of, you know, whether it's code data, you know, insert some type of, you know, modality of data um in ways that just previously weren't available to organizations or individuals.

24:24 So synthetic data really act gives you that that you know amplification of data so that one you know you can start to introduce not only net new data sets but you can start to introduce more diverse data sets you can expand on existing data sets to account for maybe gaps that you may have in it. Um so it really is kind of this not I wouldn't say it's a silver bullet but in a world of training models where you know there has been conversations of scale is dead you know training on more compute isn't the answer anymore

24:54 so you know moving to higher quality data obviously quantity is important but higher quality data um is a really big proponent of you know how good your models can be in the end. Yeah, for sure. Abraham, what how do people um follow on with this work if they want to check it out on HuggingFace? Is there a blog post you'd send them to? >> Yeah, check it on HuggingFace.

25:14 There's a you know, if you go to ibm.com or IBM research websites, you can find out, you know, where we are. We've got a number of partners, whether that's, you know, replicate, you know, we're working on some new ones as well. So, yeah, I would say HuggingFace is the best place to pull on the bottom if you want to work locally. If you want to learn about it, we've got a number of blogs on HuggingFace as well as IBM research and IBM.com.

25:34 Um, and then of course, as I mentioned, you know, if you got anything you want to share, local llama on Reddit, we're always listening. >> Well, I'm going to move us on to our last topic of the day. Um, you know, Meta has been cooking uh for the last few months. Everybody's been kind of awaiting what they've been uh getting ready to release. Uh there's been high-profile uh encyclicals from uh Zuckerberg uh about what it is that they're building, but um it is really just in the past week or two that we've gotten a

26:05 chance of what it is that they are actually building. And um it's a product called Muse, which is uh really their foray into the personal assistant and agent space. And so I guess Mihi, it's good to have you on the show uh today. Um to talk a little bit about your early impressions on Muse as someone who's worked on agents and agent harnesses um you know and and I'm really sort of interested in your impression on whether or not you know this is this is the moment right like everybody's been dreaming about the agent

26:34 that can book the vacation for you. Um and uh and I think Meta is trying to make a really serious play for this kind of space and I guess your impressions on you know uh is is this the beginning of of what's going to really flip that product area? So look, I'll give you kind of my technical impression first, which is I like what they've done with Muse and building it on top of a secure VM that has its own browser.

26:56 So you can segregate your data and it imposes some mechanism of sandboxing and isolation from your own system. So I think at least from a design perspective, they're making some good calls in terms of security and isolation. I still take issue with the concept of handing over your personal life, your personal data to an AI or an AI agent because at the moment the state-of-the-art in terms of guardrails, security, sandboxing are not yet at a place where unintended use can actually mess with your bank account, can mess

27:36 with your credit card, can mess with your personal data. What if it books the wrong vacation and instead of, you know, going to Hawaii, you end up in some other place, right? And now you're like, "Hey, I'm stuck here. I'm in the wrong I'm in Dublin, Ohio." >> Problem is you like got to get onto the plane and you realize you're actually headed somewhere else.

27:57 >> It's like, you know, this flight to Dublin is taking a very very long time. Is Dublin, Ohio. I was like, "Whoops. If I were to use this, I would still use it from a airgapped machine. that is dedicated only for its use. Even though it's in a secure VM and I wouldn't put all of my, you know, accounts and access and information and banking and everything else in between.

28:18 Maybe I would use a credit card that is, you know, a debit card that is preloaded with the right funds and has some mechanism of approval of the spend. I still think I think it's valuable research and I think the design they've taken is security aware. I would maybe let some other folks test it first. I I have a question almost more from the standpoint of like a as a consumer and from like a brand standpoint.

28:48 Um obviously technical people know that a meta has been a leader in AI with their llama models for for many years. But for those of us who just kind of interact with meta as users of Facebook and Instagram and in my case I have like an Oculus that I occasionally, you know, play Beat Saber on. Um, how should we understand why Meta, the social media and entertainment company, is even getting into the business of productivity?

29:15 Is this part of a play to, you know, to to build hardware that does everything for you as you navigate through the world? I mean, I feel like when when Meta launched their kind of characters, their sort of character AI style um AI agents or not agents actually AI conversation partners in the messaging apps that felt like a a certain sort of organic brand extension, but I'm struggling a little bit as a meta consumer user to understand why I need a meta like an agentic meta assistant.

29:50 And I'm just curious what you think is the play here. Maybe I'll maybe I'll put the question specifically to Abraham, like why why do you think someone might reach for a productivity assistant in the context of a a universe that that people mostly experience as kind of social media and entertainment? >> Uh yeah. So I I mean from the perspective of why I think it's just where what attention can they capture?

30:13 What uh you know what user journey can they capture where there is they're involved as much as possible. So if I think of like the the the meta value chain from you know from WhatsApp to Instagram and everything in between having a productivity agent allows them to be embedded in your day-to-day life even more so than they are already. Um especially if you're already using their products I would assume there is some you know simple integration between whether it's MCP server between like their already existing

30:42 application. So I think from their perspective this is just a a stack on top of the or a layer on top of the attention stack they already pull from people and um if they can start to integrate into how people spend money book trips basically access the internet which is something that you're doing all day then they just have another you know they have another opportunity to either one embed some type of advertisement or sales cycle or sticking point with that particular user that they already have.

31:13 Um so I think that from my perspective that's the why. Um and then the how is obviously through a purpose-built autonomous purpose you know personal model. >> Yeah. It also makes me think a little bit about the um you know the huge scale and scope of Facebook marketplace. You know we have sort of thought about these you know products as kind of entertainment and social media products.

31:38 But it is sort of interesting that Facebook in many cases is infrastructure, right? People are actually using it in a very kind of not really business enterprisey way, but certainly a personal enterprisey way. Um, and so I think that may be a little bit of what sort of meta is seeing here is sort of the idea that that is to say, well, it turns out we just live our lives in social media now.

31:57 Um, and so it's not completely insane that, you know, you'd also have an artificial intelligent buddy who occasionally books vacations for you. Um and so that kind of extension is kind of latent in um you know like the degree to which you know certain products like Instagram are just kind of part of the fabric of people's lives. >> Yeah, there could be a cool tie into just from like a you know point in time purchasing it through your personal agent via Instagram marketplace or what have you where you don't have to exit

32:28 their social media landscape. You can do everything within it which is again another kind of sticky point for them. Yeah, I it's interesting. You know, I I before I joined IBM, I was a tech journalist for many years at Fast Company. I remember we were talking about the Facebook phone and Facebook's foray into hardware, you know, uh over a decade ago.

32:46 It'll be interesting to see what happens, especially when I gather now um OpenAI is partnering with Johnny I formerly of Apple to sort of create a a goes with you everywhere assistant. Uh obviously my our colleague Sasha and a frequent flower I think on this show show um wears the often wears the meta glasses and I would be very curious to speak with him in a future episode about you know how and to what extent he's using Muse.

33:14 So it will certainly be interesting to see how um these products develop and um whether indeed there is brewing a sort of uh competition for um for everyone to have their kind of agent the productivity agent that they feel most comfortable with whether that's embedded in a meta glasses or in your you know your cloud application or or uh in um you know whatever openai is is cooking up with Johnny IV.

33:42 So Tim, you you spoke about how meta is essentially infrastructure and we've been living our lives on it for a you know very long time now. I did read um Emma Roth's review in The Verge and she had some nice things to say about Muse. She also had this experience of this sort of uncanny experience of it knowing more about her than she seemed to know about herself.

34:06 it began to suggest things in her feed, I guess, that it gathered uh that it had sort of inferred her interests. And I think this is speaks to this sort of um zero sum game we sometimes experience in in these products between utility on the one hand versus a feeling of privacy and comfort on the other. And I'm curious, Mihi, how do you what is your guess about how consumers are likely to navigate these trade-offs of of, you know, utility on the one hand versus the eeriness on the other?

34:38 I think at least for now, not a lot of individuals are willing to hand over their personal lives, personal data, passwords, access, and I know I know they say, "Oh, no, passwords are protected and so on, but you're still connecting them to the system." Um, to an AI model that can potentially train on your data, and I've seen, look, you can opt out, but already that's a dark pattern, right?

35:00 you should opt in, not opt out of them training or perhaps serving ads and they say they don't and there's an opt out feature but I've already seen issues with Metam's image for example a couple of months ago when they released it allowed users to what generate images from Instagram pictures or something along those lines and he caused a scandal on privacy.

35:25 So I think there might be fatigue in the general public with all these type of use cases where you're handing over your personal life to an AI assistant or anything of the kind and it's going to help you with shopping or help you consume more or help you spend more money or help you do some research. I think the tools we have today where you can do your own research and then separately choose to pay what to buy, what to purchase and make an informed decision are sufficiently convenient for the difference and the

36:00 surrender of privacy not to be worth the investment. So I don't see it necessarily taking off in the short term. long term. However, I think a new generation of consumer is being trained to do their shopping on Tik Tok or to their mobile phones or through other kind of things. I'm already of the older generation that, you know, there's a joke that if you're a millennial, you can't do any kind of shopping or anything unless you go, "All right, let me pull my laptop.

36:26 You're not going to touch your phone. You do everything. I need to see a physical keyboard. I need to be able to do my own research, spend a bit of time, look at some a couple of reviews of what I'm buying. I don't necessarily trust just the AI recommendation on what to buy. And then maybe the next day or in a couple of days, I'm going to go in a separate um browser window and I'm actually going to purchase this thing.

36:51 So I think it won't take off with the current generation of consumers. But in a couple of generations, so maybe in you know 5 years, 10 years, 15 years, this might be the new way that people consume and shop and make purchases. Technology is going to evolve. Maybe this will be the only option in many situations. It's going to be deeper integrated into the marketplaces.

37:12 The experience consumers are going to have fewer and fewer desktop devices and laptops and move more to tablets and phones. So, I could I could see it as a long-term investment. >> I mean, how do you go to a separate browser to avoid the dynamic pricing and like cookie surveillance to make sure you get the best price? Um I I do I do and for example I've made a mistake of not not doing that now and I'm getting ads for something I've watched every single time.

37:38 So if I want to you know get a surprise from my wife or if I want to do anything I really need to do it in a completely isolated VPN instance otherwise I get ads for it in everything everything that I open. So I'm I'm trying to keep some discipline. When Mihi does e-commerce, it's like a John Lere novel in terms of his opsseack and like trades spycraft to make sure he he gets the best deal.

38:03 I think that's the millennial way. But if if we have Jen Alpha on the show, they're just like >> AI bot just just beat, you know, send me to Dublin, Ohio. >> I'll pay million dollars. Whatever you say. >> Well, uh all that and more on today's mixture of experts. Uh Mihi, Abraham, David, thanks for joining as always. and Mihi will have you on in about 10 years to check some of those predictions.

38:27 Thanks to all you listeners. If you enjoyed what you heard, you can get us on Apple Podcast, Spotify, and podcast platforms everywhere. And we'll see you all next week on Mixture of Experts.