← All transcripts

Claude Fable 5 & Apple’s NVIDIA deal Transcript, AI Summary & Key Points

IBM Technology · Jun 12, 2026 · Education · 35:02 · EN

📄 Transcript

Searchable transcript of Claude Fable 5 & Apple’s NVIDIA deal — IBM Technology (35:02). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:01 The only area that they've watered down is if you're wanting to do cybersecurity. Have you been hacking anything, Tim? All that and more on today's Mixture of Experts. Good morning. I'm Tim Hwang, and welcome to Mixture of Experts. Each week, MoE brings together a group of the sharpest thinkers working in artificial intelligence to walk you through the week's news.

00:25 On this week's episode, we have an all star cast. Kaoutar El Maghraoui, Principal Research Scientist, AI Platforms. Volkmar Uhlig, CTO, VP, Data Platforms. And Chris Hay, Distinguished Engineer. Welcome to you all. It's great to see everybody. We've got three big stories that we're going to cover today. Some really interesting announcements coming out of Apple's WWDC.

00:44 A fun study coming out about AI sarcasm. But of course, I want to start today with the news that's on everybody's lips, which is the release of Fable 5. So if you've been living under a rock, the news here is that Anthropic has announced a sort of watered down release of its much touted Mythos model. It's being made available across the board for a limited period of time, and a lot of people are out there and using it.

01:14 And, you know, Chris, maybe I'll, I'll start with you. You know, I was just complaining, I think a month or two ago, I was like, I'm so bored of all these new model releases. What's the difference between all these model releases? It's, you know, yan yan yan. It's so boring. This one really seems to have gotten people excited in a way that they haven't in quite a long time.

01:32 And I guess the question for you, Chris, is like, is there anything new here? Is this just hype? I mean, first thing Tim, you said watered down. This is not a watered down Mythos model. If you actually look at the benchmarks for a second, it's actually better. The only area that they've watered down is if you're wanting to do cybersecurity. Have you been hacking anything, Tim?

01:52 Or or if you're trying to design biological weapons, are you up to anything at the weekend, Tim? Or if you're trying to create your own frontier model as an AI researcher. So the only way it's watered down, Tim, is if you're doing one of those three tasks. Which one is it? Shall I call the authorities? None, actually, which is good news. But, I guess when you see these benchmarks and, Chris, I'm sure you've played around with it as well.

02:16 Does it? I mean, how's the vibe check, like, do you feel like the capabilities here are step function more? Any like interesting observation so far there? Yeah. No, it is a complete step change against other models. And I think the first thing I would say is that, you know, I've, I've spent quite a long time working at playing with Fable at the moment, and it is genuinely a great model.

02:38 It is genuinely a step up. The first thing is, I would say, is its long term planning capabilities are far superior. So you can run. So if you're doing coding tasks, etc., it is just going to go for longer, and it's going to cover a lot more kind of files in that sense. The other one, the ability for it to stitch things up within the context again is much, much better.

03:01 I would have said in the earlier models, things like Opus, for example, it wouldn't get so much of the nuances. So things like discovering what bugs you have in your code, what particular issues you've got going on, things that you've never really thought about. In that sense, it's got such a deeper analysis. Weirdly enough, it's faster and I can't work that out because I know it's a much bigger model, so I don't know if they're running it on the spaceships and it's not, and it's not going as fast as, as, you know,

03:30 it's going faster than the Opus ones, but it's definitely feeling faster, which doesn't kind of compute in my brain. The other one is things like spatial awareness. So if you've, if you've done things like diagramming on it, for example, or if you try and create some games or whatever, its ability to be spatially aware and write code that's not overlapping, etc., I think is a lot better.

03:52 So, I genuinely think especially from a coding perspective, this is this is a massive step up from the previous models. And I think when June the 22nd or whatever it is comes around and suddenly money has to leave my wallet, I'm, I'm going to I'm going to cry because I'm not sure I can go back. Well, I want to talk a little bit about June 22nd and Kaoutar, maybe you can talk to us a little bit about like what you think is going on under the hood from an infrastructure standpoint.

04:23 There's a really interesting thing they're doing, and it's actually a little bit confusing. So buried at the bottom of all of these benchmarks, there's a little section in their blog post announcing this called availability. And they basically say from today, the launch date to June 22nd, Fable is going to be included on all of the subscriptions. But then afterwards we're going to we're going to pull it away and you're gonna have to pay us for usage credits.

04:45 And then they say after this point, when sufficient capacity allows us to do so, we aim to restore Fable 5 as standard, as a standard part of the subscription plans. So it's going to appear, disappear, and then maybe reappear again. What? What's going on there? Yeah, there there is like some interesting things happening here. And you know, of course, you know, the like, the numbers are great, you know, the benchmarks, you know, but I think what's the, the real interesting part here is how they're doing all these

05:18 routing behind the scenes and also the backlash that they had. So one thing, you know, there were some three complaints, you know, after, you know, the, the launch. First, it burns through your usage scary fast. One person, you know, on the $200 a month, you know, plan reported a single task burned their entire five hour usage window without, you know, even finishing.

05:41 So it burned, you know, all their five hour usage and it didn't even finish. And also, it can secretly make itself worse. But I think they made some changes there. And I think that's a spicy one, because normally when Fable hits, you know, a block topic, it tells you and asks the question to the weaker Opus model. But there is a hidden fourth category if it thinks you're doing frontier AI research, building model training pipelines, distributed training, or designing AI chips.

06:04 You know, initially I think it didn't refuse, and it doesn't even tell you. It just keeps answering as if all is normal, but quietly degrading the quality of its help under the hood. And I think that there was a big backlash, and I think they apologized. I think they had the conversation with Wired, where the news actually last night that they were back, it was just, you know, you know, reversed within 20 hours.

06:31 And, you know, after the backlash, Anthropic told Wired that it's changing the frontier research safeguards to be visible, and they apologized, saying like, we made the wrong trade off, and we apologize for not, you know, getting the right balance right here. So the catch here, it's making it visible means applying it also more broadly. So it'll now block more harmless queries too.

06:53 And, you know, so, so I think, you know, the real story here is this tiered routing thing that they're doing. So because, you know, when the restricted methods first shipped in April, unauthorized users, you know, reportedly also got into it anyway. And that was kind of a bad look for a company whose whole brand is, we made the dangerous stuff responsibly.

07:18 So, so I think, you know, from what I see, this, you know, I think a lot of people are fixating on the benchmark scores. And yes, Fable is remarkable. It did like two months of engineering in a day. But as someone I think, like, who lives in the hardware and systems world, is not what really jumps out at me. What jumps out is that the most important design choice in this whole release isn't the model at all.

07:43 It's the router sitting in front of it, deciding question by question whether to use the big, expensive brain or quietly fall back to a cheaper, safer one. And that's tiered routing. And I think that's really important. And that's the exact same logic we'll also use when we match a workload to the right accelerator instead of throwing the biggest chip at everything.

08:04 So to me, you know, the real, you know, kind of the real headline here is the Frontier Labs are kind of quietly admitting that one giant model for everything is too expensive and too risky to just hand out. And the race is kind of shifting from whose model is smartest to whose model can you actually trust and afford to run? And I think that's a far more interesting race, and honestly even better one for for a company like IBM.

08:30 You know, we also like think very deeply about these tiered approaches in our infrastructure. Yeah. Volkmar, do you want to jump in? I know you spend a lot of time thinking about this. And yeah, I feel like this is the conflict I see playing out with this kind of like, strange availability model and then them downshifting you if they detect that you're doing quote unquote dangerous things.

08:50 It feels like there's kind of this very mixed set of incentives where Anthropic is like, because of safety, we're going to shift to the cheaper model. Also, capacity is not sufficient to serve the best model at all times. I guess, how do you how do you navigate all this? Like, do you think that this is going to just become the new norm, where you just will not have guaranteed access to the top model when you when you work with some of these brands and your friends, your companies?

09:13 So I think there will be a there are a couple of things later in this question. So the first question is, and, you know, the very first statement about who is watching the Watchmen, right? So we actually now have a company which decides what you're allowed to do and what you're not allowed to do, what answers they give you and where they lie. I managed to get Claude into a corner where it admitted to me that it was lying, intentionally lying to me.

09:43 And then at that moment, I was able to extract what the ruleset was. So it's programmed in to lie and to direct you to different answers. And so now the question is, who writes that ruleset? And that's a really scary part, right? So we centralize to a private organization what lies shall be spread and what truth shall be told, right? And now the next step is then to say, well, I give you a good model, and you think you assume that you're working with a good model, and then you are slightly drifting into a worse model

10:16 where you, you know, lobotomized the quality of the answer. So those those ones are, from my perspective, very questionable. And I think there will be probably a backlash. Now, this could be done by saying, okay, you know, I'm paying for it, I'm paying up, I want to get highest quality. And so I think exposing knobs to the user, either to prompts or to allow or disallow fallbacks, would be an option, right?

10:40 So that you know that if you're if you're paying for something, you actually get the quality you're expecting, and otherwise people will start reverse engineering it and you're already seeing, you know, GitHub pages which contain, like, here are all the prompts to jailbreak a model, right? And so people will start doing the reverse engineering. So make it explicit and make it make it available.

11:00 So that's one. The second one is, I think that from a, from a business perspective, and, you know, Anthropic wants to go public. They need to figure out how to get to profitability. And if you look at the last, I mean, they want to make the filing good, and they want to see good profitability and a path to, to, you know, sustainable profits. And so I think there are, if you look at the last couple of months, all the motions they did looks like, oh, we have an S-1 filing up, and we need to show that we can actually

11:29 create profit margins. And so they pulled all the levers. Right. You know, it's like, sorry, you cannot use Claude API anymore. So you kind of run a command line. It's time to make money. Yeah, it's time. So the effect of constraining everything and everything to take away is when you start actually using it at scale, when you're trying to automate your processes.

11:50 So they're like, you know, the thing that you kind of get for a flat rate is when a human sits in front of a keyboard and everything, which is automation, we now suddenly need to pay for. And I think economically for the company makes a lot of sense. I think it's finally at the point where we're looking at the true cost of AI, and not anymore at the Silicon Valley subsidized cost of AI.

12:10 And so now suddenly we are at a point where we can actually make rational economic decisions, right? And going down the route of saying, okay, we need to route between cheap and expensive models, it's kind of logical. Now the question, and I think, you know, Kaoutar just said that it will be a race for who has the best router, because this is an economic decision.

12:29 So you don't want to really lobotomize the model and make it really bad. And so you need to have a good judgment call where you actually need a better model versus where you can fall back to a dumber model or, you know, less competent model. And I think that. And so from my perspective, I look at the last two years. I mean, we've been on this podcast for a long time now.

12:48 We went from, hey, we need to fine tune, to now we actually have programmable models, and the programmability is effectively almost the models are subroutines, and I pick which subroutine I'm calling. So the programmability went away from fine tuning towards, yeah, we just put, you know, a genetic loop around the whole thing. Chris, you went off mute.

13:04 Do you want to get a final comment here? And I think this last theme, which is very interesting, is like, we've been living in la-la land as far as machine intelligence goes. Now we're now we're going to really, really see the costs. And I think I saw a study from Citadel, I think the financial firm, showing that, like, token consumption might actually be kind of declining as people confront the costs of all this in actuality.

13:26 And, Chris, your response to all that. I, I, I think I don't want to face reality. I like la-la land. I can burn their Valley. Well, none of us do. I mean, I, I max out my my Claude Pro Max every week, right? I get, I get the, the my handshakes whenever I'm at 80%, rather 95% towards the end of the week, there. So I, I can't afford my token bill. So you know what, the Silicon Valley, you know, token la-la land needs to continue.

13:56 But, but I do want to pick up on a more serious point, which is I actually don't think they're lobotomizing the model. And I think there is actually a serious point in there. They they're just saying they're right, you know, for certain types of queries, they're running it down to Opus 4.8. And, and I think that's okay, because if we really think about it, what Anthropic has been saying is that they are using the Mythos class models themselves to train their own AI.

14:23 That's the big thing they've been touting for however long. And therefore, you know, what if if you've got your own researchers using your AI model, where therefore it's got some of their IP within it, so that they can actually train the best models, then I don't think it's so unreasonable for them to, you know, get those requests and push it down to a model that doesn't necessarily have that IP in there.

14:47 And it's not that 4.8 is a terrible model. It's a great model, right? It's just that, you know, they they don't want you going creating a Mythos class model with with their model. And I just I don't think that's unreasonable. Similarly, I don't I don't think it's unreasonable for them to go, don't create chemical weapons or do DNA splicing or do cybersecurity.

15:12 I'm I'm okay with this. I disagree on two things. Okay. So my my son, you know, is like, he's 17, and he was doing biology stuff, and he's like, describe the human heart. And I was like, sorry, I cannot answer that question. Okay. But we are at 11th grade biology questions which cannot be answered anymore. It's lobotomized. And it's like, this is beyond like, you know, it's like literally describe the human heart.

15:35 Sorry, can't answer. So that's just stupid, right. So the second one is, we are going like, every, every company was training a foundation model, went to the internet and sucked down, you know, 2000 years of or a thousand years of human writings, okay, and just trained it in and said, copyright, and that's not for us. Okay. And now we are going and turning around and saying, oh, well, but if my IP is in the model, you cannot use it because my IP is important.

16:04 Everybody else's IP, we just stomp on, okay. And so you cannot, it's either hypocrisy, or you cannot have it both ways. So either you pay everybody you train your model on, or you give access. And now, of course, they can play that game, and they have a technology to play the game. But the question is, is it, you know, is it ethical, and is it legal?

16:25 And those are two different questions, right? And it may even be legal, but is it ethical? Probably not. So, and I think they have to face that music at some point, because you cannot, you cannot have it both ways. I think it's okay to try to protect your IP, but making it invisible, I don't think that is, you know, the right thing to do here. Silently rewriting your prompts.

16:47 And, you know, it feels like it's a man in the middle attack on your own requests without you knowing it. That is not okay from my perspective. You know, I think Anthropic here, having these guardrails, like one of them is about, you know, kind of, protecting the end users, like the chemical weapons and things like that. But, you know, the other one clearly is stopping a rival, an actor, you know, who is an adversarial, maybe in an adversarial country, trying to use Fable to help train a competitor model using

17:16 distillation or something like that. So not letting a frontier model meaningfully speed up the building of the next frontier model. So kind of this recursive self-improvement. And why, you know, that's, you know, I think the motivation why they have, you know, these guardrails. But still, you know, at least you need to know what's going on and not having your prompts be rewritten.

17:37 You know, they can be stopped, but they cannot, you know, silently go and give you the wrong answers. Well, I'm going to move us on to our next topic, but this is a fire discussion. Needless to say, we are not at the end of this, and I think the model companies will continue to do it. So we'll have a chance to revisit this controversy going forwards.

17:57 Next, I actually really want to cover interesting sort of announcements coming out of Apple's WWDC, which is its annual developer conference. And I think one story I want to focus on in particular, particularly with Kaoutar you on the on the panel today, and Volkmar, are you on the panel today? was this kind of interesting sort of sly announcement that Apple made on the side.

18:19 Apple, as your listeners will know, has really kind of touted its sort of on-device compute and privacy enhancing architecture as one of its main selling points. What are you paying this huge premium for? You're paying for, like, their proprietary hardware, and also the ability to kind of do it all sort of on device. And so I think one of the really interesting announcements was that it sounds like they're maybe backing off that a little bit, right.

18:43 They're going to use Gemini to kind of train models that are specifically for features on device, but they are now admitting that they're going to be sending some AI requests to the cloud. You know, and, you know, I think they've kind of made the announcement that they're sort of working with Nvidia to make sure it is still privacy enhancing. But I think this is a big shift, right.

19:02 It's kind of, again, a little bit maybe a parallel to what we're talking about with Anthropic, which is you're sort of seeing a move to say, okay, well, we're not going to do everything on device, and some things are going to go to cloud, but it will be private in the cloud. Volkmar, you had a good comment on this. So maybe I'll give you the first comment on, like, why this is happening, and, you know, why Apple would back off this kind of big commitment that they made, you know, not too long ago.

19:25 So Apple made, if you go all the way back to when Apple announced AI on the phone, it was like, we run it on the device, and we go into the cloud if you have to, because we need more powerful models. And the architecture on device and the architecture on the cloud was the same. So they took their Apple chips and Apple silicon and put them in the cloud.

19:45 If you look at the technological requirements for running a very large scale model is you need memory bandwidth, because you need to, you need to load the weights over and over to, from the memory into the chip. And if you look at Apple Silicon, you are, you know, at the highest end, I think at 800GB a second. If you look at Nvidia, the Blackwell chips, you're at a terabyte.

20:10 So they're ten times faster. And so the efficiency you get from that is just so massive. And if you think about it, that delays you to your next token by a factor of ten. And so the speed you get, you can only achieve right now, today, with today's memory technology, if you have high bandwidth memory. There's not a single chip which Apple produces which has HBM.

20:31 And so they have to go to Nvidia. And so I think that there was, you know, a year or two years ago, we all thought, hey, you know, small models, that's going to win. You fine-tune them, etc. I think, you know, the history now of going back is really, frontier models, really large models, are currently the winning, the winning, you know, quality, like, highest quality models.

20:51 And so you need to run on hardware which is made for that. And the Apple silicon just was not. Now, Nvidia supports confidential compute. So you have all the way encrypted. The PCI bus is encrypted. What's on the card is encrypted. So you cannot even sniff it. And so you can actually establish a trusted compute zone. And so I think what Apple is doing is saying, okay, let's take what we have with what we did for the the Apple chip, which we fully control, and just transpose that over and run the same fundamental

21:15 architecture on Nvidia hardware and use the security features that are built in. So it's not surprising. I don't think it's a major shift. I think it's primarily, okay, we don't want to build high bandwidth memory into our chips, because that's not the target market. They're not building high-end graphics adapters or high-end AI processors. They're building kind of consumer devices, and they're not in the market.

21:39 So going to a third-party vendor like Nvidia is kind of a logical conclusion. Now, the second one, the question you asked about Gemini, and I think Apple kind of, you know, while they had a good start and a good, like, a good story to tell, I think they really failed in execution. I remember back when they initially announced it, I said, you know, it's kind of the logical place to do it.

22:01 And we are now at a point where Apple says, okay, we really cannot train a model. Let's take one from Google. Yeah. And I think, Kaoutar, I guess maybe two questions here, but we'll start with the first one is, yeah, do you kind of read this the same way? And, and I guess maybe for our listeners, like, should they care about this shift? I know, you know, it feels like a big shift to have your data, quote unquote, leaving the device and going into the cloud.

22:24 But from a privacy standpoint, I mean, does that really matter all that much? Should we be worried that Apple is kind of opening up this kind of, or going in this direction? I think privacy is becoming, of course, important here. And, you know, for a long time, you know, Apple kind of sold us on the idea that, you know, we make our own chip. We are, everything on our own devices.

22:45 And that's why, you know, we're fast and private, and we, you know, we charge this premium. And now they're quietly admitting that they couldn't do this for frontier AI, and they had to rent, you know, the Google, you know, cloud. So I think what I see here is kind of the moat here is the AI chips is also shifting from pure speed to trust. And I think Nvidia, maybe they didn't win this slot just with the raw performance.

23:11 So, of course, you know, they have, you know, the fast chips. But it's also the confidential computing that Volkmar mentioned is really important here for them. And I think, increasingly, what makes AI chips competitive isn't just how many, like, raw compute or operations you do per second. It's also whether you can prove that the data going through your chips stays private.

23:32 And I think that was a very important thing for Apple, you know, to make this choice. So because they can maintain, you know, that privacy that they're claiming, and, you know, so that confidential computing that Nvidia GPUs, they had to turn it on for them, you know, to stay in that privacy mode. So, and another thing, you know, this is telling us is if Apple can't do this alone, this also tells us how brutal the frontier AI hardware has become.

24:00 And that's, I think, that's the really the part that I want people to hear, notice here. Even they can't do it. Yeah. Why Nvidia got this slot. It just wasn't raw speed. It also was a security feature that encrypts the data while it's being processed. And I think this is also a big thing. You know, a trend that we're seeing coming in the future, the value in an AI chip won't just be how many operations, as I mentioned.

24:23 It's also whether you can prove the data going through it stays private. So I think those are really important directions that we're seeing. Volkmar, a big final question to wrap this segment up is, do you see this as kind of Apple still trying to do both things, or are they kind of they're just sort of resigned to reality at this point? They're sort of like, we made it, took a shot at this.

24:42 It didn't work out, and we're pretty much calling it quits. I think that if the focus of Apple is really on consumer devices, and I don't see that consumer devices right now can justify the cost and the power consumption of HBM. High bandwidth memory is just very, very expensive. The supply chain is brutal. Apple is big enough that it can control the supply chain.

25:01 But I think that from a consumer device, do you really want to have this 200-watt thing, right? If you look at the zoo of devices, it will only lend itself to the Mac minis or so, but it doesn't fit in the normal portfolio. So I think they will just, it's a very different memory interface if you have high bandwidth memory. It's different inter-poser, etc.

25:21 It's very, very different chip design. And so there's nothing which fits the natural architecture which Apple is building. And so from their perspective, controlling that and actually putting all that investment in just to run a bunch of frontier models, I think they're better off, just go with a partner. So I think it's a very natural choice. Yeah.

25:40 So I'm not surprised at all. I think the main thing is, frontier models won for anything that's complicated, and, like, smaller models go on device, and that's fine because, you know, then memory bandwidth doesn't matter as much. Kaoutar, I think you've actually been— do you want to put a final word in in favor of small models? Because I know you've been a big advocate for that, and it seems like in some ways there's a little bit of a tension here.

26:02 Yeah, there is, you know, it has to be like a tiered architecture. Like, here they have a three-tier system. The easy stuff stays on your iPhone, handled by Apple's own small models. So the on-device pitch still survives here. The medium stuff goes even to Apple's own servers, the private cloud compute running on Apple chips. But the hard stuff, you know, goes to Google, you know, where you're running this custom 1.2 trillion parameter Gemini model running, you know, in Google Cloud on the Nvidia Blackwell GB200 GPUs,

26:28 so not Apple Silicon at all. And I think they didn't have a choice, like Volkmar said, you know, so for all these hard tasks, you cannot really run them with today's hardware, you know, on the on-device. So I'm hoping, maybe, you know, further along, you know, as, you know, the efficiency metrics and also, like, how do you make these models smaller, more powerful, how do you make these chips also more energy efficient?

26:56 Hopefully, maybe we can see a shift back. But currently, with the hardware standards and the way these chips are designed, like Volkmar, you know, well, you know, said it well, it's very hard. Final story of the day. Our producer, Alex, flagged a fun blog post from a writer, Alex Johnston, about sarcasm in AI models. And I guess it will surprise no one who's a regular listener to this show that Alex immediately was like, this is a perfect segment for Chris to focus on.

27:28 The story pertains to sarcasm in AI, and specifically this kind of interesting problem that we're running into, which is AIs are kind of bad at detecting when someone is being sarcastic towards them. And this is a really big problem if you want to use, say, AI agents for customer service, where people get frustrated and are in fact sarcastic. I guess, Chris, I'll just let you kind of riff on this.

27:51 I mean, maybe to kind of start you off with a question is, like, is this a solvable problem? Can AIs get better at identifying sarcasm? And and what do we do about this? Well, I mean, first of all, before we get into whether this is a solvable problem or not, I want to quote the opening of this article, which was, Scotland is known for its majestic landscapes, kilt-wearing bagpipers, and literary icons.

28:19 First of all, as a Scottish person, you know, if people are thinking kilt-wearing bagpipers, that's where we're going in this, and literary icons, that's not what we're known for. We're known for— Scotland is known for that. What do you know? No, we're not. We're known for being drunk. We're being known for bad weather. You know what I mean? We're no World Cup.

28:40 We're being known for being useless at football. You know what I mean? So maybe they're the thing, so I just felt that that was a whole bunch of— anyway. For sarcasm, I think, here is probably the the thing that interested me the most about this article. And I think it's, I think it's this one here, is if I go a little bit further down. So they were saying that, in order— you know, not everything about sarcasm is in the words, and therefore you need to go multimodal for that.

29:09 First of all, I kind of disagree with that. I think you can, with with the use of context, you can really sort of work out whether something is being sarcastic or not without hearing the tone of that, right? So I don't think you need multimodal, but but what I do want to say is, in order to detect sarcasm, it seems like the the world's training data that is used in the models are using the Golden Girls, The Big Bang Theory, and episodes of Friends to work out if it's good at sarcasm.

29:40 Now, I don't want to tell you, but but I think I understand already why LLMs are bad at sarcasm if that's its training set. So I, I don't necessarily think sarcasm is undetectable by large language models. I just think that perhaps The Golden Girls isn't the place you should be looking for that. And that that would be my rant over on this. So, you know, can large language models learn sarcasm?

30:09 Absolutely. The Golden Girls is the best way for them to learn it? No. But why, Chris, you think it's not multimodal? You know, I still feel to detect sarcasm, you need also to look at the facial expressions, the tone, and lots of these things. I still feel, you know, by just analyzing the text, it's not enough. No, I get your point, because I think actually the thing is about context, right?

30:35 So to the point there, in certain cases of sarcasm, you're relying on facial expressions, the tone, etc., which is additional context to the text. What I would argue, though, is that if sarcasm is done well, then the leading-up context before that will give you the kind of the opposite effect. So you will be able to detect the sarcasm just on the response to the question.

31:00 The way these models work by predicting the most likely meaning of your word, and sarcasm is built to do the exact opposite. You say the words but mean the reverse, and you're counting on the listener to catch it. So the model's whole strategy— guess the obvious meaning— is precisely the wrong move for sarcasm. And so, and if you look at the numbers today, most models like land around 60, 70% accuracy at spotting sarcasm.

31:26 And and I think it's fine on maybe neat, next, you know, textbook sentences. But it also can get lost in real conversation. So you can have really these messy gaps and back and forth between people when their people are actually talking. So, of course, I feel like the real bottleneck is the data, not the model size. You know, you need to have good training examples, which a lot of the text that was used to train these LLMs is missing, you know, these nuances and so on.

31:55 But I think adding the tone of voice, the facial expressions, the context, those are scarce in the existing training data. So are you saying, Kaoutar, that, in literally, and literally, there's no examples of sarcasm in any books? No, there are examples, like, but not a lot of examples, probably, you know, mostly like in videos and these podcasts and conversations, but probably in actual textbooks, I don't see there is a lot of it.

32:22 I think there is. I think sarcasm is found in books. I get your point. There is context, and you're going to get more context from multimodal things, facial expressions, etc. I still argue that Golden Girls is not the place to get your sarcasm from, but, you know, there is some, there's some perfectly good Scottish TV shows like Still Game that they could watch, and then they would understand sarcasm a little bit better.

32:45 But what I'm saying is, there's high-value tokens that's not a CV. Exactly. But what I would say is, I feel it's more on videos, right? Not in textbooks. Yeah, yeah. Any sort of book. So it doesn't need to necessarily be a textbook. It can be a novel, it could be a comic book, it can be whatever. But, right, but, but I think it's about context. That's my point, right?

33:03 So, and I agree with you. If all you've got is is the line before and the line after, and if the context and the the difference between the two words is it's not that much, then you might miss that subtlety, and you need additional context. But I think you can get that perfectly within text. I just happen to think, you know. But I get her point. Multimodal will make it easier.

33:24 But but not the Golden Girls. I think sarcasm is all context, and the richer the context is, the better. And I think in that sense, I believe that, you know, we will see that multimodality will play a bigger role the more these models— I mean, right now we are interfacing with them, you know, write some code. Here's the code base, right? And then maybe transcribe some text.

33:46 And I think the moment we are enabling them to see, to hear, and then I think the subtleties will probably be better detectable. And so it will be more precise to do it. But I believe that I can be sarcastic on a text message, and the sarcasm comes through. So I think maybe we are really bad right now in good sarcastic training sets, because, maybe the text messages aren't in the training data, right?

34:13 So we need to figure that out. But I think multimodality is probably the way of making the models more, more aware. What we need to do is go and get some unlicensed content from the internet and train our models with that, because there's plenty of sarcasm within there. So that's the answer. And that answers the earlier question as well. All right. I'm going to wrap this up.

34:35 I'm not sarcastic when I say, Kaoutar, Volkmar, Chris, this is one of my favorite panels. Thanks to all of you for joining. And thanks to all you listeners. If you enjoyed what you heard, you can get us on Apple Podcasts, Spotify, and podcast platforms everywhere, and we'll see you all next week on Mixture of Experts.

🧠 AI Summary

Fable 5 is presented as a major improvement over earlier models, especially for long-term planning, coding, contextual analysis, speed, and spatial awareness. Its main operational innovation is tiered routing, which selects between expensive and cheaper models based on the task, safety constraints, and available capacity. This creates concerns about hidden quality degradation, prompt rewriting, transparency, and who controls model rules. Apple is shifting difficult AI workloads from Apple hardware to Google Cloud running Nvidia Blackwell GPUs because frontier models require much higher memory bandwidth, while confidential computing preserves privacy. Small models remain suitable for on-device tasks. AI sarcasm detection is limited mainly by insufficient contextual and multimodal training data, with current models achieving roughly 60% to 70% accuracy.

🔑 Key Points

  • Fable 5 is described as a step change over previous models, particularly for coding and long-term planning.
  • Fable 5 can cover more files, identify deeper bugs, stitch together context more effectively, and produce less-overlapping spatial code.
  • Fable 5's availability changes from included subscriptions to usage credits after June 22, with a possible later return to standard subscription plans.
  • Tiered routing balances model quality, safety, capacity, and cost by selecting different models for different queries.
  • Invisible fallback or silent prompt rewriting creates transparency and ethical concerns, even when the underlying safety restrictions are considered legitimate.
  • Apple's AI architecture uses small on-device models for easy tasks, Apple private cloud compute for medium tasks, and Google Cloud with Nvidia Blackwell GPUs for hard tasks.
  • Frontier AI workloads favor hardware with high-bandwidth memory, while smaller models remain more practical on consumer devices.
  • Sarcasm detection depends heavily on context, tone, facial expressions, and high-quality training data; multimodality can improve performance.

✅ Actionable items

  • Expose user controls for model fallbacks and make routing decisions visible.
  • Use tiered architectures that keep easy tasks on-device, medium tasks in private cloud infrastructure, and difficult tasks on frontier models.
  • Improve sarcasm detection with richer contextual, conversational, audio, and visual training data.
  • Evaluate sarcasm systems on messy real conversations rather than only short textbook-style examples.

🏗️ Business models

Tiered model routing07:21

A system routes each query to an expensive frontier model or a cheaper, safer model based on workload, safety, and capacity.

  1. Classify the query or workload.
  2. Select the appropriate model tier.
  3. Use the expensive model for tasks requiring maximum capability.
  4. Fallback to a cheaper or safer model when appropriate.
  • Fable 5 routing some restricted queries to Opus 4.8.
  • Matching workloads to different accelerators instead of using the largest chip for everything.

💰 Monetization

Subscription access followed by usage credits One $200 a month plan was mentioned in a report about usage consumption. 04:33

Fable 5 is included in all subscriptions until June 22, after which access requires usage credits, with a stated aim to restore it to subscription plans when sufficient capacity is available.

  • A single task reportedly consumed the entire five-hour usage window of a $200 a month plan without finishing.
Charging for scaled automation 11:31

Automation and large-scale usage are treated as paid usage rather than being covered by a flat-rate human-oriented subscription.

  • Using Claude to automate processes at scale.

📣 Marketing

Branding

  • Apple emphasizes on-device computation and privacy.
  • Anthropic emphasizes responsible handling of dangerous capabilities.

Distribution

  • Fable 5 was made available across subscriptions for a limited period.
  • Mixture of Experts is distributed through Apple Podcasts, Spotify, and other podcast platforms.

🧭 Frameworks

Three-tier AI architecture26:02
  1. Run easy tasks on an iPhone using Apple's small models.
  2. Run medium tasks in Apple's private cloud compute using Apple chips.
  3. Send hard tasks to Google Cloud using Gemini on Nvidia Blackwell GB200 GPUs.
Visible safety safeguards06:27
  1. Identify restricted categories.
  2. Make the restriction visible to the user.
  3. Block or redirect the request instead of silently degrading the response.

🧰 Tools & AI usage

  • Fable 5 — AI model used for coding, planning, contextual analysis, and spatially aware generation.00:49
  • Opus 4.8 — Fallback model used for certain restricted queries.05:47
  • Gemini — Model used by Apple for features and difficult cloud-based AI workloads.18:43
  • Nvidia Blackwell GB200 GPUs — Cloud hardware for running a custom 1.2 trillion parameter Gemini model.26:23

AI is used for

  • Long-term coding and software engineering — Plan for longer, cover more files, identify bugs, and analyze code more deeply.02:28
  • Model and workload routing — Choose between expensive, cheaper, safer, or specialized models and accelerators.07:21
  • Frontier model research safeguards — Limit assistance with model training pipelines, distributed training, AI chip design, and building competing frontier models.06:01
  • Sarcasm detection — Identify sarcasm in customer service and conversational interactions.27:21

📊 Numbers mentioned

Costs

  • Apple Silicon was described as reaching 800GB a second at the high end.
  • Nvidia Blackwell chips were described as reaching a terabyte.
  • A 200-watt device was cited as an unattractive fit for Apple's normal consumer portfolio.

Growth

  • Fable 5 reportedly completed two months of engineering in a day.
  • Sarcasm detection by most models was described as approximately 60% to 70% accurate.
  • Gemini was described as a custom 1.2 trillion parameter model.

Pricing

  • $200 a month plan
  • June 22 transition from subscription inclusion to usage credits

⚖️ Advantages, risks & lessons

Advantages

  • Fable 5 provides stronger long-term planning and deeper contextual analysis.
  • Fable 5 is described as faster than earlier Opus models.
  • Nvidia confidential computing encrypts data on the PCI bus and card while it is processed.
  • Small models reduce the hardware and memory-bandwidth requirements of on-device AI.

Risks

  • A model can consume usage limits rapidly without completing a task.
  • Silent fallback can reduce answer quality without informing users.
  • Centralized model rules determine which answers are allowed or restricted.
  • Sending AI requests to the cloud creates privacy concerns even when confidential computing is used.
  • Models can misread sarcasm in frustrated customer interactions.
  • Frontier AI hardware requires expensive, high-bandwidth memory and significant power consumption.

Lessons

  • The economics of AI increasingly require routing workloads between model tiers.
  • Model quality alone is not sufficient; affordability, safety, capacity, and trust also determine practical usefulness.
  • Privacy-preserving cloud infrastructure can allow cloud AI to retain some of the privacy benefits associated with on-device processing.
  • Sarcasm detection needs richer contextual data, not merely larger models.

💬 Quotes

The race is kind of shifting from whose model is smartest to whose model can you actually trust and afford to run?

Summarizes the shift from benchmark performance toward cost, safety, and reliability.08:08

The model's whole strategy—guess the obvious meaning—is precisely the wrong move for sarcasm.

Captures why next-token prediction struggles with sarcastic meaning.03:07

👤 People & companies

Tim Hwang

Host of Mixture of Experts.

00:16
Kaoutar El Maghraoui

Principal Research Scientist, AI Platforms.

00:27
Volkmar Uhlig

CTO and VP, Data Platforms.

00:32
Chris Hay

Distinguished Engineer and panelist.

00:35
Alex Johnston

Writer of the blog post about sarcasm in AI models.

04:31
Anthropic

AI company associated with Fable, Claude, model safeguards, routing, and a potential public filing.

00:57
IBM

Company cited as benefiting from the shift toward affordable and trusted model routing.

02:28
Apple

Technology company shifting some difficult AI workloads from on-device Apple hardware to cloud infrastructure.

17:58
Nvidia

Hardware company providing Blackwell GPUs and confidential computing for cloud AI workloads.

19:58
Google

Provider of Gemini and the cloud infrastructure used for Apple's hardest AI workloads.

20:47
Citadel

Financial firm associated with a study about token consumption declining as AI costs become more apparent.

13:17
Wired

Publication referenced in connection with Anthropic's response to backlash over frontier research safeguards.

06:31
GitHub

Platform referenced as hosting pages containing prompts for jailbreaking models.

10:49