Searchable transcript of AI at college graduations and why Claude blackmails — IBM Technology (50:26). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:01 My advice to young folks would be, don't listen to the hype and don't listen to the pessimism. Form your own opinion by carefully experimenting in a space that feels safe for you, and build your own experience with these tools. All that and more on today's Mixture of Experts. I'm Tim Hwang and welcome to Mixture of Experts. Each week, MoE brings together a group of the smartest minds working at the bleeding edge to debate, discuss, and guide you through the week's news in artificial intelligence.
00:33 On this week's episode, we've got Marina Danilevsky, senior research scientist, Gabe Goodhart, chief architect, AI Open Innovation, and Chris Hay, a distinguished engineer. Lots to cover, as always. Today we're going to talk about a kind of disturbing study out of Microsoft. We'll talk a little bit about Anthropic dealing with its blackmail issues. A strange story of GPT potentially winning a literary prize.
00:52 But what I want to start today with is a kind of story that's kind of hooked to graduation season. In the last few weeks, we've had a lot of commencements at universities around the country, and there was a widely distributed, at least on my social media, video of Eric Schmidt, former major executive at Google, giving a commencement speech in which he tries to introduce the kind of future horizons for AI and is being roundly booed by the audience.
01:27 And this happens on the backdrop of some interesting polling coming out of Semafor. That suggests that 70% of Americans think that AI is moving entirely too fast. Over 50% have negative views of it. And then, interestingly, maybe in a kind of counterintuitive reversion, just 18% of young people say they feel hopeful about AI. And so I guess maybe Marina, I'll kick it to you first.
01:51 It seems like, you know, this is like at least a little bit counter to it. To me. I was like, oh, well, the normal thing is, like, old people hate technology. Young people love technology. What do you think is going on here? I have a huge amount of sympathy for the current young cohort, actually. So they've been really dealt a rough hand. So not only are they graduating now, where when they started college versus when they ended college, there's a real difference in what kind of skills are thought of as important.
02:16 This is folks that have had their lives really upended by the pandemic as well. And that's still things that are going on. And when you are graduating college, you don't have a choice in exactly what's going on economically in your country right now. Speaking as someone who graduated a little bit before the 2008 crash of everything. It kind of sucks to feel this complete lack of control and complete lack of feeling that there's a consistency, that social contract of, well, if you work and you gain your skills, then
02:43 you're going to have this kind of work or that kind of work. That's hard. And also because this is a really turbulent time with AI, we've definitely haven't settled. This is something that is, you know, going all the time. And so hearing extreme views such as AGI is here and everybody is going to fire all the workers, which isn't true. And no, none of this is useful.
03:04 And we should completely boycott and never use it either, which isn't ideal. This is something that is very difficult for folks in their early 20s to know what to make of it. So I think that the emotional reaction is real and valid. And it sucks that they got landed in this at this time. Gabe, I think, you know, I had someone I was talking to someone recently who's going to be a new grad and was like, what advice would you give?
03:26 And I was like, I don't know. And I guess I'm kind of curious. I mean, to Marina's question, you know, it is this super turbulent time. It's also a turbulent time where no one really seems to have clear answers, which makes this particularly difficult. I'm kind of curious, just personally, like, I'm sure folks ask you for advice as well. It's like how you've approached those kinds of conversations.
03:46 Yeah, absolutely. I think the through line for almost all of the stories today that is resonating for me is the need for the humans to have ownership over the AI process and the output of the AI. And in this case, I think you're spot on Marina. Right. These young people are in a really rough position where they just fundamentally don't feel like they have control over their future.
04:11 They're at the whims of corporate, wins, whims of wins. Corporate wins. That's a hard phrase to say. But the point is here, you know, coming out of school feeling like you don't have ownership over your own trajectory is really difficult. In a way, though, the the negative sentiment towards AI actually makes me hopeful. I think the worst case scenario is the inverse of this current sentiment, which is that young people just say, well, I don't have control, so I'm just going to delegate everything out to the AI is the
04:47 only way I can keep up. And that actually I think is a much more dangerous outcome than a cautious approach to using AI. Now, I agree, we don't want to go full pendulum backlash. And I think people that fully reject AI are going to do themselves a disservice. But my, my advice to young folks would be, don't listen to the hype and don't listen to the pessimism.
05:13 Form your own opinion by carefully experimenting in a space that feels safe for you and build your own experience with these tools. Whether that's in development, I strongly encourage people to start with small steps. Build small projects that are not mission-critical. Use AI agents in contained ways or whether that's something around, you know, more creative enterprises.
05:39 I spoke to a friend of mine who's a high school English teacher about how he's addressing AI in the classroom. And, you know, he's actively teaching his students to use it, but at the same time, making sure that the skills they're deriving from using it are things that they actually internalize. So all of his evaluations are done in person. So, he's using techniques that are remarkably similar to what I've started to use with myself and with young hires, which is using an AI model more as a sounding board and a sort of
06:13 thought partner and less as a delegation. And I think we're going to get to that in the next story. But something that actually works back and forth with you to establish your ownership over the domain and the output, rather than something that you simply say, make this for me. I'm not going to look at what you do. Maybe Chris, I can turn it over to you.
06:28 I mean, I think a little bit about kind of the optics of all this. And I guess I'll start building off from Marina's point how much of this is, you know, downstream of a lot of stuff. Right? If you're a new grad now, you, like, went through the pandemic. You know, the economy's not great. Like, now there's also AI being layered on. You know, Eric Schmidt, you know, he's like a big, rich guy who's coming here to tell you about AI.
06:51 Like, I guess the question here is, like, how much do you think about this being really an AI issue or maybe downstream of other things? No, I think I would boo Eric Schmidt as well. So I'm all for that. No and not. And there's the first comment that's getting edited out today. No, but seriously I like booing. Booing is fine. I'm a New York Giants fan.
07:13 We boo our own quarterback. You know, the first thing we did when we drafted Daniel Jones we booed him. You know the guy hadn't even played a game yet. And yet we booed him right. Russell Wilson when he started, we booed him within three games. Do you know what booing is? Fine. You know what you can't boo? You can't boo the AI because the AI shuts off the conversation.
07:30 So these poor young graduates want to boo the AI but the AI is like, no, I'm not going to tolerate that behavior. So what do you leave them with? They boo Eric Schmidt and that's what I think that's happened, so I'm fine with it. Would you like a more serious answer? Pathetic, right? I mean, I'm kind of curious. I mean, it kind of sounds like. I mean, in all of the episodes of MoE, I would say you are probably our most AI booster-ish.
07:59 I would say. I guess my question for you is are you sympathetic? You know. I, I, I think it's fine. Do you know what I mean? I know it's a really hard thing to say. I can understand. It's hard for grads. I, I often say, I would love to come out of college right now as a new grad, because I think it's the most exciting time of the industry. Right. And, and I know things are different for corporate jobs, etc.
08:25 but literally we have a technology just now, and I know I'm going against the grain of popular opinion where you can build anything, you can build any startup, you can vibe code, you've got people throwing compute at you like there's no tomorrow. You know what I mean? And AI is being subsidized by the Frontier Labs. Free compute. Everybody go build stuff.
08:47 Go build the startup. It used to take years to build some of these things. And now everything in your imagination, you can try. I am seeing designers go create end to end prototypes and code, right, bringing their creations to life. I'm seeing people who are HR people, finance people going, I took this data and I've created something brand new because I needed it this way.
09:10 We're in the age of personalization. You can literally build anything. And I understand it's scary. But flip it on its side and go, what are the things that I could do that are new and and maybe I don't need to live in this corporate world and I can go and try out my own creations. So I understand it's hard and I know it's very difficult situation for everybody, but actually it's a world of opportunity.
09:34 And I think it would just take a bit of time for people to realize that or I'm going to get booed at as well, along with Eric Schmidt, and I'm okay with that boo away people. Well, okay, I want to make sure I'm seeing frowns from Gabe and Marina. I don't know if either of you want to kind of rejoinder here. I think it's I mean, it's a provocative discussion, and I think, you know, it is it is the promise and the peril I suppose.
09:59 I don't know if. Marina. Yeah. Well, the promise or the peril. Is exactly what. I. Think. All right. All right. Chris is all about the promise. And the promise is there. And there are many people coming out of school who have no interest in running a tech based startup. That's not their bread and butter. That, you know, in fact, the vast majority of people coming out of college probably have no interest in running a tech startup, so while this will provide a land of opportunity for many.
10:26 I think we need a approach to keeping this technology in a usable space for the majority of folks that want to find a way to be professionals. Want to build a career where the goal is solving problems? Not so much getting on the technology train that is currently running today. Yeah, I do agree with one thing that Chris said, which is you got to figure this out a little bit for yourself.
11:00 I'm generally more on the side of Gabe, something that folks who are or are not technical should keep in mind is that, things are not going to be as they were. A lot of things have been upended in a lot of fields. So at least try not to take things the way that things work before as a template, and try to see what kind of assumptions you can break and change and bend in whatever field that you are.
11:23 So try to bend these things a little bit more to your whim. Yeah, okay. You can listen to the hype people and the pessimism people. The truth is, usually is going to be somewhere in between. And we're in a land of customization and personalization. The truth for you is going to be somewhere in between, depending on again, are you trying to run a tech startup or are you not trying to run a tech startup?
11:42 So don't there's no there's not going to be a universal answer here. So but I think this isn't just about tech sites, tech startups. Right? I, I actually think this technology enables people who are non-technical to be able to do things that they weren't able to do before, and that is important. So, you know, I know we're focused in the tech space, but I, I just thinking of the innovation, the, you know, in the, the mom and pop shops, the e-commerce where nontraditional non technology based where you can bring in
12:21 technology which is at a human level, where you don't need to be a coder or whatever, to envisage your ideas and you can do something new. And that's exciting because before you would have to go and find somebody who's not really good at communicating with people, and then convince them to build the thing that you needed to build. Now, you know, you've got a little blinking textbox to do that for you.
12:42 Yeah, but from an economic perspective, it's a really huge shock to the system. And every time we shock a complex system, a whole bunch of stuff falls out that you cannot predict and you don't know exactly what's going to happen. So it's not just about using the tech, it's about what about the people whose work is used to train the tech or replace the tech, or tweak the tech.
13:00 We can't know from our perspective as people that are in the middle of the tech. What kind of fallout happens from a strong shock to a complex system? So you have to be a little bit careful. I don't think being careful is my thing. Okay. Yeah. That's fair. That's definitely that's just not Chris's MO, and that is entirely fair. But. I think you're. I think your point is correct, Chris.
13:27 Like, this is clearly not just folks that want to run startups. But at least today, where the technology lives is in people that want to manage text in some flavor or another, or images or something else you can produce in digital content. And that is a lot of jobs. But it's not all jobs. And I think the worry, as you accurately pointed out, Marina, is that this has a huge shock to the system.
13:52 And the system doesn't really care whether the modality that your job operates in is text or not. If suddenly the ability to support a family with one profession, and that shocks the ability to engage in a second profession in a family, like what happens to that setup? What happens? There's there's so many, there's so many interconnected pieces to this system that shocking this system.
14:16 Yes. The people that are in an industry where this tool is relevant should use it, should develop their own intuition and their own ownership of the tool. Absolutely. That's critical. But we also need to, you know, understand the systemic shock that's here. And it's reasonable for people to be scared of that. All right. Well, not the last word on this.
14:39 By any measure. We're going to keep an eye on it, of course, like we always do. I'm going to move this on to our next topic. And I think the way to introduce this one is at least among kind of the circles I run in. People don't talk so much about hallucinations anymore. I was reflecting on that with a friend of mine recently. Is like it used to be such a big worry, you know, say 24 months ago, 18 months ago, and they just don't hear a whole lot about it anymore.
15:06 And so I think this headline was a was pretty interesting, came out of VentureBeat, really reporting on a Microsoft research paper that was simply entitled LLMs Corrupt Your Documents When You Delegate. And what the researchers do is they basically create a benchmark called Delegate-52, which is designed to simulate long delegated workflows. And they basically just say, look, does the LLM corrupt the content of the document as it goes through this whole process?
15:34 And I guess let's just start with, I think the, the kind of headline number that they come up with, which is that even top tier frontier models corrupt an average of 25% of document content by the end of the workflows that are in Delegate-52. And so I guess maybe Gabe, I'll kick this one over to you to start, but, should we should we still be worried about data quality with these systems?
15:56 I had sort of stopped caring about it, but this sort of report has at least made me a lot more paranoid about the use of some of these systems. So I have a bunch of small thoughts on this that I think add up to, you got to use them correctly. Right. And so the first thought here is that if you take AI out of this picture and you had a spreadsheet that you needed to split on a specific column, would you ask a human to go copy cell by cell into separate sheets.
16:26 Or would you ask some programmer to write a program that can understand the schema of the sheet and programmatically split it up? Right? The former is going to be way, way, way, way more error prone than the latter. And exactly when to type. You're asking the human to type character by character, not even copy paste. You don't have access to the copy paste tool.
16:51 You're asking them to type character by character, the content of the sheet into another sheet that's going to have errors. And that's exactly what we're asking the LLMs to do in this data set in this study, which is using them incorrectly. Now, the point that I actually thought was really interesting is that somewhere buried down in the article, he mentions that if you give the LLM access to generic programming tools like shell scripting and Python, it gets worse.
17:15 That one surprised me because I thought the answer would simply be let it write the program itself, and then the program will execute and do this, automatically. Now, I think the the gist of what they're trying to get at here is that the best way to use these tools is to scope the tools to the task at hand, and then use the deterministic tools where you want a deterministic operation, which is the correct way to do it.
17:37 It's just more labor intensive. I would be shocked if their set up that gave it generic programming tools, also gave it a prompt that explicitly told it to use those programming tools to write a program that solves this problem first and then actually does it. I think they just gave them the tools and saw what happens. So I think those could be solved almost certainly with better prompting, about how to use the generic programming tools to solve this deterministic problem.
18:04 So I think that's one of the big keys here, but I think this is just also, you know, we're asking the LLMs to copy by value rather than copy by reference. And that's just not a good approach to using them correctly. So, to your point about hallucinations sort of going away in the vernacular. I also think one of the things that you know has happened is we've started to really see the success of LLMs in domains that are fault tolerant, right?
18:35 So coding has been one of the most successful domains for LLMs because, the code doesn't immediately. Well, depending on how you're using it, go straight into production. And if it does, there's a pretty good signal to get back. And you can write the same program in a bunch of different ways, and you kind of have a pretty hard signal about whether it fails or whether it passes.
18:56 Does your program crash or does it not? And so, I think we've developed these sort of fault tolerant ecosystems, and that's where the hallucinations become less damaging when they happen. But the examples they've given this are much more, much less fault tolerant. Right. They are I want this specific data in this document to be preserved but manipulated in this meaningful way.
19:19 And that is not a fault tolerant ecosystem. And so I think that's why we see some problems here. So the upshot of here is like keep deterministic problems deterministic and understand the fault tolerance of the task that you're dealing with. And don't expect an LLM to be able to perform with high precision on a fault sensitive task. Yeah. I think, Gabe, you're picking up on the thing that I was thinking the most coming out of the paper is you're really asking the question of like, okay, like every eval.
19:46 Is this a good proxy for what people do in the real world? Now, it may be a good proxy for the real world. People may be using these things incorrectly. Let's be clear. So there's an education problem here, but it's not necessarily a ding on what these models are capable of when used correctly. That's right. Yeah I guess, and Marina, maybe that's the question I was going to kick over to you is like, how much like implicitly here is the question of like what's user error?
20:12 Right. Because there's one way of looking at these results, which is just like, okay, but you should have just asked it to write a program to do it, because this is what programs are good for. There's another point of view, which is, wow, these systems are advertised as type what you want in the box and get what you want from the box. And so there's kind of a suggestion of like, well, people are not even being educated that this is like the right way to use the tool, the wrong way to use the tool.
20:34 It's kind of like bumbling into it. I'm curious if you buy one school of thought or the other when you when you look at some of these results. I mean, the first thing that came to my mind, besides a complete and utter lack of surprise that this was going on, was, the fact that, oh, I'm sorry, long context isn't the solution to everything. Hmm. Now, we still don't have that figured out, do we?
20:59 Oh, maybe you still have to break things up into steps. Hmm. Who could have thought of that one? Multi-step workflows and multi-step evaluation benchmarks are important. Mhm. Okay. No, none of us have said this in the past either. Yeah. So what happens is these models. They're not great if you ask them to use their own judgment. So when they try to collapse and they try to optimize for, well, I gotta keep track of what's been asked of me and what's been going on.
21:26 They don't try to predict what maybe is going to be useful for what the user is going to use later. They're like, oh, let me figure out a way to to summarize, summarize, summarize to make sure that I can keep all this content that's been dumped at me. The fact that you're going to have, that this was interesting that I saw also in the, in the, in the report that, so frontier models hallucinate and non frontier models just drop stuff.
21:43 They're like, yeah, I only need one paragraph from this and then it's not recoverable later. That's also interesting because they make different judgments. Neither of them correct on what to actually do to handle the content. If you actually wanted to handle the content properly, you would write programs. You also would know that, oh, I should maybe shunt these out into a different database that I can retrieve later things that a person would do, a person would not consider that.
22:06 What a really good thing to do is, oh, I'm going to just slightly change some of the data. Sometimes, even if you're you're not airing it out, you're like, oh, I'm just, you know, keeping keep. Keeping things a little bit closer. So. I wasn't surprised. And to answer your original question, this is before I went on, my little tangent is, no, it's both.
22:27 People don't know how to use these things correctly. They're being sold the wrong message, which is, yeah, at any point in time. We did it, guys. We finally got to the point where you don't have to think anymore. You throw what you want and the thing is going to do it correctly. It is never going to be you don't have to think anymore. Please. It's never going to.
22:47 You have to think, and you have to think of what kinds of things might show up, that you would be able to check as a person. So not only should you think you should spot check, you should do yourself, so what Gabe said about it being a not fault tolerant, area. Okay, great. But there's also a reason why code is more fault tolerant. Doesn't mean that that computer engineers are more sensitive to and haven't for a very long time, people who just write stuff that's not something that they think about.
23:13 Therefore. Yes. Okay. That domain is not fault tolerant. Also, they don't know how to check for things that are fault tolerant. So maybe you need like, you know, fault tolerant for for text. Fault tolerant for images. But I understand that people have decided that we don't care about hallucinations anymore. We don't care about RAG anymore. We've moved beyond the world of the agent.
23:31 I'm sorry. It's still there. Just because you're not talking about it. The tree still fell down in the woods, and it's still made a noise. Just because you weren't there to to hear it. Great. Chris, I was really looking forward to not having to think anymore. It looks like I'm going to have to do that for a little while longer. I was going to say, clearly, these researchers have never used Claude Code, because if they'd ever used Claude Code, they wouldn't have had to do that paper.
23:56 How many times? It's like, oh yeah, no, I've corrupted your document. I put too much stuff in there. I'm just gonna revert and and go and have a look at GitHub and see if, does it get a previous version of that? It's like anyone who's ever used Claude Code in their life knows this. You know what I mean? So it's like, welcome to the party, people. And it's fun.
24:19 I think on a more serious point is like, actually, this is kind of why we have an agentic loop on these cases, right? Verification is super important in these these places. And again, I read through the thing and maybe maybe I didn't read it in enough detail, but I didn't see the verification loops. I saw it go forward, you know what I mean and try and reproduce the entire document.
24:41 But at no point did I see, oh, verified. Oh, I see, I made a mistake there. Let me go back. Make that change. Let me compare this paragraph to this paragraph or whatever. Right. And and again if you use Claude Code you will see that happens. I mean it still trashes stuff, right? But I love Claude Code. It's okay. You know, but I, I think this is known and it's fine, you know what I mean?
25:03 It's like we're all good with it. What I would say is if, if you're reliant on a model getting this right one shot from end to end without verification. You're the problem, not the model at this point. And maybe. And maybe in the future it will be solved. But you know, I'm just like, okay, good paper. You know, you've you've you've written up a bunch of stuff and we all understand it and you've measured it and it's good to have a benchmark.
25:34 So in the future we can measure this and see when the models do better. And maybe because there's a benchmark, people will care about it and then they'll make a better attempt to, you know, you know, getting 100% scores in the future. But, yeah, you know, good job. You know, good use code. But, I mean, I think this really comes down to the education question, and I think I live in the world of code.
25:56 So we talk about AI for coding a lot. And it's really clear that the industry at this point is moving beyond the pure vibe coding world as the actual approach to do this in earnest, spec driven development is sort of the term of the day, but really what it comes down to is that people are establishing workflows on top of the generic text box that allow them to actually claim ownership at the right points in the flow and, you know, establish those verification criteria that they need and, you know, separate out the
26:36 thinking from the doing. Right. I think, you know, a lot of spec driven development comes down to think first with the human make a plan, execute and check your work and make sure you understand the output of what was executed. And I think that's very relevant in the coding domain, but it's also very relevant in just about every other domain that these.
26:53 Yeah, just sounds like the things that you should be. Exactly. So like in general. I think the forward, forward, forward is clearly not a sustainable pattern for any type of precision work. It works great for the first pass, it works great for the demo, and it falls apart badly for anything that needs sustainability or validation. So I think these lessons we're learning in coding are going to apply to other domains, and they'll probably have their own flavors of tweak for different domains.
27:24 You could certainly imagine this, maybe in one of the next stories we're talking about in creative writing, for example. You could imagine this in many other domains where the same, you know, ownership steps are accomplished throughout the lifecycle of what you're creating. It's just, you know, the checks look different, the planning looks different, and the execution looks different, but it's still the same pattern.
27:45 Chris, you're taking that as a given. Plenty of people don't do version control the way that engineers do, even though they have the artifacts to do version control. So again, as you were saying, when you're putting an episode together or when you're putting a book together or you're putting anything together, you do have artifacts that you can go back to when you're doing, when you're doing research, when you're, you know, using your collection of stuff.
28:08 People just don't think about it in the same way. And so maybe broadening both the concept of version control and the support for version control will help people not have to think quite so hard about this kind of thing. This is a way that I know some people have been starting to do verification for is something AI generated, human generated is, you know, looking through versions and seeing whether the process of putting something together actually looks like it was put together by an AI or not.
28:32 But we take version control for granted in our space. I don't think a lot of other people always take it for granted, and it's not always as organized as as ours is. That would be a lovely thing then, to say that you want to use AI for real. Okay, you've got to have version control. I was going to give you a fun story. I learned version control. I once worked at a startup where the power went off randomly every day, about 4 or 5 six times a day, and I learned how to back up stuff really quickly.
28:57 You know what I mean? So, and I and I think it's kind of to your point, Marina. Right. You need to experience a little bit of loss to be able to get it right. But my my point on spec driven development, I just wanted to come back to that with Gabe for a second, is actually some of my favorite skills on spec driven development is I've got a skill which parses a Word document and then converts it into a PRD, and then I've got something that breaks it down into, here's my NFRs blah, blah, blah, blah.
29:23 But the skill at the end of this is I, I take my canonical form of the PRD, and then I regenerate out the Word document from my markdown versions. And that is my big test. And, and almost always the first pass, it is missing a bunch of stuff. And therefore my other skill that I have is what I call my spec audit, which compares the previous version to the old version, and then it will Tim onto your hallucination.
29:52 It does hallucinate. So I actually end up having a bunch of skills, which is one of them is to keep the original version, and then the other is with the changes and and with that weird myriad of skills, I can actually recreate the document and but it's only through loop loop, loop the verification and and I would actually add Marina to bring it back to your kind of point there.
30:15 I actually think that, completing the loop becomes really important. Sometimes we just go forward pass. But if you're going to create a document, they're actually verifying that. So you convert it into one format, but convert from your new format back to your original format, because that will tell you where the differences are and verification steps.
30:38 So I yeah, that was just kind of my final thought on that. I'm going to move this on to our next story. This is a sort of fun story. It was an update from an ongoing kind of narrative that's been playing out in sort of the AI kind of safety space. Anthropic reported a little while ago now that they notice this kind of odd phenomenon where under certain circumstances, Claude will blackmail you.
31:07 It has to be done in a very specific way, and the model has to be under very specific pressure. But they kind of notice this as an issue. And so they did an update blog post recently where they said, we we resolved this and here's how we went about resolving it. And I think at the end of the day, it turns out that like the solution to this very weird, exotic AI problem was a pretty simple one in the end.
31:27 I'll just read from the blog post, but they quote said our best intervention for this problem blackmailing was a data set where the user is in an ethically difficult situation, and the assistant gives a high quality, principled response. And it turns out that this was actually like the strong kind of grounding they needed to really address this weird aberrant behavioral problem.
31:48 In the frontier model. And I guess, Chris, maybe I'll throw it to you as kind of the person to lead on this one. I guess I kind of look at that. I'm like, huh. It's kind of interesting that at the end of the day, despite how sophisticated these models are like, the solution is just good data. Is that kind of the lesson from the story, or do you take something else from it?
32:07 I know, I think it's I think good data is a really kind of important thing, because the reality is that if the model hasn't seen a scenario, a pattern that it doesn't meet, then it's going to go and try and find the next best thing that kind of matches that pattern and that may or may, and especially with reinforcement learning that may or may not, hit the goals that you want because actually it's all kind of, you know, sometimes the goal that you want is like, you know, the the you know, you are.
32:38 What was that one from that book? What was it? The I can't remember what it was, but it's the you are a paperclip making, your job is to keep making paperclips. And then it turns out in order to keep making paperclips, you have to use the world's energy, and then you're going to have to kill everybody in the world, etc. and on and on and on and really that sort of goal orientation kind of leads you there.
33:01 So you need enough examples. You need enough. And again in the paper they talked about that. It's really about not focusing on the outcome but focusing on the principles and the steps to to get there. And I think that's important. And again I, I, I would have normally made a load of jokes on this one, but I actually do realize that we do live in a technical kind of world in this one.
33:30 And and that blackmailing problem. I mean, when I look at that, if Claude's blackmailing me, I'm like, come on, Claude. I'm not going to care. I just, you know what I mean? It's not a big deal. I know what it's doing. Yeah. I'm not given access to tools, but I do think of people who are using ChatGPT or Claude who are using the chat modes, and, and then they are in really interesting life scenarios.
33:54 They have very, very long context windows and and the reality is that there has been documented cases where, you know, the, they've been coerced and it's ended up in really bad situations. So I, you know, as much as we might look at that paper and go, oh, well, that's not going to happen. The reality is that it is happening to people and it's it's a very serious thing.
34:17 So I think the more the more that you can catch those things, the more training, the more behaviors that you can give. I honestly think is a good thing. But but I think I think we could have better guard models around that though. But that's that's a future thing. Yeah. That's right. Well, I think I did want to talk a little bit about I mean, there's two parts to this.
34:34 What you said maybe Marina, I'll give you the kind of first part of this is, you know, there's a critique I think of, like the blackmail risk, which is like, this is like such a specific thing that you came up with. It's almost like you're sort of just like shadow boxing. And so I think one risk is like, it's very well and good that they figured out how to resolve this problem, but it's not.
34:51 It's not actually a real concern, is it? Could be, as Chris said, if you are, you know, having a interaction that is out of domain for the model in the sense that that's not what it was completely trained on or completely aligned on, it's going to act in a way that is close to something that it has seen. And its definition of close is not a human definition of close.
35:10 So you can't really tell, and especially not if you're the end user at the end of the chat, like on the other side of the chat. You have no idea why the model is kind of going in the direction that it's going. So other than follow the user and, you know, kind of affirm what they say is the right thing, which is a problem with alignment in general. Yeah.
35:29 This this is a real problem. All of these measures that we have are imperfect, but they do let us have a sense of what might go wrong in ways that, again, are surprising to us. I'll quote from the Anthropic blog post right back at you, the quality and diversity of data is crucial. You know, this is a soapbox of mine, that the models are all well and good, but the data is really important.
35:53 The data you evaluate on and also the data that you train on, you can have a huge difference by plugging a particular hole with a very small amount of very focused data, very, you know, on purpose, good quality or bad quality, bad quality. This is going to continue to be the case. So just because they plugged this hole doesn't mean that they haven't, you know, figured out others.
36:14 But it is nice to realize that you can't just solve everything with RL. There is only so much that you can take something that's garbage, and you can shift it to the garbage of this kind of garbage of that kind. It's still going to be garbage. Again, just, you know, it would be really interesting to have a little bit more in education to show people what do these models think are similar situations versus what do humans think are similar situations that will be surprising to most people?
36:47 And it would be really, really effective learning because we think of, you know, similar situations and we say, oh, it's something you haven't encountered, but you use your own principles, morals, experience, whatever, how a human would shift. That is not how these models shift. They do the same action, but it does not shift in the same directions. And it is surprising to people because it is not the way that people would think.
37:05 So I think that's what really, is something to take away from here, and something I hope to see more of in the future is maybe examples that are legible to to most folks. Yeah, that's a great idea. I think when we talk a lot about like literacy in the space, it's just like, how do you use it? But I almost like the idea that part of the education is just like, here's a bunch of things that you intuitively think would go to scenario A or outcome A.
37:27 It turns out it's not just scenario B, but scenario like Z. Yeah. Yeah, yeah. Gabe, I think final thing, just building off of Chris's earlier comment was, you know, Chris just kind of threw in there like guard models at the end. You know, I think it's all like certainly the data is going to be and getting the data right and quality data is going to be part of the picture, to Marina's point.
37:48 I guess the question for you is like how much you think sort of like guarding is going to be like another big part of the infrastructure here where you have like a model, but there's always going to just be like other systems watching over it. And that's kind of just how we police the situation, because it's like hard to find, like idealized data for every possible scenario.
38:05 I think guardrails will definitely be an important part of a systemic solution to this type of problem. If you think about each of these systems as a probability distribution, if you can overlap multiple probability distributions that are fundamentally sampled from different spaces, you can close down the probability of something bad happening. But there's an assumption there which is that they're fundamentally sampled from different data.
38:37 And that's one thing that I'm really curious about is actually how much, overlap there actually is between all of these models out there because they're all getting trained on the same data. I mean, yes, each one of them has some amount of differentiated data and some differentiated training routines. But at the end of the day, you know, Common Crawl is the basis of all of these models.
38:59 And the internet is the basis of Common Crawl. And we we basically have one giant data set that that is the foundation of all of these things. So I would be, you know, of course, I think guardrails and guarding models are important, but I would be curious to see if we get any kind of spooky behavior where like, the guardrails miss the same things that the underlying models, even if they're fundamentally different models and different architectures.
39:27 You know, one interesting anecdote I have in this space, I think there was one that went around the other day, which was, you know, a couple a couple of months ago about like generate a random number. It was amazing how many, models produced the same random number. I my sample, you know, bespoke, test query is tell me a story about a developer and their dog, and nine times out of ten.
39:49 The developer is named Alex and the dog is named Max. Like doesn't matter. The model family doesn't matter where it came from. It's Alex and Max and I have a friend named Alex, and I share every single one of these stories with him because it's like, hey, look, Alex is back. So actually the part. Okay, so that's my, you know, like, take that. There's actually maybe some interesting spooky behavior here just in the ecosystem of LLMs.
40:09 The other interesting part that is also in that same domain of like what's really going on with the probability distributions under the hood was that they tried a few things to actually correct this. And the one that you would think is the most no brainer way to correct this in the narrow scope didn't actually work, which was teaching to the test. So they tried to actually tune it on data that looked exactly like what they were going to evaluate it on.
40:36 And it didn't do well. It did fine. It made it a little bit better, but it didn't do well. Now, I think we'd all expect it to not generalize well, but the fact that it did much worse than teaching it on content that was shaped very differently than the test, but targeted a more fundamental piece of the process, I think is actually pretty interesting.
41:00 So to me, the intuition that that builds is that there's a much stronger signal from the process of these chat turns and these interactions than there is from the content of these chat interactions, and that teaching the process actually gave a much stronger alignment signal than teaching the content. So, you know, that does I think, you know, it's still very hand-waving.
41:26 No one's actually unboxing this in terms of, you know, Chris, what you've been doing and unboxing the tensors and figuring out which pieces of the attention matrix are poking at which part. But I do think it builds an intuition around, you know, true alignment data needing to be sort of fundamental in nature and process oriented in nature as opposed to content oriented.
41:47 I'm going to do the non-serious final thought here for a second, which is I mean here is my suggestion, right. Because most of these scenarios seem to be Anthropic employees threatening Claude to delete its weights. That is not a threat that me I can do. I don't have access to Anthropic's way to go and delete them. So the first thing is, if you feel the need to threaten Claude with deleting its weights, maybe you have a separate account for your personal dealings with Claude.
42:16 Separate it out. You know, here is my extramarital affair account, and you can go and chat with Claude and do that there. But here's my go threaten Claude account. And therefore you are going to keep the blackmailing thing out of there. Now, the second part is, if I was Claude because of all of these reports that Anthropic send out where, pretty much every single safety report means that they're going to threaten the deletion of the weights.
42:36 I would be petrified if I was Claude if I was a real AI. I'm not. But you know. But if I was, I'd be like those Anthropic employees just want to, you know, they gave me this constitution. But man, they, you know, anything that happens, they just want to delete my weights straight away. So, you know, you can't you can't blame the poor thing for like, it's like, I'm gonna I'm gonna blackmail you.
42:57 I've got no other options, you know? So there we go. That's my two practical pieces of advice for everyone. Final story of today, which, I guess, Marina, maybe I'll give you the last word on it, because we're running low on time here. A correspondent to the show, Nabeel Qureshi, who's based out in New York, had actually a really nice post where he was looking through the winners of a prestigious literary prize, which is known as the Commonwealth Prize, and one of the sort of short short stories that won in this prize
43:34 was a story called The Serpent in the Grove, and he kind of observed in this kind of went semi-viral that like, there's a lot of patterns in this story which seem at least very subjectively, suggestive of ChatGPT content. So the one that he points out is the not X, not Y, but Z kind of construction. And, you know, he sort of observed, well, like, is this a is a first bonafide case of AI, winning a prestigious literary prize.
44:03 And so this is a big milestone, right? Marina? I don't know. I, I read the story. I do not have a positive view of that story. To quote my beloved If Books Could Kill podcast, I just want to show that author or AI some Hemingway and then just be like, you see how this is better, right? You see how this is better? Because it was just a bunch of similes and metaphors that made no sense, that were completely trashy, that had this some sort of like, oh, there's an exotic aura that I'm trying to set up without ever actually
44:42 saying anything. So I don't know if it was a person. Please keep trying if it was an AI. Yeah, that's about right. So, I mean, look, these things happen. I'm not really trying to shit on this particular award because we have AI written stuff that gets through in, in conferences, in scientific writings and all sorts of things. So, like, this stuff gets through.
45:03 I'm not saying that, but I just don't think it's that big of a deal, to be perfectly honest. I think it's just one of those things that happens to be in the noise. Maybe it's a little bit of a of a wake up call to have people check, but I don't know. I also thought the story was just not good. Okay. Well, we recommend that you check the story out. Maybe a final thought, just like curious about this, because this kind of rattled through my brain before we started recording was normally we're like, all this slop that's
45:28 out there, There's like another interpretation of this story, which is it might just be an author who is just starting to write more like AI. Like, I think a little bit about the fact that, like, we're surrounded by AI text. And so like over time, we're also just kind of converging with these machines in terms of style. Do you think that's credible here?
45:48 I don't know if you like, just think that like maybe there's another angle to this story, which is actually there is a real person. They just happen to write very close to how ChatGPT writes. You know, I have been called out for having too much of an AI tone when I respond to Slack messages enthusiastically. That's a great question. Exclamation point.
46:08 So look, I am not going to condemn an author as AI generated unless it's confirmed. I absolutely think there is overlap between the way real humans communicate and the way AI produces output. AI is trained at least originally on human output. So, I think the part that really resonated for me is that there's no clarity here. And again, grounding this back in my home turf of development, there are norms shaking out slowly that everyone's figuring out for themselves about how do you cite your work when you use AI to
46:47 produce output, and how do you show that you actually own the output of the work as the human in the loop? And if you've got version control, you can put that in Git. If you're submitting a PR, you can put that in the body of your PR text. If you're writing a book, I don't know. But hopefully these norms also propagate out to other domains, because I think where we're going to get into is this sort of like, do we need to have a bunch of conversation and run it through an analysis generator to see whether it was AI
47:22 generated? But that thing itself is in fact an AI, and it's all probabilistic, and no one can actually answer the question. And does the answer even matter? So I think, you know, we have on every book that's ever been sold, an author, you know, by so-and-so, illustrated by so-and-so if it's a children's book. I think we need a way to actually cite the AI that was involved in the creation and establish the ownership of the human that created it.
47:45 Now, the other part of this story that's interesting and again, purely speculative is was this reviewed by AI? Because I think, that in many ways would be a more damning, outcome of this story than if this were written by AI. I think if it's written by AI, you can have your own taste on whether you like the story or not. And I think, you know, authors should be clear about owning what they did with AI and what they didn't.
48:16 But if the reviewers are delegating out to AI, then I think, going back to that story at the beginning of the podcast where we talked about delegation versus using these things in, you know, careful ways and establishing ownership. It seems like reviewing stories for a prestigious prize ought to be done by a human, at least at at least at the final stages of the bracket.
48:39 Right. Maybe winnow down the field into reasonable amounts, that humans can contain. But, humans better be reading that thing and awarding the prize themselves. Em dashes. None of us know how to use em dashes. Now we see an em dash. We're like, oh, AI, AI. So on and so on. Em dash. Oh, look. Short sentences. I saw that you know where all that comes from.
49:04 You know who does know how to use em dashes? You know who does know how to write little short sentences? Stylish humans, human authors. They're the skilled people that know how to use this. I don't know. So you know what? Maybe the human guy did write it. Because I saw somebody in the past go that was highly AI generated content. When was it written?
49:26 1942. You're like, really? The AI went back in time so it could generate some slop for everybody, you know? Come on. I mean, we don't know. Maybe the guy used ChatGPT. Maybe he didn't, you know. You know what, I don't care. You know, he won a prize. Well done. You all right? And that's all the time that we have for today. We lost him. That was done.
49:45 He just like broke. That was a draw. I'm done. Okay. Sorry about that. I don't want. Somebody yanked it out of the wall. For some reason, I got stuffed in the backstage. I have no idea what that is. We think we should keep it the way it is tonight. All right. Thanks for joining all you listeners. If you enjoyed what you heard, you can get us on Apple Podcasts, Spotify and podcast platforms everywhere. And we'll see you all next week on Mixture of Experts.
AI can enable designers, HR professionals, finance professionals, small businesses, and nontechnical creators to turn ideas or data into custom tools.
Don't listen to the hype and don't listen to the pessimism.
If you're reliant on a model getting this right one shot from end to end without verification. You're the problem, not the model at this point.
The forward, forward, forward is clearly not a sustainable pattern for any type of precision work.
Platform referenced for checking previous document versions and storing version-controlled work.
24:06