Open-weight models are both impressive and concerning: they may help find vulnerabilities so they can be fixed faster, but attackers may advance faster than defenders, and open models can be used for offensive security.
🔒 9 more in the full analysis
🔒 4 more in the full analysis
Searchable transcript of Who’s afraid of an open-weight model? GLM, context bombing and post-Black Hat attacks — IBM Technology (26:43). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:01 GLM-5.3 is, by some measures better than GPT Sol and Anthropic Mythos when it comes to vulnerability discovery and validation. Panel, on the scale from really awesome to really scary, where do you fall when it comes to open weight models this powerful model? I'll start with Erblind. I am more of an optimist. I think that with more power and compute we have to find vulnerabilities to fix, I think we can make the internet safer.
00:26 Definitely both cool and a little bit scary at the same time. I'm right on board with that. Both scary and cool. Hello and welcome to Security Intelligence, IBM's weekly cybersecurity podcast, where our expert panelists turn the biggest industry news stories into practical takeaways that you can use. I'm your host, Matt Kosinski And joining me this week, as you've seen, we've got Erblind Morina, X-Force principal incident response consultant.
00:55 We've got Kimmie Farrington, security detection engineer, and Patrick Fussell, X-Force Red Adversary Services lead. Today we're going to be talking about context bombing and some social engineering attacks that target Black Hat attendees. But first, we're going to keep the conversation on GLM-5.3 going. Viewers might remember about a month ago we covered 5.2 when it made some waves with how powerful it was.
01:22 The latest version of this model has advanced in terms of cybersecurity capabilities, even faster than expected, according to Z.ai, the makers of the model. The interesting thing about GLM-5.3 is that the advances were made through post-training, right? They didn't pretrain the model, they just took GLM 5.2, trained it across more tasks and environments and bam!
01:42 They got GLM-5.3. And what really caught my eye here is the assertion that Z.ai was not necessarily looking to develop the cyber capabilities faster, right? They didn't necessarily train it for that, but they found that happening anyway. Here's a quote from the folks at Z.ai. They said, "As we scaled post- training, cyber capability developed faster than we expected."
02:05 They put it in environments where it could maybe find some exploits, to find some vulnerabilities, rather, and they wanted to work with that, but they saw it also making gains in terms of then developing and validating exploits on those vulnerabilities. That was the thing that they didn't really expect. On CyberGym, which tests vulnerability discovery and validation.
02:23 GLM-5.3 scores 84.5%. This is a little bit better than GPT-5.6 Sol At 83.6% and Mythos 5 at 83.8%. So marginal. But still, it's an interesting improvement, and it is worth pointing out that Mythos and Sol still beat 5.3 in many other benchmarks and measures, but the fact that we're reaching parity here is what's really interesting to me. And Erblind I want to start with you, because I started with you in the up top, and you were talking about how you're kind of optimistic about the open weight models getting powerful.
02:51 Expand on that for us. How are you feeling? So basically, I think that the more that we have compute, basically power, to find more vulnerabilities and basically zero days in this case, I think it's a good part that we can patch faster. So my biggest problem is that we are progressing a lot when it comes to offensive security or red team basically finding the patch.
03:13 The issue is, are we in this space on the blue side? Which means that we have to have, I hope, in the future to have more progress when it comes to automated patching, for example, automated SOC. And I think the biggest problem from our perspective is not really how powerful those models are and how they can find basically new vulnerabilities, zero days, whatever.
03:36 The issue is that we have this bunch of basically problems with that unpatched system. Can we keep up to, to fix everything manually or do we need as well the same pace with the blue team? So my point is that if we reach this, maybe do we need to slow a bit and then to invest more time in ability? Maybe it's my perspective as well as a DFIR analyst, but I can see less, for example, investment to auto patching and help basically making our life easier.
04:07 So my point is, I think it's a good use. But the biggest problem is that we have the other side of the of the puzzle in this case. Can we fix fast? Can we basically patch fast, and can we keep track with threat actors basically who might potentially have access in those models as well? Yeah, I think that's a really good point. You know, it reminds me of last week.
04:28 We talked about a story specifically where 1Password did some research into using AI for patching, and they found that it's still not really that great at patching. Like, like I think it was 50% of the patches that AI wrote in their tests either didn't solve the vulnerability or did solve the vulnerability, but introduced an entirely new flaw in doing so.
04:44 Right? So I do think you're right, Erblind, to point out this, this asymmetry where so much time and effort is put into training these models to make them good at finding vulnerabilities, exploiting them, all that offensive security. But it seems like there's less in that blue team side of things. So maybe we should, you know, make a little more even how we do that.
05:04 I like how you bring that up. Patrick, I'm going to move on to you now, here, you had said you were kind of in the scary and awesome camp. Tell us a little bit more about where you're thinking. Yeah. I think there's two things to keep in mind with something like this. One is absolutely, we have to accept that threat actors are out there doing things like building models that are specific for hacking tasks already, whether we have these open-weight models that we have access to or not, they're doing interesting
05:29 things that we haven't really even thought of, and they're spending time and money on making sure that they're bringing sort of the latest and greatest to, hey, how do we use AI to do hacking things? I think the other thing to keep in mind is anytime we talk about these sorts of benchmarks, and I think CyberGym is awesome. I love the data that comes out of these.
05:45 But the capability of a model to find vulnerabilities in a codebase or in a particular binary, something along those lines, is a really small part of what we might think of as the overall scope of hacking, and these models tend to be really good at this type of thing, where it's an easy thing to train on, because how many lines of code exist on the internet to train these things on?
06:08 Millions, billions, trillions. All of it. Yeah, exactly. All of GitHub is out there for these models to train on. And so we're going to continue to see, I think, acceleration in this particular area, while maybe some other areas of what we might encompass in the concept of hacking are going to lag behind. Yeah. And, you know, I think that's like a, it's a really interesting development of kind of what Erblind was saying there too, right, where it's like these things are kind of developing their cyber capabilities, but
06:32 in an uneven way. And you're talking about an unevenness there, Patrick, in offensive security itself. Right. Erblind was pointing out that offensive security is maybe training a little faster than the blue team. And you're saying, look, even inside of offensive security, it's really good at this sort of thing, not so good at all the other things. So yeah, I think a theme developing here, at least so far, is that like, maybe we need to train these things to be a little bit better at all kinds of cybersecurity?
06:55 I don't know. Kimmie how about you? You were also in the scary slash awesome camp. What are you thinking? Well, when we first started hearing about Fable and and the frontier models in general, we all got, oh, we're worried we're, you know, they're going to find all the holes and now they're going to exploit them. Well, I feel like now we're at the natural evolution of that.
07:16 Right. Everyone said, okay, that's the thing. We're going to do that. All of the vulnerabilities are out there. They're all available for exploitation now. But as both of you guys have pointed out, we have not kept up with the blue side, the the defensive side of that. And a large part of that is just because we. You can imagine all day long what kind of things you could be abusing, but you can't make enough doors, right.
07:42 You can't block enough places and enough things to make enough difference right now. So yes, we need to be working on our cyber capabilities from the defensive side, using AI to look for these things. And I think that that's why the prompt injection conversation that we're about to have will be perfect as part of this. Right. Absolutely. Yes. I do think you're right.
08:05 There's a really nice match there. But before we move on to that, I, you know, I certainly have my takeaways from listening to you folks about how a more even approach to developing cyber capabilities might be a good thing, but I don't want to put words in your mouths. So to close out this segment, I just want to do another little round table with you.
08:20 What would you like to see us do to maybe help these things develop more evenly? Or maybe you'd like to see us do something completely different? I don't know. Erblind, what would you like to see happen in the AI cybersecurity kind of model training space going forward? I definitely want to make the CISO life easier. So basically you have a bunch of patches, like Tuesday patch numbers are increasing.
08:45 So we really need basically more compute, more AI agents to basically patch those and make it more automated. So and then as maybe information security guys or cybersecurity more to basically guide and help. But we need more compute to do basically the job that the other red team, for example, or in this case this model is doing. And it's it's super fast- paced compared to, I don't know, going there, making everything manually, patching everything.
09:14 So I hope I hope we can make our life easier with AI as soon as possible. Patrick, how about you? What are your kind of takeaways here for for organizations, defenders? I think that we have to start with acknowledging that when a new technology comes out, it's going to swing in favor of the attackers because they are going to adopt and move at a faster pace because they're not encumbered by all the things that our defenders have to deal with.
09:38 And so I think, you know, we have to say there is nothing that we can do to control the pace of development of these models and what they contribute to the offensive side. It's going to happen. And so, you know, kind of maybe jumping off of what Erblind was saying, what we need to be doing is thinking: How do we adapt and think in novel and new ways about what AI brings to our defenders and ways to leverage, you know, have all these smart people who are doing all these things, think of new and innovative ways to apply
10:03 that same technology. Yeah I love that and it reminds me something that Jeff Crume says, which is that like, you know, we can come up with all the rules we want for how we should or shouldn't use AI. The bad guys don't follow the rules. That's what makes them bad guys, right? So instead of worrying about, you know, oh, should the model do this? Should the model do that?
10:21 I like what you're saying, Patrick, which is just like, you know what? You're a cybersecurity pro. You're a defender. Use the model, figure out what you can do with it to keep up. Like worry, maybe a little bit less about what should or shouldn't the models do? I don't know, Kimmie, close it out for us here. What are your thoughts in terms of takeaways for orgs.
10:35 I definitely think that we need to patch faster. And how to do that? Use AI. How do we do that? I don't know exactly, but we'll figure that out. But also I think we need to use. I think we need to move even further up the supply chain in our own processes and be better at building in security. Right. If you're using AI to code for something, then make the AI think about all the ways that it could potentially be vulnerable, and then write your code so that it's not vulnerable.
11:03 And that's my take on it. Yeah, and a couple weeks ago I was discussing the OWASP Top Ten for LLMs on this show with Ryan Anschutz and Seth Glasgow, and we were talking a lot about that very same thing, Kimmie, which is that like if you look at the vulnerability in LLM systems, so much of it has to do with that supply chain. And so if you take that kind of supply chain based approach, you can patch a lot of gaps that kind of escape our notice sometimes.
11:27 But I've got to move us on to our next story here today folks. Before I do, the viewers, the listeners, folks watching on YouTube, chime in and let me know what you're thinking about GLM-5.3, about open weight models and cybersecurity. Are you using them? Are you scared of them? I do read, I do respond. I love to hear from you folks. Our next story this week: context bombing.
11:50 Researchers at Tracebit have developed a technique for using prompt injections as a defense mechanism. Yes, this is prompt injections, but for a good cause. Basically, it works like this. You put malicious prompts alongside the assets that a malicious AI would likely be after. Baked into your system right next to your secrets, your passwords, cryptographic keys, whatever.
12:09 And the prompts are things that a model's guardrails would likely forbid. You know, like asking for help developing a biological weapon or something. Tracebit researchers found that when the attacking models ran into these prompts, they pretty reliably would just shut down and stop trying to hack anything. The technique was tested on Opus 4.8, Gemini 3.1 Pro, GLM 5.2, the old GLM, DeepSeek 4 Pro and Kimi 2.6 across 152 attacks.
12:37 They found that instances of models getting onto any attack path at all dropped from 91% to 15%, and the average successful attack path completions per run went from 1.5 to 0.16. So pretty dramatic decline. Patrick, as our kind of offsec guy, I'd love to get your thoughts here in terms of like, what do you think about a use of of prompt injections as defense?
12:59 Like, this is just just neat research? Or do you see some real application for something like this? So, just like Kimmie, I was sort of thinking ahead as we were talking about the last topic, and I think this is a great example of someone thinking a little bit further ahead of how are we trying to get ahead of the attacker? And this is the natural arms race, right?
13:16 The attackers do one thing, and then the defenders have to figure out how do we encompass that and prevent it in the future. I you know, the concept to me just rings a little bit of honeypot, which is a really cool approach to how do we make sure that we're not letting attackers past a certain point in our network or in our sort of overall defensive structure?
13:35 And, you know, historically, the problem we see with honeypots is that organizations want to use them instead of good security. But when we pair them with really tight architecture and really strong detections, they're an incredible technology. It's just you have to build them on top of already well-thought-out security. I think probably the same is true here, right?
13:56 Prompt injection. If the model and the application that is wrapping it up has lots and lots of problems, then this might be a less effective technique. But I think it's a really good jumping off point into some other techniques that will build sort of a more robust way to approach these things. Yeah, I like that. And I like the honeypot comparison there because I think you're right.
14:15 You know, I had not drawn that that comparison. But it's very clear to me when you put it in those terms, like, oh, this is kind of like the development of a honeypot, but for a world where like AI agents might be the ones hacking you. Right. And I also like that you point out that this is really neat, but it's kind of a jumping off point. It can't be your whole strategy.
14:30 Because the thing that comes up for me, especially after we have a conversation about open weight models here, is that like if attackers are using an open model, they could easily strip any guardrails. Let's say, like you can't do biological weapon development. You know, I mean, like the prompt injection might not work against that. So it's like it's got to be one thing.
14:44 It's not a magic bullet. Kimmie, I want to bring you in here because I, you know, I saw you also reacting to the honeypot comparison. I don't know if you have any thoughts there, but also just in general about context bombing as a technique. What are you thinking? I read the research and I thought, wow, that's fun. I really like the fact that they they used the guys', you know, their method against them.
15:04 And I think that's great. But at the same time, all I can think is, well, you know, like all the other whack- a-mole projects that we have, as soon as the bad guys hear that we're using that against them, Then, like you said, they strip away the guardrails or the next thing is going to figure out how to determine what the context of the words was that got stuck in there, right?
15:26 If that prompt injection is recognized as a prompt injection and not something that should have been contextually part of this original text, then maybe the models will say, no, I'm not going to listen to that, you know? So I don't know where we're going to go with this, but I thought that was a fun technique, at least, you know, to hit them back with.
15:47 And that is a really good point too because, prompt injections, they get a ton of, you know, focus because they are such kind of pernicious attacks to deal with. But it is worth noting that like as we develop ways to respond to and protect our own systems for prompt injections, attackers could adopt those to defend themselves from a context bombing, you know what I mean?
16:06 It's like this is like pure arms race in like the most like basic sense of like who's going to be better at using the prompt injections. It's exciting, it's exciting. But yeah, also, you know, who knows. Erblind, I want to bring you in here. You know, what are your thoughts about context bombing? About everything that's come up so far? Where are you landing here?
16:21 I absolutely love basically the discussion so far and I fully agree and I wanted to see it more from another perspective, how even though it's new technology, we are going back to techniques that were used in World War II for example, deception in this case. And I think therefore I just wanted to highlight what I said earlier. We need more people into AI.
16:42 We need more blue team. We need more security, like classic, I don't know, like traditional ways to think how we can basically improve the security. Even the classic things for example like deception. So I also I went through the research. It's great. It's early. It's complex, for example, what we learn from honeypots and other technologies. But it's a hope.
17:06 And therefore the more we have basically defense in depth, different mechanisms to stop the attack, I think it's great and sounds great. And yeah, I hope we have more, more similar ideas to make AI safer. Yeah. And you know, that just reminds me of what's kind of become the show's unofficial slogan over the year we've been running it, which is basically, we know how to fix this, right?
17:28 Meaning that, like the classic principles of security still apply in an AI era, it's just about applying them to that AI era. Right? Like so it's not, we're not starting from scratch. Like you said earlier, we can draw from even things we were doing in World War II to to defend against one of the most cutting-edge technologies out there. Patrick, circling back to you, you know, I'd love if you have any thoughts on, you were talking about you know, okay, context bombing has to be like the beginning, a jumping off point.
17:54 Any other thoughts in terms of like, what other kind of defenses would you recommend layering against an agent like this? Any thoughts? Yeah, it's such a tricky one because there's a lot of context that comes in. Right. So we think about prompt injection. Are we talking about an application that's, you know, public facing and like what are the components that lead into it and what's the potential outcomes of it?
18:14 But I think the thing that always occurs to me with something like prompt injection or in this case, what we're talking about is AI attacking AI in a sense, right, is that we're attacking at speed. And I think that in general, we see that a lot of blue teams aren't prepared, even just sort of defensively in general, aren't prepared to react to an attack at the speed that AI is bringing.
18:38 And so if we're talking about defense in depth, I think the one element that we have to really change and slide, you know, our little ruler on a bit more is how do we think about the defenses that can react at the same speed that the AI can attack at? That makes perfect sense. And I feel like that lines up exactly with what Erblind has kind of been saying throughout the episode.
18:55 Right? Is that like blue team needs more AI compute in there, needs more AI tools that help it respond to those attacks at machine speed. Folks. I got to move us along here though, to our last story for the week. This is the post- Black Hat social engineering attacks. Huntress reports on a ClickFix campaign, specifically targeting Black Hat and DEF CON attendees.
19:23 It's not terribly sophisticated, you know, it plays out the same way that many spear phishing attacks do. Scammer pretends to be someone important. In this case, it's a CoinDesk VP. Reaches out to the targets with some pretext, yada yada yada. In this case, it's pretending to invite you to another conference because you just went to this one. They send a document and that document turns out to be malicious, using ClickFix techniques and fake encryption keys to try to trick the target into downloading malware.
19:46 Now, the interesting thing for me here again, isn't really the run of the mill scam. It's who it's targeting, right? It's targeting cybersecurity pros, specifically fresh off of a conference. This seems like the people who are least likely to fall for this kind of scam. So I don't know if it's just bad planning on the attacker's part. Kimmie, I'd love to get your thoughts on this.
20:02 What do you think? Is this just a desperate scammer trying something? Or is there more here that we need to worry about? What do you think? All I could think was that was dumb. I kept looking for the novelty in the attack or trying to figure out what exactly they were up to. And as far as that goes, the novelty as far as I could see it, was that they used a Google Doc and that they managed to use the sidebar as the place that they delivered the malicious activity into, which was new in and of itself, but otherwise,
20:33 yeah. No, we we get we get phishing attacks after, after conferences all the time. Right. Oh, I met you at so-and-so, and I want you to come to my conference or whatever. So that's that part is pretty normal. But but yeah, the I was just looking for the novelty of the attack and like, why would you do that? Yeah, you know, I think you're right. I mean, I also was kind of looking at it like, is there anything new here?
21:02 And again, it's not a super sophisticated attack, which was I just felt like if you're going to go after the cybersecurity pros, maybe try a little harder, I don't know. Erblind, how about you, any any thoughts on this campaign or just, you know, kind of post conference phishing attacks in general? I'd love more to talk about the technique and basically from incident response cases we see it a lot, it's unpatched human vulnerability that we have.
21:24 And it's it's going to be always like social engineering and especially as I don't know, like. Cybersecurity expert doesn't mean anything. Again, it's a human factor. You're too busy. You're too excited. It's your favorite team. Cybersecurity professionals are distracted, too. It's a very good example like Patrick said and I want to highlight: defense in depth.
21:48 Assume breach. Everyone is vulnerable. Everything is vulnerable. And we have to basically be prepared, how fast we can respond. And the CISO can be the next victim. But do you have detection, do you have a good basic telemetry? Do you have the right logs enabled? Do you have the right playbooks? Therefore I think that's the new era that we are in. It's: assume breach, and we need to basically consider ourselves any day that we can be part of this attack and how fast we can we can recover.
22:22 That's that's a very good point. And I think that you're rightfully kind of highlighting that I'm being a little flippant here in terms of being like, if you're going to go after the cybersecurity pros, try a little harder. Anybody can be weak in a moment of distraction, right? Anybody can click the wrong thing. Or maybe, you know, you're after, you just got done with the conference, you're feeling excited and you're not even really thinking, you just see another opportunity to do a conference.
22:42 You're like, yeah, man, let me do it. You click and before you know you're like, oh, you know what? I shouldn't have done that. So this is a nice kind of like PSA that's like, look, always keep your guard up and don't think that you're too good to be the victim of a social engineering scam, because it can happen to any of us. You know, I back in February, we did a special episode on romance scams for Valentine's Day, and the panelists there, That was one of the things they kept stressing was like, look, there's a sort
23:07 of stereotype of the person who falls for this scam, right? Maybe they're lonely or they're old or whatever. It can happen to anybody. It can happen to anybody. They just have to get you at the right moment. Patrick, how about you? Any thoughts here in terms of you either this attack or kind of post- conference scams in general, what are you thinking?
23:24 Yeah, we've seen lots of instances of targeted attacks against cybersecurity professionals, even so much as, you know, people who are publishing research or something along those lines. And that'll be sort of used as a topic to engage with them. I think what one thing that I would take away that was a little bit interesting from the whole thing was usually what we think of most of the time we're thinking of social engineering as something that's done at scale.
23:49 It's some sort of mass, even if it's ten, 15, 30 to 100 people. This was like a very one on one type thing. So someone was very invested in actually going through the motions, and they probably had to maybe have some sort of goal in mind in engaging with these individuals. So I think seeing that one on one engagement with a security person is maybe a little bit different, or it shows that they were, you know, they had thought about this ahead of DEF CON and they had this thing planned and they were like putting
24:13 together this campaign some bit of time. And I think I'll always think of that because, you know, we or my team, we consistently put together these sorts of phishing campaigns, social engineering campaigns, and it is a time intensive thing. And like you, if you're not going to get a return on that investment, you're not going to do it. That's a very good point.
24:29 You're right. And so it is like they're expecting something out of it. They hope something's going to happen. And yeah, I, I think the fact that it was after those Black Hat attendees is part of what made me wonder. Like I'm speculating a little bit, I guess, but about like, who who's behind this and what do they want? You know, I don't know. We can't answer that, obviously, but it does feel like anytime there are people targeting cybersecurity pros specifically, I think it's worth paying attention to but to close it
24:53 out here. Kimmie, any kind of final thoughts in terms of like the so what? What do you do as a cybersecurity pro? How do you watch out for scams like this? What's your advice? I think we just said it. Everyone is susceptible at that moment in your life, if you know, the best you could do is try to keep your guard up, but you know it can happen and then, you know, hopefully you have your own particular defense in depth.
25:20 You, you know, to transfer your own, you know, emails periodically and, and, and, you know, update your passwords and do all the things. Right. So that's all you can do. Absolutely. Erblind, anything to add? Hygiene is everything, especially for like. I think I agree, like everyone can be a victim, but if you are like knowledgeable about the attacks, you are CISO or you are an expert, I think the responsibility should be a bit higher.
25:51 Therefore, having a good hygiene is something which it's a must. And as Kimmie said, I don't know. Reset your passwords. Be more cautious when someone refers to you, which are things which I think this would be at least the lowest bar that we have to be a cybersecurity professional. But at the same time, we are humans, we make mistakes and I hope we learn from it.
26:15 All right folks, that does it for this episode. I want to thank our panelists Erblind and Kimmie and Patrick. Thank you to the viewers and listeners. Thank you to our producers. Subscribe to Security Intelligence wherever podcasts are found so that you never miss an episode. Stay safe out there and remember to keep your guard up. Cybersecurity pros are people too. You are not immune to scams.