← All transcripts

AI Hackers Are Faster Than Your Pen Test — Eli Cohen, Snyk Transcript, AI Summary & Key Points

AI Engineer · yesterday · Science & Technology · 19:29 · EN

Watch on YouTube

Answer

AI hackers are faster than annual pen tests: the average successful AI attack takes 24-34 minutes and the fastest took 4 minutes, so pen testing once or twice a year leaves ~350 untested days — security must shift to continuous AI pen testing on every code change.

AI Summary

Security has to move to continuous offensive testing in an agent-first world: developers now ship 218% more lines of code per developer, 62% of LLM-generated output is insecure or broken, and AI-powered attackers chain low-severity vulnerabilities into critical ones in as fast as 4 minutes (average successful AI attack: 24-34 minutes). Static scanning (SAST) misses runtime and authorization issues, dynamic testing (DAST) misses business logic, and human pen tests are expensive and happen only 1-3 times a year — leaving ~350 untested days. The answer is continuous offensive security: AI pen testing on every PR, agent red teaming for multi-step attacks (prompt injection, data exfiltration, goal hijacking), and a team of agents — orchestrator, recon, vulnerability hunters, an exploit-validating judge, remediation and reporting — fed with context from other scanners. Snyk Evo is the product implementing this. Four questions to ask any AI pen testing vendor close the talk.

Key Points

  • Developers produce 218% more lines of code per developer than before, and more than 85% of developers use coding agents
  • 62% of LLM-generated output is not secure or is broken, and it is reaching production, growing the security backlog faster than it can be remediated
  • Attackers use frontier models to chain low-severity vulnerabilities into critical ones, attacking especially where application context creates vulnerabilities
  • One operating model identified more than 600 vulnerabilities and breached 600 firewalls in more than 555 countries
  • The average time for an AI attack to succeed is 24-34 minutes; the fastest is 4 minutes; 43% of all MCP servers have vulnerabilities
  • Static scanning (SAST) is cheap and runs on every change but misses runtime issues such as authorization and configuration problems
  • Dynamic testing (DAST) finds runtime issues like broken object level authorization (BOLA) but cannot reason about business logic
  • Human pen testing is the golden standard for business logic but is expensive, slow, and typically done only 1-3 times a year, producing a PDF report with slow retest cycles

AI in practice

Used for

Agents

  • Snyk Evo AI pen testing agent team — Continuously penetration test applications on every code change and PR, find vulnerabilities including business-logic flaws, prove exploitability, and produce actionable remediation and reports. 2 held

Tools & resources

2 items

ENo. 5307
AIAINotes.us AI product

Evo by Snyk

snyk.io/evo/

Evo by Snyk is a security and governance platform for AI-native applications, agentic development, and AI systems. It provides continuous visibility, governance, testing, and real-time control across the AI software lifecycle. Its Agentic Development Security capability discovers models, agents, MCP servers, datasets, plugins, tools, and services used by agents across repositories; generates an AI bill of materials; applies policy and guardrails to agent actions; and validates AI-generated code before it reaches repositories, pipelines, or production. Its Continuous Offensive Security capability autonomously stress-tests applications and AI systems like an adversary, with agents coordinating reconnaissance, vulnerability hunting, exploit validation, remediation, and reporting, and an independent validation model intended to confirm exploitability before reporting findings. Testing is intended to run on code changes and pull requests.

Mentioned in
2 videos
Kind
AI
SNo. 5309
AIAINotes.us AI product

Snyk

snyk.io

Snyk is a security company and developer and application security platform that secures AI-generated code, development agents, and AI-native applications from development through production. Its platform provides continuous validation and integrates with IDEs, CI/CD pipelines, and AI coding assistants, while addressing automated AI attacks, untrusted agentic development, and ungoverned AI applications. Snyk also presents continuous offensive security and AI-powered penetration testing as part of its approach to finding and validating vulnerabilities as code changes.

Mentioned in
1 video
Kind
AI

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of AI Hackers Are Faster Than Your Pen Test — Eli Cohen, Snyk — AI Engineer (19:29). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:12 Great. Hi everyone. We're going to talk today about continuous offensive security. Okay. I'm going to speak quite soon. What does it mean? Uh but before that, a little bit about me. I'm very excited being here today. I'm Ellie. Uh used to be co-founder and CEO of Helios, a runtime security company. got acquired by sneak and in the last two and a half years I had the privilege in working in sneak primarily on product management for AI products but more recently founding field CTO meaning I'm working with developers and

00:41 with security folks to better understand and to help them shape their strategy around AI security. So how do you take your AI agents and making sure they are um secured? How do you work with your agents in production? Um, back there, can you set up the timer for me, please? The timer. So, I'll know the time. Great. Okay, now we have the timer also. Great.

01:08 Um, so just so I'll get to know a bit better. Please raise your hand. Who is here more from the development side of the table? Developers. Okay, great. From the security side of the table. Yeah, I already know you. Great. and other people like sales, marketing, other people running agents in production that are not from developmental security. Great.

01:30 Okay. So, we have everyone here. So, I'm going to tailor this uh session towards everyone. So, obviously this is an AI engineering conference, right? So, we already know that the amount of code generated today is incredibly bigger than it used to be before. Okay. According to our data, it's 218% more lines of code per the developer than it used to be before.

01:56 Okay. And more than 85% of the developers are actually using coding agent now probably running here today. It's more than 100% right because I guess that each one of you probably have five different session of cloud code running simultaneously. Otherwise, you wouldn't be here today. And the funny thing about those agents running in production and generating fork for us is that it's far less secured than we are used to.

02:22 Okay, so 62% of the output generated by LLMs is not secure or it's broken. Okay, so just imagine this data combined together. It's meaning that we are generating more code than we ever did and this insecure code is getting into production. So that's a huge challenge for us as a community and what it actually means is that our backlog of security issues is becoming longer and bigger and much more challenging to remediate it.

02:55 Okay. So we are actually creating many more issues at a pace that we simply couldn't remediate or mitigate. And not only that, AI is helping our attackers also. So they are able to chain low vulnerabilities together and to form critical vulnerabilities from those um low low severity vulnerabilities. And what's actually is going on is that our attackers are leveraging those LLMs and they're attacking us especially where it's very difficult using um application context.

03:32 This is actually where our LLMs are struggled the most because they are creating many more vulnerabilities around those areas. So obviously we are leveraging AI but our attackers are also leveraging frontier models. Okay. And the numbers are simply incredible. Okay. We are familiar with an operating model that actually was able to identify more than 600 vulnerabilities uh to breach 600 firewalls in more than 555 countries.

04:01 Okay, this is simply unbelievable. And the time for an AI attack to be effective is constantly getting faster and faster. Okay, so the average time for an AI attack to be successful is 24 34 minutes. Okay, well the fastest time is 4 minutes. Okay, it means that since the moment I began talking, there is probably an AI attack out there that was able to successfully breach a production system and 43% of all the MCP servers actually has vulnerabilities.

04:39 Okay, so combining all of this data together, what we realize is that we are producing much more code than ever. Each one of us has a few sessions simultaneously of cloud code of codex. Those this code is insecure. It's getting into production and the people we are facing the attackers are actually leveraging the same tools and they are acting faster than ever in trying to actually leverage those vulnerabilities and our backlog is simply getting piled up and we are not able to get it.

05:11 And this is exactly what I want to talk about today. what we should be doing to to help improve our security posture. And this is ties heavily to how we can leverage offensive security not just for the attackers but for us as defenders. So why traditional security can't really keep up? So there is a mixed audience today both people from the development side also people from the security side.

05:42 So I want to cover first a few basic concepts okay to make sure everyone is familiar with. So what used to be really easy about operating was static code scanning okay SAS. So imagine that you are pushing new code to production. There is a static scanner that is scanning that and is looking for SQL injection cross-ite scripting. So it used to be relatively cheap and relatively easy to run this across every code change.

06:07 And this is great. Now what it doesn't catch, it doesn't catch things that are going into runtime. Okay, so if you have any authorization issues, if you have any configuration issues, for that you will need dynamic application security testing. Okay, you'll need something that test application dynamically. One example is bola uh broken object level authorization which basically it means that user A can access the data of user B.

06:36 Okay, those kind of vulnerabilities you can never find with static analysis. For that you need to prop the application in runtime. So what actually a dust scanner is doing it just takes a predefined set of tons of payloads and it just sending to the application to the different API endpoints and by that it managed to find vulnerabilities. So this was also something that could find new things but it wasn't good with the biggest challenge we have with the code that LM are producing and this is the business logic.

07:09 So for that we've been you what we've been using as an industry it's the pentesting. Okay. So what's the idea about pentesting? It's a human pentester that really understands the concept and the business logic of the application and is trying to break it and to find all the vulnerabilities. Now it is the golden standard in application security testing but it has some challenges right.

07:36 First of all, it's cost a lot of money, right? Because you actually needed to have a security researcher, a human, not a human in the loop, a human that is actually doing the testing and trying to break the application. And that's cost a lot of money. Now, the second problem is because it cost a lot of money and because the process was so slow, companies can only do it once a year, two times a year, maybe three times a year.

08:04 Okay? So combining the fact that nowadays our AI attackers are trying to continuously attack us like doing this pan testing two times a year is simply not good enough. Okay. So what you going to do with the rest of the 350 days you have a year that no one tested your application. Not now not only that the results you got the result you got this crazy and glorified report in PDF that you had to send to your developers.

08:32 there to take it back to understand what's real and what's not and then to start fixing that and then you had to retest. Okay, it sounds like a movie from the 90s, right? It doesn't make sense to do this anymore. But this was the Gund standard. This is what we were doing. Now the thing is this is the best method out there to actually test your application for its business logic.

08:56 Okay. So what changes today? Today we have a new technology and this is exactly what I want to talk how we can leverage this new area of AI tools and LLM and frontier models to actually do pentesting on a continuous manner in a scalable manner and this is why we're introducing EVO continuous offensive security. Okay, so this is obviously very tied up to what we're doing within SNIK, but what I'm going to tell you today is general.

09:29 Okay, you can take the principles of what I'm telling you and you can actually verify that with any vendor out there in the expo. Okay, so I want to I want you to think about the following things. Okay, so what you need to do today, you need to prepare your production system from the unknowns unknowns. Okay, and for that you need a new paradigm. Okay.

09:49 So first of all obviously you need the AI pentesting capabilities which I will double click on soon again. But you also need to start taking the concept of red teaming. Okay. So agent red teaming the idea is to take multi-steps attacks. Okay. So to try to simulate prompt ejection or data excfiltration and goal aenting and obviously to combine it with dynamic testing that is good at identifying authorization issues.

10:14 So we are taking all of those capabilities together. We are building one suite out of that for offensive security as part of our EVO offering. And the idea is to help you actually mitigate or help you uh handle the fact that the attackers are leveraging the same tools that we are using and they are doing that in a continuous manner 24/7 finding vulnerabilities faster than before and your backlog simply cannot can stand up with that.

10:44 So what we're doing differently within AI pentesting. Okay. So we I told you before the great promise of pentesting was always the fact that the human and the end really understand the business context. Okay. So now we are taking this to the LLM and if there is something that LLM are very good at is really trying to understand the business context. So behind the scene we are leveraging the LLMs and we are doing that in a continuous manner.

11:12 Right? So no longer you're going to do things like do pent testing twice a year or three times a year. We're going to do it on every code change. Okay? Every PR that you send, every delta, we're going to run this pentest for you and we're going to tell your developers what do they need to fix and how do they need to do it differently. Obviously we leverage our understanding of the application context, but we also help reduce the false positive.

11:38 For us as a industry, we have a severe problem of false positive. So we are actually proving you that the vulnerabilities that is that we are finding are actually exploitable and we just just flag them as bugs. We actually help you prove that they are exploitable and we take contexts for many different machines and for many different environments and we combine them together to help guide our LLMs.

12:01 Okay, so this is what we're doing differently and I think this is really what's makes this solution so unique and so well positioned to the fact that each one of us have a few coding gadgets running in production and pushing more code than ever. Now how do we do it? How does it work behind the scenes? Okay, so what I want you to understand is we are not taking the static scanners or the dust scanners and just combine him with LLM.

12:32 Okay, we build it from the ground up using LLM. So we have a series of a few agents working together. The first one is the LLM orchestrator. Okay, imagine it as the brain. This agent is working according to a dynamic plan that it follows and it knows how to operate all the different sub agents. The first agent or actually the second agent is the recon.

12:56 Okay, you can imagine it as your elite intelligence unit. The idea is to collect all the data it can about the different network endpoints, the different APIs, the different fingerprints, collect all of this data together, formalize an architecture out of that and hand it to the next agent. And the next agent is the vulnerability testing. It's actually again a series of a few specialized agents.

13:22 Some of them are targeting the well-known classes of vulnerabilities. Some of them are actually hunting for vulnerabilities in the wild and they are trying to break your application to understand the business context and to understand where can attacker can get from and then they hand it into the judge as I like to call it the exploit validation. So actually every vulnerability that the previous agent find are now being handed over to this judge and the goal of the judge is to tell us whether this is a real

13:56 vulnerability or or not to tell us whether this can be exploitable or not. Exactly because we want to reduce the cognitive load. Exactly because we want to reduce the amount of false positive we have in the system. Coming up next is the remediation agent because everything we do we want to make actionable. Okay, we don't want to report about a vulnerability that eventually the developers they don't know what they can do with that.

14:21 And coming after that is the reporting agent because everything needs to be reported. So what makes us different and those are the point I urge you to check with every vendor out there trying to provide you with the solution or whether you want to implement a solution inhouse for AI pentesting. This is a huge industry right now. So I think the most important thing you need to remember is and you already know that especially the developers in the room context is everything.

14:53 your LLM are as good as the context that you feed them with. So what we're trying to do, we are not just taking our AI pen testing solution as a standalone. We actually feed it with data from all the previous scanners. Okay, so the the statics code scanner, the dynamic code scanner, all of the different agent engines are being combined together and they are being fed into a our AI pentesting solution because that way we know how to tailor and we know how to guide the LLM where to look for and as you know also for

15:30 mythos or for frame when those models know where to look the results are significantly better. So that's one thing we're doing that is you know very unique and very differentiated. Um and also what I want you to remember is that one solution is not enough. Okay. The great thing about it is when you combine those things together and based on the data that we see both from the labs and from real data from customers combining the data together is really what's important and you can't really you can't really use just one

16:05 of them. Okay. There are certain classes of vulnerabilities. This dust is the best scanner for that. Okay? Because you have a set of really exhaustive pre um a set of really exhaustive uh payloads that you can send to your application. So it's very good for example with quite scripting. Okay. So dust for crossite scripting is the best solution. But if you want again like business logic then you should combine it with uh pentesting because it will be the best solution.

16:37 So I think what I really want you to take from this, if there is one thing I want you to take from here is to realize that the world is changing. You already know that. That's what made you come in. But you know we used to test our application once every few months because the change and the velocity of the change was so slow and we're we had people at the end that were reasoning about the application and about the change.

17:07 But that world is does not longer exist. Right? Right now we are pushing code to production 24/7 and we have LLMs that are trying to reason about the application context and they are trying to break it continuously. Some of them can do it in like as fast as four minutes. Okay, that's almost five time of the lecture of the session that I just did. So when you're coming to look for new solutions for this problem, you have to ask yourself four questions.

17:35 Okay. First of all, do we run this solution in a continuous way or is it just a point solution that we run every once in a while? Okay, that's a major difference. Second, do they really reason about the application context? Okay, because given the nature of those LLM, application context is really where the new vulnerabilities are coming from. And third, what more context can they digest?

18:03 Okay, how can they enrich the data? Because the richer the data is, the better and by the way cheaper the LLM operation is going to be. And the fourth part is do they really help me prove that there is a bug and to make sure that this is exploitable or they they just flag it as a bug. Okay, so our goal is not only to help you find the bugs faster, but to actually do it better and cheaper and to help you as a developer or as the security people here focus on what matters most.

18:40 Um, so thank you very much and uh I used to do it in every lecture. You know, my my my dream was to be a rockstar, but I had to be a cyber entrepreneur. So instead of that, I'm going to take pictures. So hands up everyone. Okay, great. Thank you so much. And I'm here for questions. If you want, go for the boot. It's very close here. sneak.io. Meet Evo, our new AI security offering. Thank you so much.