← All transcripts

AI Is Exposing Your Data: An AI Security Problem You Can't See Transcript, AI Summary & Key Points

IBM Technology · 4 days ago · Education · 11:29 · EN

Watch on YouTube

AI Summary

AI systems can expose sensitive data through training data, prompts, uploaded files, context and policies, RAG pipelines, models, agents, tools, databases, endpoints and downstream agents. A recent study found that 31% of organizations had a data privacy violation due to an AI-related incident. Shadow AI and public cloud chatbots can bypass required security controls, and sensitive information entered into public chatbots may be used to train models and become available to others. Managing exposure requires continuous AI-aware data classification and discovery, end-to-end data lineage, visibility into transformations and propagation, unified policies across workloads and employees, intelligent investigation, proactive risk identification, and compliance monitoring.

Key Points

  • 31% of organizations had a data privacy violation due to an AI-related incident, according to one recent study.
  • Shadow AI projects may lack the security controls required by organizational policy.
  • Public cloud chatbots may receive sensitive spreadsheets or other information, which can then be used to train models and become available to others.
  • AI data flows through training data, prompts, RAG sources, context and policies, models, agents, tools, databases and other agents.
  • Sensitive data can enter through training data, user prompts, uploaded spreadsheets and documents, policies, context and agent tools.
  • AI security requires tracking what sensitive data is used, where it came from, how it moves, whether exposure was authorized and governed, and where it ends up.
  • Workload visibility includes AI applications, RAG pipelines, vector databases, data transformations, data sources and destinations.
  • Workforce visibility includes file uploads and downloads, copy-paste operations, derived child files, employee consumption and sharing.

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of AI Is Exposing Your Data: An AI Security Problem You Can't See — IBM Technology (11:29). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 It's AI time. Do you know where your data is? Well, even if you think you do, you probably don't. If you're like most organizations, your sensitive data has already been exposed and AI is just making the problem worse. And it's happening right under your nose and you can't see it. Here's what I mean. AI adoption is moving faster than traditional security approaches can actually protect.

00:23 One recent study found that 31% of organizations had a data privacy violation. Due to an AI related incident. There are things like shadow AI projects where it's basically unauthorized AI that people are deploying inside the environment. Well, those things often lack the security controls that our policy would specify they should have. And then there's use of public cloud chatbots, these things where people go in and ask their questions, but they may also include a spreadsheet of information that's sensitive.

00:57 Once that goes to a public chatbot, it basically becomes public information because they can use that sensitive information to train their models, and then it's available to everyone. So also, we've got some questions that we have to answer about how AI is using this. So first of all, what data did the AI use? Where did it get that data? And ultimately, how are we going to manage?

01:26 The AI data exposure in a proactive way. Traditional DLP data loss prevention and AI tools can't answer these questions. We're gonna need a different approach. It's not enough just to know what AI tools are in use, although that certainly is a good place to start. You also need to know how sensitive data is flowing through AI systems and where it might be exposed.

01:48 So let's take a look at a high level architecture to see what I mean. So it starts off with training data. That goes into train or two in a particular model. Then we have a user who enters a prompt that goes in to the AI, which is gonna have it then do other things. Along with that prompt, we may augment that with other data sources, so a retrieval augmented generation.

02:12 Then we probably have a context or a policy that's overriding all of this, and that might have some sensitive aspects to it where we... Put other information telling the AI how to operate and how to come up with its answers, then if it's an agent, it may also reach out and use different kinds of tools. Some of those tools may write code, they may access databases, they may, in fact, spawn off yet other agents, which then could spawn off other agents.

02:40 Okay, so that's the general flow of how the information goes and all the different components and how they're interacting. How about from a data perspective? What could be sensitive here? Well, we might have sensitive information here in the training data that then is going into the model. We might have some sensitive information that this user has that they are entering in with their prompt.

03:04 They may have used a spreadsheet or a document that has sensitive information and that is also going along with the prompt and going into the AI. We may have sensitive information. In the policy and context, maybe ways that we do things that we consider to be competitive advantages we wouldn't want our competition to know about. So this also is feeding into the model.

03:28 Then the model is reaching out and using tools. How about these tools? What are they doing with the information that they receive? Are they guarding it or not? Do we have control and visibility over those particular tools? If they're writing it to a database, do I know where that data is gonna end up? And maybe it's going other places downstream. And then these agents, where I'm taking potentially sensitive information and feeding it into the hands of these agents.

03:57 And then they spawn off yet other agents. So you can see when you look at this that sensitive information is all over this thing, and it's flowing all through the architecture. So we've got a lot of things we have to consider in this case. So first of all, we need to know what sensitive data is being used by AI. And then we need to know, what is the exposure?

04:22 What data is being exposed? Was it authorized and was it governed? Then we wanna know, how do we investigate? What happened as the data moves through the AI? And then ultimately, how to we enable AI adoption without creating data security gaps? So in order to do that, we're gonna need to be able to do some different functions where we're looking to see where the data is going.

04:49 I'm gonna look at this from a couple of different perspectives. There's a workload stream, which is basically inside the AI systems themselves and the workforce stream where this is the employee's use of the AI. And what do I need to do? I need the monitor, track, and then be able to show how that data has moved through the system and how it's been used and how its been exposed.

05:10 So let's take the workload example first. So this is inside the AI systems. I've got to look inside AI apps and see how they're using the data. I need to look at rag pipelines, see what information is coming into the system that way. I'm going to look at vector databases, which are the heart of a lot of these generative AI models. Then I'm gonna need to be able to track all of this.

05:33 I've gotta data transformations that are gonna occur. The information looked like this, but now it's in some other form. I still need to know about that. If I'm looking for it in only this form, I might miss it and not realize that the data has leaked out. So I need to be able to realize when the data had been transformed. And then I need to be to show all of this.

05:53 I need be able show the sources of where the data came from. I need to show these transformations that I just referred to. And ultimately, where does all this information go? So that's the inside the AI view. How about inside the workforce view? The workforce or the employees. In this case I want to look at the file uploads and downloads. I want see what stuff they put into the system, what stuff they took out of the system.

06:21 I wanna see when they did copy paste operations. Maybe they pulled some sensitive information and then pasted it somewhere else. So you can see the data now is living in a different home. I've gotta look at child information. So I have a main file and then there are child files that could be derived from that. And I need to be able to track all of those just as I did the original file that had the sensitive information.

06:47 I need be able show how employees are consuming this information and how they're sharing it with other employees as well. So all of things need to viewed and I need to see this in a unified end-to-end visibility protection platform. What kind of things do I need from that? Well, I need able to see the lineage. I need to be able to see where did the data come from and how did it move through the system.

07:15 I need be able see things like the source, then it went through this AI, and then it ended up on this endpoint system. I need policies that are going to cross all of these. So policies that shared across both of these and monitored and enforced. I also need a single pane of glass where I can monitor all of this. I can't afford to have a whole bunch of different monitor systems that are not integrated, because then I won't know.

07:43 And then ultimately, I need to be able to identify risk. I need a risk identification capability where I can see this proactively, not just after all of the data has escaped. But how can you get the kind of visibility I've just talked about here? Well, let's take a look at three different lenses that we could look at this through, three different tools that will give us different aspects.

08:06 So we'll start off. With agentic platform discovery. And tools that do this would be able to answer questions like who prompted which particular model and which agent actually ran, which MCP tool was called. So that's what those are particularly good at. Then we could also look at this from an endpoint DLP data loss prevention discovery tool perspective.

08:29 Those tools exist and they could tell you things like which agents and extensions and MCP servers exist on each of these devices. It could tell you files, memory, and other digital residue that's on the endpoint device. Then the third perspective here is the cloud or could be on-prem discovery. And this is gonna tell us which data sources exist on these particular platforms.

08:54 We could also figure out the classification and sensitivity of the data that's on there and maybe even the ownership. Okay, but all of that's good, but did you see the problem? Each one of these only sees part of the picture. So what we really need is a holistic view that takes all of these into account and integrates them all into a single view. So what's really needed here to manage data exposure from AI?

09:22 Let's take a look at the requirements. First up, we're gonna talk about an AI aware automated data classification and discovery system. It's gonna have to be continuous because the system is changing constantly. It needs to recognize sources across all of our platforms and it needs to be aware of the sensitive data that we have, things like personally identifiable information, personal health information, financial data, intellectual property, all of that kind of stuff.

09:52 Next thing we need is something that shows the lineage-driven risk visibility. So, data gets transformed by RAG. AI agents and systems and things of that sort may also do transformations. And then we need to see how the data propagates through the system, and it will be changed as it propagates in some cases. And then finally, we need an intelligent investigation capability.

10:18 It needs to be context aware based on the user sensitivity, the destination, and a number of other different factors like that. And we've got to be able to take the investigation time which is typically weeks down to minutes. We've got to operate with speed. And then ultimately compliance reports. There's a lot of different things we have to consider.

10:42 The GDPR, Generalized Data Protection Regulation, the EU AI Act, SOC 2, ISO 27001, HIPAA, the list goes on. We need to be able to see all of those things and how our AI is in compliance. Data is the lifeblood of AI. It has to keep moving. Or the patient dies. But if we can't monitor where it's going, we could be hemorrhaging and not even know it. Like a human body, the system is complicated, but there are tools that can help you manage your data exposure from AI.

11:13 That's the good news. Want to learn more? See the link in the description below.