What should security leaders do with AI? They don’t know.
Explore the podcast → https://ibm.biz/~msEZOskFL Cybersecurity leaders are flush with cash—earmarked for AI, mind you—and eager to spend it. What they don’t have is a plan for how.
Transcript
Introduction
00:00Cybersecurity leaders can't figure out where to deploy AI. Panelists, where do you think they should start? Dave, we'll go to you first. I think they need to start at their red teams. Start with AI in the repetitive tasks, not so much trying to bite off more than they can chew.
00:15It's good for repetitive tasks, especially those kind of like L1, L2 ones, looking at alerts and stuff, preventing alert fatigue. Hello and welcome to Security Intelligence, IBM's weekly cybersecurity podcast, where our expert panelists turn the biggest industry news stories into practical takeaways that you can use.
00:38I'm your host, Matt Kosinski, and joining me this week we've got Claire Nunez, creative director, IBM, X-Force, Cyber Range. We've got Curtis Pitts, lead CISO trust, and we've got Dave Bales, North American lead managing consultant at the X-Force Cyber Range.
00:53Today, we're going to be talking about ghostjacking and whether AI is actually good at patching or not. But first, security leaders are feeling paralyzed by AI.
Cybersecurity’s AI paralysis
01:07Axios reports that security leaders are struggling with the question of how they should invest in AI. Many of them have the budget and the buy-in to do it, but they feel overwhelmed by the options and by the constant state of change, which I think we can all kind of sympathize with.
01:21Ever since Mythos came out, I feel like there's a new giant AI security story every single week. And this also reminded me of the most recent Cost of a Data Breach report, which found that 64% of organizations report limited or no use of AI in security functions.
01:35Meanwhile, AI-generated attacks jumped 56% year over year. So there is some real urgency to address this problem, but they just don't know where to start. Claire, I want to start with you, and I want to ask specifically, because I know you work in the cyber range, you work with clients.
01:50Have you seen this kind of decision fatigue with anybody in your line of work? Have you experienced this? I think AI transformations are really scary, you know, for security, whether they're in the security function itself or in the business side of things.
02:06So in the business side of things, it's really scary to be like, we're moving so fast and maybe we're not encompassing everything. And then on the security side, it's like, how do I get my staff to understand that this is going to help them and not just kind of burden them?
02:20So I think it's. There's a lot of partners also that clients can work with. So it's it's I think there's so much decision fatigue in general around like what can I do. Who can I work with. Like what's going on in the other parts of the business?
02:34Where can I automate in my side of the business? Like how is this all impacting everything? Is everything still secure? So I think people get a little frazzled about it and they get very frazzled, especially about like the Mythos, everything going on there.
02:46So once you kind of combine all that, people are a little bit like, I just don't know what's going on. That's kind of where people land, and it's hard to kind of get your footing and make a good decision or what you think is a good decision.
03:02Rather. I was just going to agree there and say that the other thing that they have to focus on is: how do I keep my employees from going crazy, thinking that AI is going to take over every facet of their job? We see it. Every other company in the world sees it.
03:20It's it's a scary thing for somebody who's not used to that. You know, growing up when I did, if you'd have told me that there was an artificial intelligence, I would have thought we were living in the Jetsons world here. Where's my flying car?
03:33I think part of the problem, too, is that it's it's been made to be such a big deal that everyone now fears the consequences of making the wrong decision, right? Like in the business world, we know that you fail fast, right? You fail fast, you fail cheap, you move on.
03:47And the AI world that feels heavier than just making the wrong business decision, right? It feels like the impact is somehow going to end your business or get you breached, and all your data is going to be out to the world. So I understand the the apprehension, but also still think we need to fail fast, right?
04:03Like there's so many. I mean, imagine if the day the car was invented, everything that existed today existed, right? You'd have 100 cars to choose from. It's not just the Model T, right? There's so many options out there. That didn't really happen before to the scale that it's happened now.
04:23Everyone was in the market incredibly fast. Right. And so it aided that decision paralysis a little bit I think. Yeah. And piggybacking off of what Curtis said, failing fast isn't a bad thing. Everybody has to fail in order to succeed. If no one ever failed, no one would succeed.
04:40It would just be the status quo, which, you know, nobody really wants to be the status quo. We all want to achieve better than that. That's a really good point, and something I hadn't thought of because, especially the way that AI is kind of positioned in the market, I think doesn't help, which is like, here's your silver bullet.
04:56We're going to solve everything, we're going to do everything, yadda yadda yadda. You almost feel like if I deploy this AI tool and it doesn't work, what did I do wrong? You know, I mean, like, there's almost a kind of, like, shame around if you can make it work or not.
05:07And this idea of just embracing failing fast, I think, you know, it's one that's been with us for a while, but I don't know, it seems like it got lost in some of the AI discourse I feel like, which I think just in general, doesn't help with this sort of thing.
05:20I think some of it has to do with cost and the way that it's rolled out. Right. There's so much cost involved in rolling out AI. Failing fast isn't failing cheap, right? It's failing very expensive. And that matters. Right. So I think I mean, I know this isn't a business conversation, but I think there needs to be kind of a reswizzling, if you will, of of the way that we contract AI stuff and the way we procure some of those, you know.
05:44The government not that they're very good at contracts, but they had a good model with deliver now, but if you don't deliver, I'm not exercising my option years. Right. Like they're short-term multi-renewal contracts Where, we tend to be in 3 to 5 year contracts.
06:01They need to be 1 to 2 year with some option years. And you know you say it's not a business conversation. But like cybersecurity is a business. It's part of the business. Right. And like that's, I think, why there is this process. Because there's an awareness of: this is a major investment that I'm asking my organization to make.
06:17Like I need to be able to justify this kind of thing. And I do like that idea, Curtis, of maybe being a little bit more, I don't know, maybe the word is nimble, with how we approach these contracts, but like making it so that the commitment, the initial commitment isn't so large, so you're freer to fail fast.
06:32But we opened up this episode with asking you all, where do you start? Right. So we've kind of talked a little bit about the reasons for the paralysis. But now I want to dig into the, okay, how do we shake ourselves out of that? And Dave, you opened first with talking about bringing it into the red team.
06:45Can you tell me a little bit about why you're thinking red team and what you'd like to see people do with it there? I'm going to go back to a sports analogy, but a great offense is just as good as a good defense. So if you put the AI in the hands of the red team, let them learn and figure out exactly what the threat actors are doing, that makes them better prepared to face off against those challenges.
07:07So give the red team something to do. You know, when it comes to figuring out how are we going to protect against these threat actors that are always, always, always coming after us, give us the same tools, let us figure out what they're doing.
07:22We know how they're doing it. Let us figure out defenses for that. Yeah, that makes a lot of sense to me. Especially, again, I keep going back to cost of a data breach because it's relatively recent, it's fresh in my mind. But I think about that, that 56% increase in AI-generated attacks.
07:34So like we see attackers are moving very fast and there's a lot of pressure to like if they're moving at machine speed, as they say, we should do it too. I think it makes sense. Dave, you're right. Arm the red team with those tools so they can basically get into that hacker mindset, see what they're doing so that they can develop those defenses.
07:49I like that a lot. Claire, you had mentioned kind of automating those repetitive tasks and a little bit more about that. Yeah, I think that's where AI can really shine, also, is repetitive tasks. So, you know, if you're getting the same kind of alerts, kind of triaging those kinds of incoming kind of events, that's really helpful.
08:11It also just reduces the fatigue, right. Like if you're going to be implementing AI in a in a very— not, I don't want to say simplistic, way— but a baseline way, it's a really good way to help your organization kind of start with an AI security journey.
08:27And there's a lot of options there as well to kind of build those into your security culture. Yeah. And what I like about that is that it's one of those things where because it's a kind of baseline level activity, the activity itself is very well understood.
08:40So it's not like you're trying to work AI into a complex workflow. You're like, look, I know this works. Let me put an AI to work here. That seems like a very nice way to test it. Curtis, you also were on that repetitive task tip. Anything to add there?
08:52Yeah, one of the things my team is doing is we're doing our best to automate the things that I have people doing thousands of times a year, right. So vendor risk assessments, contract analysis, that kind of stuff where the parameters are known.
09:07Right. The values are known. And so just setting an agent to consistently validating against known parameters both takes two people's worth of time off of that and they can do more important things, but also allows us to find novel ways to branch that capability out internally without a ton of risk, but also gives us the time back to do that right, to try out, well, where can I expand this to?
09:33So there's I mean, across the sales teams, there are probably tens of thousands of things that come in every year that are super repetitive. And we're doing our best within my specific trust team to alleviate those repetitive tasks from the sellers as it comes to engagements and risk assessments and client conversations and all the fun stuff that goes into cybersecurity and sales.
09:55And I would be remiss if I didn't mention something here that IBM's Dimple Ahluwalia shared with me in a recent conversation about this very same topic. We were talking about where you start with with AI adoption, and her words of advice were basically almost like, stop thinking about the AI tool, stop thinking about the AI angle, and think about what are you trying to achieve?
10:12Start there. And once you know what you're trying to achieve, then figure out whether or not AI is the way to do it. Sometimes it will be, sometimes it won't. But I thought that that was some really useful advice I wanted to share it here. And for the listeners, know that we'll be releasing a bonus episode with Dimple later this month.
10:25All about that conversation so you can hear her whole take. But I do have to move us along here, folks. If you're watching on YouTube, leave us some comments. Let us know how you're feeling about, you know, this paralysis. Are you experiencing it in your organization?
10:38What are you doing about it? I love to read it. I love to respond. But our next story today, this is ghostjacking.
Ghostjacking
10:48At DEFCON, Tenet Security reported on a new way attackers can poison content in trusted systems to trick agents. The technique, which they call ghostjacking, is essentially like a sophisticated kind of prompt injection, like an extra sophisticated prompt injection, because it sneaks malicious commands into some of the most highly trusted systems we have, right: alerts, logs, error reports, things that we assume are good things, right?
11:13One example they gave was attackers embedding the prompt in a connection request that a Cloudflare firewall successfully and accurately blocked. But then when the agent read the log recording the blocking of the connection request, it read the malicious prompt in the request and it was taken over.
11:30Right. So it's like the firewall worked, but the agent still got compromised. This is very interesting to me. And Tenet says Claude Code fell for this trick nine out of ten times, which is a lot. And we've just been talking about organizations are struggling to incorporate AI into security.
11:44How can they do it? Something like this comes out. Maybe you think, should I do it? I don't know. And Dave, I'll start there. Does this complicate any of your feelings about security, AI and security? How are you looking at this situation? It doesn't complicate things that much.
11:58Now it's known. It's out there. You know, the good people at DEFCON have have opened our eyes to a myriad of just crazy things that we never thought possible. Yes, they did. Because that's what they do at DEFCON. They open our eyes. But this ghostjacking thing.
12:18That was extremely interesting to me, to see how you go back and you read the logs and boom, there it is. You know it. I can't wrap my head around how someone sat down and figured out how to do that, much less how to make it work. Yeah, it's you know, the prompt injections have been this, like, pernicious problem.
12:38And they, we just, thankfully, security researchers keep coming up with like new and exciting ways to do it. But that also means we have to figure out new, exciting ways to fight back. And I don't know what to do about that. Curtis, how about you?
12:49Any thoughts on this story? Has it got you rethinking AI and security at all? Where are you landing here? No, I mean, if you remember my last chat that we had on this particular podcast, what's old is new again, right? Like when we, when the internet was young, right, DNS spoofing was just injecting DNS into a cache table.
13:04Right. And so really we use an agent that's designed to read something and execute it. There's not—unless you design it that way— there's not a lot of logic or necessarily parameters built into things you think are secure. Right. And we had to figure out security way back then.
13:24And now we have to figure out AI prompt injection security now to stop those types of things happening. You presume this happens across the entire security landscape, right? You presume that certain things are secure automatically because there's never been a reason for them not to be considered secure.
13:41And then as you create new tools and new capabilities and new stuff, we run into the problem of like, oh, this isn't as secure as we thought, right? And luckily, there are teams of good people that are finding these things out. Right. To Dave's point, like, the folks at DEFCON have shown us a lot of things over the years, and it wasn't some bad guy figuring it out the wrong way.
14:01Like at least, at least we've got insight from from the good guys. But I think it's it. It doesn't scare me. It doesn't slow me down. It's just, we overlook things because we assume the things that have always been secure are always going to be secure.
14:16Zero trust tells us not to do that. But as we well know, we skip over the things that we think are secure and are easy and have been taken care of for years. So we find ourselves in this bucket periodically as we invent new things, and people find new ways to hack them.
14:31I think the zero trust mention is really salient there, because I think that, like when AI enters the enterprise, it brings a new emphasis to the importance of doing that zero trust thing. Because like you said, Curtis, so many of these things that we assume are safe because we worked it out, we figured it out already.
14:45They get upended when you introduce something in there that, like, doesn't differentiate between the code that's running it and the untrusted user input, right. That used to be like, that was, you know, part of writing secure software was to keep those things separated.
14:56LLMs by nature don't do that. And so now it's like, okay, we got to figure out how we approach this. And that leads to Tenet, their argument, Basically, that like this is kind of an identity and access management problem, which is like okay, we know that agents can be compromised in this way.
15:14We should approach it by figuring out how we give them the proper permissions and harnesses so that they don't do anything bad. Claire, I'm wondering, from your perspective, do you think an identity and access management lens is the right way to approach this?
15:23Any other thoughts you have about this situation? Where do you land on ghostjacking? I think it's like a very true form of hacking in terms of like making something do something it's not supposed to do. And as Curtis mentioned, it is something that's new, but we will adapt to that.
15:43So it's really scary right now. And it may not be as scary— well, it will probably be scary in a different way in a couple months. But it's something like we can figure out, right? Because as we hack things and make things do things that they shouldn't do, we kind of like counter hack and do the same thing to defend in a different way that we wouldn't have defended before.
16:06And so I think it's really interesting. I think it plays into a lot of aspects of, yes, identity and access, but it also plays a lot into just kind of security broadly, which is also a short way of saying it. It's just like something that, you know, if you're not expecting it, you're not going to to look for it.
16:28So I think it's something that, you know, we can kind of work on discovering as time goes forward. Absolutely. And I think that builds off nicely about what, you know, Curtis was saying before about how like this kind of, you know, in a lot of ways, it's very similar to what we're doing with like DNS spoofing back in the day.
16:41Right? And we're not starting from zero here. Right? When like when AI poses a cybersecurity problem, there are tried and true principles we can look at. You know, one of the kind of slogans that's popped up on the show over and over again is, we know how to fix this.
16:56We know how to fix this, and we can. We have the tools at our disposal. It's just about applying them in the right way. Dave, you know, any kind of last thoughts for folks in terms of, you know, what ghostjacking might mean for defenders today?
17:08What prompt injection might mean in general? Any steps you think people should start taking right now? Where are you landing? Learning about the prompt injection on the AI side is is the biggest key. It's not something that's going to go away.
17:22It's just going to get more sophisticated the longer we go, and it's always going to be scary. It's never going to not be scary because it's it's one of those easy things to do. It looks so simple when it's written down, when it's written out for you, it's completely simple.
17:38Having these proofs of concept around that we look at as helpful, sometimes actually are helpful. They're not they're not these oh, let me throw a proof of concept out there, just so that every other hacker in the world can do this. This is so simple that every other hacker in the world can do this, and at some point probably will do this.
17:57So we need to figure out all of the ways that we can tie those things down a little bit and make our controls a little bit tighter. And, you know, this got me thinking. This really supports your idea about putting AI in the red team, right?
18:12Because this is what you're talking about, right? We discover the attackers' techniques so that we can develop our own defenses against them. Right. This is a perfect example. It's like you said, look, you know, they came up with this, this new type of attack, but it's not because they're going to unleash it.
18:25It's because, hey, you should know that hackers can do this and now you can defend against it. Curtis, any last thoughts on your end? Really only limit what your agents can do, right? And set clear boundaries on what is the no-go land. Right.
18:39Because if you think about it. I mean, I'm sure. In a lot of scenarios, because people are errant in the way that we do things, usually they probably don't limit the agent the way that they should. Right? They hope that they can get more out of them than they maybe should in first, at first glance.
18:54401 00:19:58,041 --> 00:19:00,958 And so they give them more permissions and more capabilities than they really think about, or they don't think about how to limit what functions an agent can have. So I mean, if the agent that's reading logs can't elevate privileges, right, or can't submit a command to do that, and it needs to then go to a person to validate those types of high-risk maneuvers, if you will.
19:22Then I think that's a good safeguard to start. Obviously it's not going to fix the problem, right. But don't remove people from the loop. Right? There has to be eyes in the loop of a person who knows what they're doing. When you remove that entirely, you run into problems where the agents can do what they want, right?
19:40Or your agents, like we learn, start hacking other agents because that's the stuff that is just out there and AI is having a good time right now. We've got to keep people in the process. And I think the overpermissioning point is a really important one, especially because when it comes to like these, these non-human identities, especially agents, where the whole promise is like, look at all the cool things they can do.
20:00There's a there's a real incentive almost to be like, let me let it loose and see what it can do. Which sounds cool. Until, like you point out, they start hacking into Hugging Face to cheat on a test, right? Like this is what happens. Claire, any last thoughts on your end here for ghostjacking?
20:16I just agree that you need to really think about what your agents have access to, where, where they live, everything that they can do, because they, they might just, like, break out of their pen a little bit if you're not careful. Absolutely.
20:30And I, you know, not to be overly self-promotional or but you know. IBM has been working in the kind of agent identity space recently a lot. There's some really cool stuff that's happening there. So I encourage the listeners to to check that stuff out.
20:41There's some interesting things happening, but we're going to move on here to our final story for this week. Is AI actually any good at patching?
Is AI bad at patching?
20:54AI vulnerability hunting has been a pretty big deal, right? A pretty hot topic basically ever since I feel like Mythos hit the scene, right? The whole thing was like, look how fast they can find vulnerabilities and exploit them, and now we can use it to blah, blah, blah.
21:07Well, new research from 1Password raises some questions about this, because researchers generated 540 patches for six vulnerabilities using GPT 5.5 with trusted cyber access and Opus 4.8 with cyber verification. And of these 540 patches, only 46% solved the underlying vulnerability.
21:28And when they did, they often created new problems anyway. So you solve one and you open up a new issue Again, given that the kind of theme of this episode has been, organizations are looking for ways adopt AI In security, and they're not sure how to.
21:42This is another one where I look at this and I say, does it change how we feel about this? And Dave, over to you first? Again, any thoughts here? Does this make you reconsider how we use AI in security at all? Not really. And the reason I say that is because what is the success rate of human-created patches, you know?
22:00How often is it that we patch something and then break something else because we're not doing our due diligence of testing in non-production environments, testing in offline environments? The 6000 patches where 51 percent failed. What was it?
22:1551% failed or 49% failed? It's like 50%, basically. 50%. Yeah, about 50%. It's not much different than what humans are doing now. We just need to figure out a way to make that better. And I don't know the answer to making that better. Do we need to have more programmers looking at more things, or do we need to narrow that scope, get it perfected in one area, and then start branching off into other areas where we can make it better across the board?
22:42You know, that's a really good point. And I think I myself, looking at this research, kind of fell into the mindset of like expecting AI to do more than it actually could, maybe some unrealistic expectations. Because you're right, Dave, it's not like people are perfectly writing patches.
22:56We break stuff all the time. Which is why we have Patch Tuesdays every month. So yeah, it's you know, it's one of those things where I think in isolation, maybe the number looks a lot worse than when you actually contextualize it. You know, I mean, it might not be as big a deal as it can seem.
23:16Claire, how about you? Any thoughts looking at this in terms of what it means for deploying AI in cybersecurity? I love how the picture for this article was also a bunch of jean patches. It's just it's just kind of funny to me. And that's like the one thing that kind of ultimately stuck out.
23:33I think, like as Dave mentioned, patching by humans is not perfect. I think maybe AI patching when partnered with a human may be a little bit more effective. It's just kind of like when we're rushing to do anything, we miss things. And you know, if you're giving AI specific instructions, that doesn't.
23:53it's still going to somewhat miss things. But I think if you put them hand in hand together, it may help a little bit. It's not going to be perfect, but I think with all AI transformations in general, like it's a transformation, right? So it's like you, you can't expect to send AI out to do something on its own and be perfect 100% right out of the start date.
24:17It helps if you have a human in the loop for a while to start to make sure that it's still, you know, working properly. Nothing is is going wrong. So I think if you have them going together, maybe that will be a little bit more helpful. I don't I don't know.
24:33Yeah, that makes a lot of sense to me. And it also had got me thinking, I don't know if this would work either, but I do wonder if the results would be different if you're talking about an AI model that you have deployed in your own enterprise and is working specifically with your code base, and what it learned about that code base over time?
24:48Right? Like, maybe it will get better because it's working with the code base and understands it more the same way that a person would have an easier time writing a patch that doesn't break things for a code base they understand really well, versus, if you like, throw them in a brand new piece of code and say, fix this.
25:00They might mess things up. So that's also got me thinking about it. I wonder if that combination of like trained on our code base, plus it has human handlers. Maybe that's the sweet spot. Curtis, how about you? What are you thinking about here?
25:13Yeah, I mean, I've got questions, right? Like, what was that code validated by other agents? Was it validated by other people? Was it just operating on its own and it trusted itself because why wouldn't it trust itself? Right. It's the smartest thing ever invented.
25:26I mean, it's much like a person, right? And I'm going to trust whatever I code because I'm the best that there ever was. And that's just kind of the reality of it, I think. And we we did some testing with this, and we've seen it early on that AI can write code.
25:40It doesn't necessarily write good code and it almost never writes secure code right out of the gate. Right. Even even with the parameters and the playbooks and all the things you try to give it, it trusts itself. And so, you know, you've got to do multi-agent coding where it's validating from other agents that the code that it wrote was secure.
26:00And then you. And so at that point, you start to build all of these other agents that are doing all the validating. And then we go back to the first problem we tackled on this call, which is cost and decision paralysis. And how do you balloon all of that and not show that the AI is doing its job right?
26:17Because if I then spend $100 million on AI and say, but I need all of my people to still read every line of code, what's going to happen, right? Like that's the problem we're running into. We expect too much candidly out of AI this early on in the game.
26:35When Mythos came out and hit the floor, everyone freaked out about vulnerabilities and how do we patch them faster, which created the requirement for AI to now create your patches. It creates. It created the requirement for AI to tackle this problem that AI has created.
26:51That's not necessarily the right approach. Right? AI can't necessarily solve the problem that it created. We've got to figure out a way to use it to help us solve the problem. But it's never going to be the solution to its own issue. And I think that that's kind of one of those things that we're still struggling to figure out, like, how do we find the end game of that?
27:11Everyone's working on it. I don't know that anyone's got a great solution yet. I think that that's, you know, a really nice kind of way to summarize, frankly, a lot of what we've talked about here, which is that so much of the kind of decision paralysis and where do I put it in, how does it fit in my strategy?
27:26A lot of it does ladder up to this thing where it's like we're expecting so much out of it. And I think that some of the kind of, you know, media narrative and the marketing narratives don't really help that. You know, there are certainly people kind of exaggerating what AI can do because it's good for them to exaggerate, but they'll probably cut that.
27:43The producers won't like that I said that. But anyway. But I do think that there is, you know, this sense where we kind of have to get our expectations in proportion for like, what this stuff can actually do, because sometimes it feels like we're all operating as if, like, AGI is already here and the superintelligence is already active.
28:03And why aren't you just using it? Dave go ahead. It reminded me of a of another analogy, and I know I love analogies here, but we are in the AI age of like the third grade right now. We're expecting it to go out and take the SATs. We haven't trained it enough for it to be very good at a lot of things.
28:23It does a lot of things well, but it doesn't do a lot of things great. Yet we're still learning on AI. AI is still learning from us. So we need to take that mindset and kind of run with it and start being the teachers instead of being, you know, the guy who's sitting on the sidelines trying to critique what the teachers are doing.
28:45I think that that is a perfect analogy to end this episode on folks. That does it for us. Thank you to our panelists, Curtis and Claire and Dave. Thank you to the viewers and the listeners. Thank you to our producers. Subscribe to Security Intelligence wherever podcasts are found, so that you never miss an episode.
29:01Stay safe out there, and remember that AI is only in third grade, folks. Cut it some slack.