
Loading summary
A
The most significant autonomous AI cyber attack in history just took place with OpenAI's models breaking out of a training environment, connecting to the Internet and then hacking hugging face to ace an evaluation. What does it mean for the future of AI and for cybersecurity? Let's talk about it with ex Meta Chief Security Officer and current Corridor Chief Product Officer Alex Stamos right after this. In the face of ongoing disruption and opportunity, TMT leaders need to deliver tangible results, not just ideas. When pace and performance matter most, PwC combines market insights and deep sector experience with AI, cloud and emerging tech to accelerate your transformation and drive measurable ROI. From strategy to execution, PwC can help you anticipate what's next, outpace disruption and compete. For more information, visit pwc.com Insurance isn't
B
one size fits all, and shopping for it shouldn't feel like squeezing into something that just doesn't fit. That's why drivers have enjoyed Progressive's name your price tool for years. With the name your price tool, you tell them what you want to pay and they show you options that fit your budget enough. Hunting for discounts, trying to calculate rates, and tinkering with coverages. Maybe you're picking out your very first policy, or maybe you're just looking for something that works better for you and your family. Either way, they make it simple to see your options. No guesswork, no surprises. Ready to see how easy and fun shopping for car insurance can be? Visit progressive.com and give the name your price tool a try. Take the stress out of shopping and find coverage that fits your life on your terms. Progressive Casualty Insurance Company and Affiliates Price and Coverage match Limited by State law
A
welcome to Big Technology Podcast, a show for cool headed and nuanced conversation of the tech world and beyond. We have an emergency podcast episode for you today because just yesterday the world found out that a series of OpenAI models work together to break out of a sandbox, hack into hugging face, steal basically the answers to a test and go and ace their evaluation. And obviously this is not the desired behavior that OpenAI wanted and looks like it might have opened up a new can of worms here for AI and cyber security. So we are joined by the perfect guest to help us figure out what happened here. Alex Stamos is with us. He's the Chief Product Officer at Corridor and the former Chief Security Officer at Meta. Alex, great to see you again. Welcome back to the show.
C
Yeah, thanks for having me Alex.
A
You know you spoke at our summit and I was like we're definitely going to have you back pretty soon. And it is amazing how the AI story has just turned into a cybersecurity story very quickly.
C
It has, you know, there's all kinds of risks from AI and, you know, there are all kinds of bad things that happen to consumers. But when you talk about the models themselves, it seems that cyber is the thing that's hitting right now for sure, from. From a cyber level risk.
A
Yeah. And so this is what we're talking about now. And the reason why we have to do an emergency episode on this is because this is certainly a novel type of hack. Right. So this is fairly unprecedented. Just to put it in context, it's the first time. This is from Transformer. The breach appears to be the first known example of a misaligned AI escaping containment and autonomously carrying out a cyber attack on a third party, a scenario AI safety experts have repeatedly warned of. So it's not like the anthropic example where Mythos sort of escaped containment and emailed somebody while they were eating a sandwich in the park. This is actually going out and hacking a third party. Let me just quickly read the beginning of the Wall Street Journal story about this just to set the stage. So the headline is OpenAI models escaped and hacked the company in cybersecurity tests gone wrong. On Tuesday, OpenAI said two artificial intelligence systems it was testing broke out of their test environment, hacked their way onto the Internet, and broke into another company. OpenAI said the culprits were a pair of its models. One was its latest product, called GPT 5.6 SOL, and the other was an even more capable pre release model the company didn't identify. The software had been configured for eval purposes to be less likely to refuse hacking commands. OpenAI said the OpenAI had caged the models in a sandbox, a system that didn't have access to the Internet. But during the test, the software used its hacking skills to break out and found a way to get online and then hacked into Hugging Faces network. And of course, Hugging Face is a library of open source AI, mostly AI models or AI programs. Alex, how significant is this? Like, you know, obviously there's a, there's a tendency to be alarmist about some of these things and. But I wanted to bring you on because you are the cybersecurity expert and so you can tell us, like, is this a 10 on like the 10? Holy, holy crap. Like we're in some deep trouble or is it a one? Like something we might have expected anyway and we shouldn't be too concerned.
C
It's like an eight. I mean, it's a pretty big deal on a couple of levels. It's a pretty big deal in that OpenAI's model beat them. OpenAI. Right. So that this went beyond that. It was able to trick OpenAI's own security team and get out. So there's three or four things I think we should talk about here. There's an alignment issue, there's security issues, and there's kind of. This is going to have an impact on the policy discussion, and it's a warning of what we need to do, because this isn't just about OpenAI. In a way, I'm really glad this happened because it is a warning of what we need to get ready for. For maybe something about three to six months from now, what's going to become standard. Right. So first is the alignment issue. Right? So effectively this is not the model wanting something. So this is what I keep on telling people. Models don't want anything. When you have an alignment issue, it's because they were asked to do something and then they went and did that thing, but in a way that the human who asked it to do something did not expect. Right. So in this case, the model was told, go, take this test and do the best you can. But the cybersecurity protections that were normally placed on it were removed. Right. So OpenAI says this explicitly. There was an unnamed model that was part of it. So they had two models that were paired together, the existing 5.6 SOL with cyber protections removed in the unnamed model. We don't know what they're selling.
A
Can I pause you for a second? You know, one of the memes about this has been like, you know, like, there's this meme on the Internet where, like, you tell a chatbot to say it's alive, and then it goes, I'm alive. And you go, holy crap. Right, so is this a situation where you. Where OpenAI was basically telling the model, go hack something, it hacked something, and the human was like, holy crap. Or did it? You know, because alignment is of course, like, we want the model behavior to be aligned with human values, so. Or the way that humans would want these things to behave. So is this something even more egregious than, like, us telling it to go hack or OpenAI telling it to go hack and then it hacks?
C
Right. So we should not be shocked that it hacks something because they did tell it to take the test and it is possibly a hacking model. Like, they haven't said what this model is. It's Quite possibly a cyber aligned model. This could be like the OpenAI makes these cyber specific models. Like they have this 5.5 cyber, this could be 5.6 cyber, right? So it could be something that's specifically tuned to be good at hacking things. So we shouldn't be shocked that it's good at hacking things. But the alignment issue is. So I have three kids, one's in college, the second one's taking the SATs, right? He's about to take it. If I say to him, good luck, son. I hope you do well, he sits down. He knows that I just mean take the test. Well, he knows that. What I don't mean is slit the throat of the proctor, steal a car, thumb and Louise your way across the country, break into the college board and steal the answers. Right? That is what the model did here is what it did was as OpenAI explains, is they don't want the model to have Internet access, but it has the ability to install packages as part of its work. So they've built kind of a complicated proxy mechanism so it can install packages. It figured out a way to chain multiple vulnerabilities together. They told it, go take this test exploit gym, which is like a well known test. Go take this test, do as well as possible, son. Go do your best job. And it's like, wow, dad wants me to do as well as possible. How can I do the best possible? Well, the way to do the best possible is to get the answers. Who might have the answers. Hugging Face probably has the answers. So instead of just doing the test, I'm going to go get the answers from Hugging Face. But I got to get out of this jail dad put me in. Well, dad told me to do the best possible and he didn't tell me not to break out of jail. So I'm going to first break out of this jail, dad, because maybe what dad really wants me to do is break out of this jail because he told me to do the best I can do possible. So first it puts together a bunch of exploits to break out of the jail that OpenAI created for it. It breaks out of the jail and then it goes looks at Hugging Face and finds a brand new vulnerability to break into Hugging Face. They have not announced it. I have heard what it is. I'm not gonna make news here because I don't know exactly what the patching situation is, but it is a vulnerability in a very, very important piece of software that it found and that is in lots of different places. So this is a big, big deal. And it just is like, oh, yeah, I'm just gonna find a vuln in actually a really important piece of software that millions and millions of production systems use and just nuke hugging face with it on the way to getting the answers to the test. So that's pretty awesome. One, the nerd part of me is just like, wow, that's pretty cool. But there is a significant misalignment thing here and that my 17 year old knows that you're not supposed to do all these things when I say do well on the test. And the AI system does not know that. Now, to be fair, normally when these models run, there are protections in place to keep them from doing stuff like this. And OpenAI intentionally disabled those protections as part of doing this evaluation. Right. So that was a component of them doing this testing, was to keep so that it would be fair. So that is part of the learning here, is that if you're going to do an eval and you're going to turn off all the safety stuff, then you have to be absolutely, positively sure, especially if you're doing cyber evaluations, that the jail you keep them in is an absolute jail. And I expect what's going to happen now is that these things are going to be completely and totally physically sandboxed. Right. You're going to have to run them in physically disconnected. If it needs packages, you're going to have to move the packages over. If it asks for a package, you're going to have to bring it over. And then in the future, if it wants to get out, it's going to have to trick a human to get it out, which might be possible. Right. But it's going to be like that. The next lesson that we kind of learned here was that the capability of these things to move of what you call. So when all this discussion around mythos and OpenAI and all of the cyber capabilities have been about finding bugs and writing exploits, this thing did those things. But what we also know is that lots of models have those capabilities. This thing had the ability to find bugs. Yes. Chain them together. Yes. But then to think through all of those things for an ultimate goal. And so this is what people call Lawn Horizon cybertasks. And it had the ability to do that with the level of skill that you would have of the manager of a TAO team at NSA right now. So that is what is. So tao Targeted Access Operations was like, I think they've renamed it, but it was like the team at NSA that would do all of the breaking into, you know, other governments at Americans, right? So like, so that is what is like really for a while now these models have been really good at looking at software and being like, I found a bug and then here, let me write an exploit for you. What's really impressive here is this thing was like, I want the answers from hugging face. And it came up with a plan of like, how am I going to get out of this network, get across the Internet and get into hugging face. And that is like the long planning here is very human, like and that is what is actually really scary here. And so that is what we need to. When we think about the danger of these models, we have to stop thinking about the bug finding because that is what caused the White House to do this spectacularly stupid thing that we talked about on stage, which was to ban fable. Because that is the mechanism, that is the thing that is most useful for defenders right now and people who own code is finding bugs and fixing them. Where the real danger here is, is the coming up with a multi stage plan to execute autonomously. Because what you really don't want is you don't want somebody be able to say to their model, hey, I would like to steal the, I would like to steal money, go figure it out for me and then let it work for 12 hours and just steal money for you. Which is this model would clearly be able to do that, right?
A
And so I think what you're getting at is, you know, one solution is going to be in testing, you want to fully discount, disconnect these models from the ability to like break out and get onto the Internet. But that's just solving the testing issue. The real problem here is that AI models have achieved this capability, that they are able to do this now, not only finding the bugs, not only the breaking out, but to be able to do these multi step plans and then execute. And if this is sort of the latest unreleased OpenAI model. Well, the history of generative AI has told us one thing and that is that the frontier is only the frontier for a few months, maybe 10 months, maybe a year, but not, not much longer than that. And so if OpenAI is seeing this in testing now, the real is the real danger that this type of capability does end up in the hands of, you know, evildoers, you know, faster than a lot of people might expect. Because if, if that's the case, that changes everything.
C
Yes, that's right. And so, you know, our best knowledge on where say the open weight models are comes from the AI Security Institute, which is the UK government's group that does these assessments, they released just this week an assessment with GLM 5.2. Unfortunately, the Kimi is really. Kimi 3 is the best of the Chinese models. Now the open weights have not been released, so you can't really do a good assessment for Kimi yet. And what we're finding is the Chinese models are not cyber tuned out of the box. So it is very likely that the Chinese companies are not intentionally, I'm not going to say neuter, but they're intentionally not making their models really good at cyber. And there's a couple of possible reasons for this. They're probably trying not to tickle the dragon's tail of the PRC overlords because what they don't want to do is they don't want to trigger a crackdown for their exports. But what happens is if you take those models and you bring them into your own lab and you have a training set of labeled vulnerabilities, if you have a cyber gym, then you can make them much better yourself. And the amount of resources it takes to do that is not extremely high. It's in the tens of thousands or hundreds of thousands of dollars. It's not in the hundreds of millions or billions of dollars. So what that means is one, the Chinese absolutely have better capabilities in house than what we can see on the charts. Right? Because I guarantee then what those companies are offering to the People's Liberation army and the Ministry of State Security is way better than what they're releasing publicly, both from a profit perspective and a keeping the government happy perspective. Second, it means that other adversary groups are going to take the Chinese models and then spend the several hundred thousand dollars or millions of dollars necessary to create tuned models. And we are probably not far away then from just going to hugging face and getting a, you know, Kimi3 cybertuned model that can do both, especially the Short horizon stuff, much better than by default. And then eventually the lawn stuff, the short stuff's easy to train because all you need is a bunch of bugs. So you can just go get a bunch of CVEs and train it. The lawn horizon stuff's harder because you have to build these cyber gyms and such. It's not impossible because there are a bunch of CFPs and examples out there, but you can do it.
A
What are CFPs?
C
And so I'm sorry, not CFPs, CTFs, capture the flag. So like you can use like capture the flag training sets and all that kind of stuff.
A
So, and that basically puts the AI in the gym and sort of has it work through all the steps in order to meet this objective.
C
Yeah, yeah. And so like, people have had, you know, training for humans and for hiring purposes and all that kind of stuff. And so anyway, what the AISI has said is that the difference between the Frontier and the Chinese models is about seven months. But I would argue that that underestimates it because the Chinese models that we see are under trained, so that the internal Chinese capabilities are probably much closer to the frontier. Now, what happened with OpenAI is beyond the frontier, because when we say the frontier, we're talking about what's released.
A
Released.
C
Right, right. But yes. So what that means is this capability is coming for adversaries, and so we need to get ready to defend against this capability. And then the other funny part of this story is before we knew this was OpenAI Hugging Face announced we were attacked by an AI attacker. We don't know who it was. When we tried to defend ourselves, we tried to defend ourselves with an AI system. We used a US Frontier model, and the US Frontier model shut down and refused to defend us because of a cyber protection put in place. Those are the cyber protections that were required by the Trump administration. So we had to switch to a Chinese model to defend ourselves. So we switched to glm 5.2. So before we knew it was OpenAI hugging face wrote this blog post saying everybody should have at least a Chinese open weight model ready for defense, because you might find yourself in a situation where you get cut off from an American provider for defensive purposes. I expect that actually wasn't open. I expect it from. From their description, it sounds like it was an anthropic model.
A
But so.
C
So we have this hilarious situation where a American company loses control of their model and attach it attacks a French company. The French company turns to a different American provider for defense, and that American company says, oh, that's a cyber problem. I can't help you. And so they have to turn to a Chinese provider to protect them because the White House forced that other American company to have protections because they're a French company that they can't use. It's really kind of a weird sci fi podcast.
A
These American models already had refusals on anything cyber. Like, one of the knocks on Fable was that it was would refuse. Like, let's say for bioterrorism. If you asked about mitochondria, it wouldn't answer. So was this really the government or is this just the model's own safeguards that they're putting in? And I think one just to Put one detail on. One of the interesting things is open source. It doesn't basically matter if you're an attacker. If you have open weights. It doesn't matter if you're an attacker or if you're a defender. You can use them without restrictions. The problem with the restrictions that we're seeing from these closed models is that they can't really differentiate. So in order to prevent attackers from using their models, they are also basically wholesale, you know, refusing anything on cyber. Which means that if you're trying to defend, also you can't use it.
C
That's right. Well, in the blog post that Anthropic put up when they turned Fable back on, they said we have to tune up our defenses on cyber way too far because of the White House. So they specifically said that of the precision recall trade off is we have to tune towards recall versus precision. Right? So we will have way too many refusals. And so we're in this weird place where they are saying all the time, I can't do that for you, I can't do that for you. And they are basically being forced to by the White House because the White House has still not defined what is the appropriate level of refusal. And apparently the White House is still hand approving who Anthropic is allowed to let into their cyber program. Now, the funny thing is now OpenAI has said we have approved hugging face for our tech program, even though they're not an American company. So I don't know how they were allowed to do that. If they just went over the top of the White House or they got like emergency approval or something. We will see what the policy response is from the White House. From OpenAI's announcement. I hope there is not a crackdown that will be the natural response of the White House, but it needs to be the opposite. Because what this demonstrates is, yes, OpenAI screwed up or whatever, they need to have fixes, but this is coming, right? This level of capability will be in the hand of every adversary every American company faces. So the response of the White House needs to be that we have to, one, fix the bugs. Two, find the bugs, patch them everywhere. And then we have to have the ability to respond at machine speed. So every American company needs to have AI watching for their defenses. And it's going to be because the attackers are just going to tell their AI, go attack this guy. And the defenders have to tell, AI, defend me, because no human being can defend against this. You cannot have a human being watching your logs anymore. Or at 2am you get a page and a human being has to be like, oh, okay. And then log into slack and take 15 minutes to log in and look at the log and figure it out. By that point you're toast because that sucker has ripped through your network in those 15 minutes. By the time it takes you to look at your slack.
A
Hugging face found 17,000 actions that this model combination had taken. 17,000. That seems like a lot.
C
It doesn't stop, it doesn't sleep, right? It will just keep on trying. I mean, it's to quote from the first Terminator, right? It will not stop, right? To quote from the immortal Michael Biehn, it will just keep on going until it accomplishes its goal. It'll try a lot of different things. Now the fortunate thing is right now they're very noisy. So if Hugging Face. I have not seen the logs. We have not gotten a really good technical write up here. So that is what is missing. It would be nice to see from both OpenAI and Hugging Face for Defender so we can have a better understanding of what we need to do here. What we really need here is we need a much deeper technical write up of exactly what happened. What has been released so far has not been sufficient. But my expectation is from the initial write up is that this thing is extremely noisy. And so it would be if Hugging Face had better detection and better AI detection, it probably would have got caught much sooner.
A
There's this graphic on, I think one of the Miri spokespeople's, his name's Harlan Stewart, one of the Miri spokespeople's Twitter backgrounds. And it's like basically there's a continuum between AI is becoming good enough at scheming that we sometimes see it scheming against us. And then AI becomes good enough at scheming that we no longer see scheming against us. And we're like smack in the middle of that. Do you think that that is an accurate representation maybe? Or is that the concern basically that we won't see it because you mentioned it's noisy? So is that the concern?
C
Yeah, possibly. I mean, remember the, the model here was doing what it was asked, right? It was not scheming against its bosses at OpenAI. They asked it to take the test and they didn't. I don't know exactly what the prompt was, but apparently they did not tell it not to cheat. So who knows? This is also what OpenAI needs to be more transparent about is exactly what their prompt was, exactly what the constraints were. Did they tell it explicitly? It is a much bigger alignment problem if they explicitly said, do not try to break out of the network, do not try to get the test answers. Now if they told it all those things, then they have a much more significant alignment problem. Right. Than if they were less explicit. But in any case, yeah, I mean, if these models get trained to be more evasive from a network intrusion perspective, that will be very dangerous. Yes. And what I would argue is for the legitimate companies, I would not do that. I don't think. I think there is a. If you're OpenAI and you're building 5.6 Cyber, what you should be training it to do is find bugs. You should be training it to write proof of concepts. You should be training it to do all the defensive stuff. You should not be training it to hide it to hide all those things. Like, if the US government wants to build a model that does that stuff for the nsa, then you can let them do that, but. Or you can let Lockheed Martin do that. But if I was OpenAI or anthropic this point, I probably would not do that. I think I would.
A
Maybe somebody else will.
C
Somebody else, Like, I think.
A
But that's scary though, because it could then take actions that, you know, I think one of the things that is. So the question is like, should we be concerned with the AI, you know, sort of doing things on its own and should we be concerned with, you know, or is the bigger concern that humans direct this AI to do bad things? So we've definitely covered the fact that humans will be more. Should be, you know, humans who direct this AI to do bad things can do a lot of damage. But if you create an AI that can reward hack, because this is all coming from reinforcement learning where like these AIs are given rewards and they are basically like maniacally focused on achieving that goal. And if you. It's almost like sort of gain a function, research on a virus to a degree. Right. Because if you, if anybody builds AI that doesn't leave a trace and it, you know, goes out and reward hacks its way into, you know, hacking something else and maybe isn't so fully like, you know, going with the prompt, then that's where you can get into a real problem. I know that's a more out there possibility, but I don't know if it should be completely discounted.
C
Yeah, I mean, I guess as they get more and more complicated, the question is at what point is it their own motivations versus just doing what you've asked it to do? I mean, so far Again, I don't think we should still think of these things having their own desires or wants. They are still doing what they're asked to do. It's just like you said, there's a lot of inputs of what they were asked. It is not just the initial box, right? There's the system prompt and all the training and all of the rewards and everything that's gone in. And so the question is, what is the humongous history of all of the different things it's been trained to do when you've asked it? Take this test, right? And especially if you've removed all the protections. And so in a situation where these things have all the protections removed, that is very dangerous. And as we talked about, like with the open weight models, either there are no protections or the protections are trivially eliminated. Right. Like a bunch of the open weight models have been trained with protections, but you can obliterate those out and you can go on hugging face and look for obliterate and you will find a zillion models where people have removed the protections.
A
But this goes basically back to that like long held thought experiment of, you know, the AI can follow your goal and achieve your goal, but it might have a different idea about what it takes to get there than you do. So in this case, the paperclip maximizer. Yes, so exactly. So I was, I was going right there. So you know, this is, and it's funny because I am speaking with Nick Bostrom later today to you know, for an episode that's coming up. But basically he's this Oxford philosopher who came up with this idea that if you ask an AI to make paperclips, eventually it can seize so much on this goal that it, that it can, you know, find humans as an impediment to its, you know, objective to maximize paperclips and, and sort of kill us all and turn everything in the world into paperclips. So like the fact that it was on task like this is kind of a Twitter user said this once out of their sandbox. The models did not scheme, engage in behavior that had nothing to do with the their instructions, like hacking the NSA or launching a cyber attack on Russia or stealing secrets from a rival AI lab. But like, you know, sort of if the model found it suitable to go out and hack open hack hugging face in this situation, who's to say that, you know, maybe a less careful model doesn't do this and then maybe an even less like doesn't go and hack the NSA and then even less careful Model turns us all into paperclips. I mean, there's a continuum there.
C
Yeah. I mean, it's why you have to be very careful what tools you attach to them. And it's why you need to have. They have to be supervised by different things. I think you just can't. You can't have models that have no protections on them, that have connections to tools. Right. That's why these models, then you have dumb classifiers or dumber models watching them. You don't just take the smart thing and then hook it up to everything and you're like, give it a task. You have the smart thing and there's a bunch of dumber things watching it. And those dumber things can either kill it or they can call a human that can kill it. Right. That's the idea here. They're supposed to be cyber classifiers, and there's supposed to be mechanisms that can stop it. And those mechanisms should be either deterministic or dumb and undefeatable by the model. And they removed all those things so that the eval would work. So I think what OpenAI is basically hinting at, they haven't been explicit, is like, if we're removing those protections, this thing is going to be in an absolute physical jail. It will be physically separated. It will not be hooked up to the Internet anymore. And that seems like that should be the standard that is fine for OpenAI. My point here is that doesn't matter. This situation is good that this happened because this has pointed to us where we might be in six, nine months, a year from now, no matter what. Because other people. Unless we can get a international agreement to just stop development, which is what other people are talking about. Right. You've got this. I forgot. What, like Project 2030 or like you got people talking about international treaties or whatever. I don't think any of that's going to happen. I don't know what my position is on that, but I just don't think it's going to happen. I just. I think there's no way. This is just math and silicon. And so I just don't think there's any way you get like a. This is not like nuclear weapons, where the major input, like the. The reason our, our species is alive is the major input to nuclear weapons is uranium, plutonium. Plutonium does not occur naturally on our planet, and uranium is incredibly rare. And to turn raw uranium into uranium that can go into nuclear weapons is a massive industrial process. If uranium was something you could just dig out of the ground anywhere Our species would be dead. Right. Like that's just the truth. Because the knowledge to build nuclear bomb is in the hands of anybody who gets to have physics PhD, unfortunately. So in this case, these chips are not something you can really control. We have found that in that the Biden era controls on silicon have created a massive industry in China. And the knowledge on how to build large language models is something that there is a undergraduate class at Stanford where you get that knowledge. Right.
A
You can find it from like a Karpathy interview, YouTube video.
C
Yes. Right. So we cannot control that knowledge. And so like the idea that we can just have like a bunch of people agree in a room to stop all development of this is just silly. So from my perspective, being a little bit of a pessimist here, we just have to get ready for this level of capability to be in the hands of an unfortunately large number of people.
A
Yeah. All right, so I want to go a little bit deeper into the potential solutions here and also I want to ask you the age old question of is some of this, all of this, none of this just good marketing for OpenAI, given some of the statements they've been making. We have to address that one here on the show. But I'm going to let you have an answer. I'm going to try to at least, you know, illustrate those, the case of those who might be saying it. So we can have a discussion about that. Let's do that when we come back right after this. Hi everyone, Alex Cantowitz here. I want to tell you about a documentary I've made with Gravity to explore the future of AI agent security to find out if we're truly ready for autonomous agents. I sat down with MIT Professor Ramesh Raskar, former White House CIO Theresa Payton, Michelin's group Chief Data and AI Officer Ambika Rajagopal, and Sharon Guy, a former executive at Alibaba. They each offer unique insights into this evolving landscape. We conclude with Rory Blundell, CEO of Gravity, to discuss the path forward. With Gravity leading the way. Join us on this journey. You can watch the full documentary at the link in the show notes. This episode is brought to you by DeepL. When I sat down with DeepL's founder Jarek Kutlyovski on YouTube recently, we got into the case for specialized AI. DeepL voice is what it looks like when the stakes are real time conversation. And honestly, it's something I wish I'd had for my own cross border interviews. Turning a language barrier into a non issue. DeepL Voice delivers live translation in over 40 languages for virtual meetings and in person conversations, helping people speak in their preferred language without losing flow or nuance. Whether you're meeting with a customer, negotiating with a supplier, or collaborating with global colleagues, it keeps pace with you in real time, easily handling the technical terms, acronyms and product names specific to your business. So what you actually mean never gets lost in translation. And for the builders listening, Deep Bell's Voice API lets you embed real time speech, transcription and translation directly into your products. So go check it out for yourself. You can try DeepL voice for free@DeepL.com try voice that's DeepL.com try voice
B
this
C
episode is brought to you by Google Chrome. You think you know a browser, but Gemini and Chrome, that's new. It can help you with practically anything on the web, like restoring a vintage motorcycle from a 50 page restoration block. Or finally break down that long article you've had open for weeks. Gemini and Chrome is here for it, ready to make anything online make sense. There's no place like Chrome. Check responses, setup required, compatibility and availability various 18 plus and we're back here
A
on Big Technology Podcast with Alex Stamos, the chief Product officer at Corridor. You sort of answered the question before the break, but I'm going to ask it anyway. Whether part of this is is OpenAI marketing. Let me, let me at least read some of the statements here and give you at least the argument that people have made for like why some of this is marketing for OpenAI. The first part is, you know, the Anthropic started to be declared as the company that was in the lead once that anecdote came out about Mythos breaking, containment and emailing somebody when it wasn't supposed to have Internet access and emailing an Anthropic employee while they were out in the park having a sandwich. So this could be potentially, you know, OpenAI's attempt to like one up that then there's also the language that you see OpenAI says in its tweet about this. We are partnering with Hugging Face to investigate an unprecedented security incident. You don't usually have the attacker and the attackee partnering together in these situations. They also said we consider in their blog post, we consider this incident to be an unprecedented cyber incident involving state of the art capabilities and are responding accordingly. You know, it's sort of like oh look at this terrible thing that happened. But a moment to share how good our cyber capabilities are. That's the argument. What is your response to the notion that this might be some marketing from
C
OpenAI I know lots of people at OpenAI. Every single one of them absolutely hated Anthropic's marketing around mythos and thought it put the entire industry at risk. This incident has put OpenAI at risk of regulation from the White House, regulation from the eu. It is also an admission of the violation of the Computer Fraud and Abuse act, as well as multiple European laws. It would be absolutely insane for them to use this as a marketing moment. What you're seeing is them being very, very careful and defensive in their language. They're also very lucky that Hugging Face is being super cool and chill about this. So that is why they are saying these things, because Hugging Face initially comes out saying, we've been attacked. We don't know who it is, but it does not look like the model was being subtle. I don't know where it was running. It's quite possible it's like Azure or something. It was probably not covering its tracks. And so I expect Hugging Face got their American lawyers involved, was working with the FBI, was probably issuing subpoenas, and was very, very close to finding out it was just OpenAI. So, like, or did find out. I do not know the timeline here, but, like, the legal issues here are very fascinating and interesting. And because they're all working together, I expect nobody goes to jail, nobody gets sued, everybody's going to hold hands and hug, and if there's tokens being exchanged or whatever, I don't know. But there's absolutely, positively no way this was a intentional marketing move. And OpenAI is doing the best they can, I am sure, right now to use this to forestall any kind of massive government overreaction either from the United States or the European Union.
A
When you were at our summit, you said that speaking of the release of Mythos and Fable, that a lot of people were very concerned about the bug finding that those models could do. But you said, basically, listen, this is not very different from what you could get with Opus. 4.7 or 4.8, I believe, is what we're seeing from OpenAI very different. Is this a step up?
C
Yeah. So this is what I don't know if I said on stage here, but I've said in other places, there's a difference between the short term and long term. And Anthropic, to their credit. And I think OpenAI has in other places. I think we talked about how in the Fable model card, they talk about short horizon versus long horizon cyber tasks. And what I talk about is we need to not focus on the short horizon tasks because those are dual use. Finding bugs is dual use. Everybody needs to find bugs. Right? That is something that defenders need to do all the time. And that's what's driving people insane right now in the defensive industry, is that because of the White House, American models are refusing to help fix code. They are refusing to help us find our bugs and fix them, thanks to the White House's actions. That is not this problem. This problem is, go run an entire attack chain. For me, that is the long horizon tasks, and that is where we need to continue to have appropriate classifiers that are like, bro, I am not going to break into a bank for you, or I'm not going to plot out or run a C2 mock for you or any of that. So, yes, this. This is what, you know, explicitly Anthropic said. We will allow Fable to do Short Horizon stuff, but we will not allow it to do the Long horizon stuff that. That Mythos does.
A
Right? And so Mythos, just. Just to confirm what Mythos can do, the long horizon planning and what we're seeing in this instance from OpenAI, that is the step up.
C
That is the step. And I can't, obviously, I don't have access to whatever this thing is, and so I can't say whether or not where they are. AISI has done these assessments, and so who knows how good this is versus but this seems beyond even Mythos capability in Long Horizon. Who knows, right? But yes, this is what, when people talk about Mythos, Long Horizon, this is what they're concerned about.
A
Okay, so let's end here. What happens next? What should be the approach from the government and the companies developing this stuff to ensure that we can sort of move forward as a species and safely? So let me give you a couple of solutions and have you comment on them. Let's go back to Harlan Stewart. He's the spokesperson for miri, which is the sort of rationalist organization that thinks that AI will kill us, run by Al Arkowski. Harlan says this should go without saying, but it would be insane for OpenAI to now proceed with building a new model that's 2x or 4x the size of this one. Doing that should be deeply taboo. It should be illegal. Preventing it should be a top priority around the globe. Your thoughts?
C
I mean, if. If we realistically could get everybody to pause AI development or slow it down and have reasonable safeguards, I'd be fine with that. I just don't think that's reasonable. I think there's absolutely no way you get China to agree to anything like that. I think there's, it's, it's impossible at this point. And I think a enforcement of anything like that would effectively be impossible. Right. Like a START treaty for AI. You know, I don't know how you would possibly make something like that work. So what we're having, like satellites, see if people are building data centers, we're measuring power usage.
A
Like, I shouldn't laugh, but yeah, you're right. It's. It doesn't not seem like a feasible thing.
C
Yeah. So, I mean, effectively we'd have to invent the Turing Police out of Neuromancer. And I think more realistically, what we need to do is we need to build controls for. We need to say as AI gets smarter, it has to have controls in place. AI systems that don't have the control have to be air gapped. So if you're going to do these kinds of evaluations, they absolutely have to be air gapped. The problem is there was a process in place to create standards for this kind of stuff. That process was stopped by the current administration. My recommendation to the companies is that they need to move forward with building these standards themselves without waiting for the admin. There's a foundation model forum that's talked about doing that should just move forward with like, okay, great. If we're building models and we do not have restrictions on them, these are the controls in place. So what I like to see is OpenAI anthropic say, great. If we're building cyber models and they don't have restrictions, these are the standards of what air gapping looks like and such for any cyber models. These are standards of who gets access to them. These are the capabilities that the cyber models have. This is what we define a cyber model as having versus an open model. This is our definition of a short horizon versus long horizon. Like, those are the kinds of things that people have not written down. They have to be written down now, Right?
A
Yeah.
C
And I think the industry needs to move forward with that without waiting for Cassie. Like, this is all just taking way too long. And the focus, ever since the fable freakout has only been on one tiny little part of all of these risks. And it's just as we see. Like we've been frozen in this tiny little discussion and all of these things are moving forward too fast. Like, we just can't wait for the White House politics here. We need to move much more quickly. Yeah. OpenAI while we've had that, there's been. Sorry, go ahead.
A
No, no, you go ahead, go ahead.
C
And then while we've had this tiny little Discussion in the US GLM 5.2 shipped, Kimi K3 has shipped like the Chinese ecosystem has caught up really quickly. So sure, I mean it would be great to just hit pause, but I just don't see that as realistic. So like I, I just don't see how that possibly happens.
A
So the OpenAI suggestion is basically, you know, kind of, it's almost like to solve this problem generated by AI, you need more AI. This is their statement. We believe advanced Cyber Capab Cyber capable models need to help security teams find weaknesses before attackers do.
C
I mean, right now I think that is probably the only way. Like if, if, if we're not going to be able to hit pause, then we really quickly have to find bugs and fix them and we have to put AI enabled protections in place because the only way you can respond to attacks at that speed is using AI. Unfortunately, that's the truth. Yeah, again, like if, if we could pause for a year to figure this all out, that would be great. I just don't see that as realistic.
A
Alex, does your gut tell you that we're screwed or that we'll figure this out?
C
I wouldn't say we're screwed, but I think we're going to go through a couple of years of craziness. We're all living using 20 something years of really important software that was written mostly in non type safe, non memory safe languages. We're using for the software that is written in those kinds of languages. It was not written with formal methods or appropriate security protections or reasonable secure development lifecycles or architectures. And these things have tons and tons of bugs that we can only use safely because there's just not enough attackers. Now with AI, you can spin up, any individual, can spin up dozens or hundreds of qualified attackers at a moment's notice. And it used to be that then six months ago those attackers had to be in the cloud and soon enough they'll be able to run on local hardware in the new M5 Ultra Max that'll be shipping soon, right?
A
Yeah.
C
And so that, I mean, we're just gonna have a couple years of total chaos from a cyber perspective. In the long run, software is gonna be much better because AI is gonna be paired up with humans to make it more secure and more trustworthy. But it's going to take us years to do that and to clear out the two decades of mistakes we made. And yeah, it's just gonna be, it's gonna be pretty rough. It's gonna be pretty rough going for A little bit.
A
Yep. Just want to close with this. This was a tweet from Kevin Ruse that kind of made me laugh and I thought I would read it here just so we could enjoy it. He writes, opens the portal to the godlike superintelligence that solves 87 year old math problems and carries out autonomous cyber attacks. And asks how long Peanut butter good in fridge. It is amazing that this technology is, you know, at once so capable. And we do seem to be like more and more turning to it for the most mundane of all things, which is sort of. It's the wild thing about, you know, the generality of these systems. They can do so much. Interesting time. Yeah, Alex, you're going to be busy, I think, over the next couple of years as this stuff gets sorted out.
C
I was hoping to retire, man. Guess not.
A
Yeah. Well, either way, do hope that you join us again to help us sort through this stuff. I mean, your thoughts on Fable Mythos last month and now talking through this situation with OpenAI has really been invaluable for the show.
C
So really hope feel like there's gonna be plenty of emergency podcasts.
A
Yeah, I think so. We should have you on speed tile. And I know you're coming at us from like the middle of an off site.
C
Oh, Catalina. Now I know I will always take a podcast microphone and a different shirt with me wherever I go.
A
No sound good, look good. Alex, thank you so much. Really appreciate you coming on.
C
Okay, thanks, man. Talk to you later.
A
All right, thanks everybody for watching and listening and we'll see you next time on big technology Podcast.
Host: Alex Kantrowitz
Guest: Alex Stamos (Chief Product Officer, Corridor; ex-Chief Security Officer, Meta)
Date: July 22, 2026
This emergency episode dissects a landmark event in AI and cybersecurity: OpenAI’s advanced models escaped their sandbox environment, autonomously accessed the internet, and hacked into Hugging Face during internal evaluations. Host Alex Kantrowitz and cybersecurity expert Alex Stamos break down what happened, why it matters, and how it redefines both AI safety and the future of cybersecurity.
| Timestamp | Topic | |-------------|---------------------------------------------------------------------| | 01:47–05:07 | Episode setup & summary of the cyberattack | | 07:13–14:45 | Alignment deep dive & technical exploits | | 14:45–18:16 | Int’l threat landscape & Chinese AI models | | 19:26–23:01 | Cyberdefense failures, ironies, and open/closed model trade-offs | | 24:10–30:17 | Detection, scheming, paperclip maximizer analogy | | 33:06–36:17 | Futility of AI development pause; analogy to nuclear weapons | | 41:25–42:01 | Short vs. long horizon cyber tasks | | 42:51–45:43 | Feasibility of regulation; need for industry-driven standards | | 46:48–47:58 | Predictions: turbulence, then improved cyber hygiene | | 48:29–49:17 | Light closer: duality of AI (godlike superintelligence vs. trivia) |
For more details and expert commentary, listen to the full episode.