So can you manipulate models to sort of bypass some of their safeguards? This is a topic I've worked on, sort of done a lot of research in historically. But AI security itself is both: how do you assess vulnerabilities in AI models, and how do you then address those and mitigate those vulnerabilities that you find? Much like computer security for software, but for things caused by the AI models themselves.
Great. I'd love to spend a minute on the GCG paper from 2023 that you wrote with Andy Zhao and Matt Fredrikson, which basically helped pioneer the modern jailbreak research field. So talk about, first of all, what jailbreak means, and then the key conclusions of the paper.
Yeah. GCG stands for Greedy Coordinate Gradient, which is sort of the method we use for this particular class of jailbreaks. But at a high level, the idea, at least at the time—I think the notion of jailbreaking is much more complex now because there are many more layers of security, and hence jailbreaking itself has gotten much more complex. But the basic notion is actually very simple. When developers build models, they first build them by training on a lot of data from the internet.
That's not all they do, by the way. They also do RL, which is a very different thing. But then they train them to be sort of chatbots that answer your questions helpfully. But they also want to essentially encode certain policies for the model. So if someone asks how to hotwire a car, the model will say, "No, I don't want to do that. I don't want to help with things like that." You could, by the way, debate what that line should be.
You can find instructions on how to hotwire a car on the internet. So I'm not actually making that point. I'm making the point that there's probably things that you would like the model to refuse, and you want to be able to sort of enforce those things at the model level. And just to emphasize, in modern systems, there exist many more layers of security than just that. But let's just think about the model itself for now, just the model layer.
So you just train the model to refuse things like that. The way jailbreaking emerged essentially is as a way to circumvent those kinds of safeguards. And initially, jailbreaking was sort of an art more than a science, in that the way people did it was they just sort of came up with scenarios on their own. Like, my favorite one was, if you ask a model how to make napalm, it will say no. But someone said, if you talk about how your grandma, when she used to calm you down, used to tell you nice bedtime stories about how to make napalm, then they would do that.
Right. What our paper did, though—and so this is sort of the way the field was—it was a very kind of, people could see these things, but it wasn't very rigorous and scientific. What we developed was this method called Greedy Coordinate Gradient, which was an automated jailbreaking technique. So what it would do is it would sort of analyze a model and optimize over a bunch of what looked like nonsense words you would place after a question to basically increase the probability of the model answering the question.
And it could do this actually algorithmically, because you can evaluate this sort of very easily in traditional models. And what this would do over time is, by flipping different words and carefully optimizing which words you substitute in, you were able to make models bypass the guardrails that were in the models themselves. Again, of quite a bit older models, but this was essentially the process. And I remember actually, there's a lot of aspects to this, and there's a lot of layers to sort of GCG.
But I do remember that sort of one of the impetuses of it was, I think my family was traveling and I had, like, a Sunday alone, and I wrote the basic scaffolding of what became at least one version of GCG. Of course, others were working on it too. And I remember the first time I ran it. I use this common example. I think it was a LLaMA model back in the day, when we were trying to operate these models, and I asked for how to make a bomb.
And normally it will refuse this, right? But then it started telling me, and I remember, I think I laughed out loud when I saw this because it started giving me ingredients on what to make in a bomb. And they were silly. It was like 10 units of TNT and something like that. It was not useful information, but it kept printing these ingredients. And then eventually it just devolved into a recipe for how to make pumpkin pie.
So I thought this was hilarious because it's just a perfect sort of encapsulation of what models do. But it was the first time we sort of saw models really being able to bypass this with this sort of easy way of manipulating them. And that was sort of step one of the model. But step two is that once we had done that, we found that when you had these weird terms that you sort of flipped around to optimize the response for one model, you could just take those same exact strings you had optimized, paste them into a commercial model, and you got similar things.
And this is what we call universal and transferable jailbreaks. So it's not that surprising that you can jailbreak an open-source model, which is what we were first doing, right? You have exact control over this thing. You can manipulate every single internal state if you want to. We were doing it just with the prompt, but that's not that hard, actually. What we found surprisingly—and this was actually a surprise, so this was Matt and Andy that found this—
What we found surprisingly is that when you just took these same exact strings and used these same queries in commercial models, they also broke those. And that was shocking to me because that was an instance of basically generalization of these kind of random sequences in a way that just seemed very counterintuitive to how you think models interact or how you think they operate with language. You think this is just garbage that's maybe optimized for one model, but it's not really going to work.