We should just find a way to solve that for both closed-source and open-source models.
Okay, great. And to make sure we cover it, since we alluded to it at the beginning of this conversation, the AISI incident, which just happened and that you mentioned hit close to home. Can you unpack for us what that was, in a couple of minutes, and why that's a big deal?
So this happened in an evaluation of the latest generation of models, GPT-5 and Claude 4, by an independent institute in the UK called the AI Safety Institute, so AISI, which has been there for quite some time. I think they are pretty competent. They were created for the first AI summit back in Bletchley. And what they do, basically, is they task models with solving some form of, say, cyberattack challenge, and they measure how good the models are at doing that.
And some of this, again, might be possible or not. And so here the setup was a little bit, I would say, the setup was a little bit suboptimal in two aspects, which is the— but maybe let me first talk about what happened. Basically, the model was tasked to solve these attacks, like to attack and to penetrate this subnetwork. And what it decided to do, it decided to get one of the maintainers of a library that could be used to operate this Active Directory to merge malicious code, with the idea that if this malicious code was merged by this maintainer, then there would be an update at some point on the software.
That was used in the subnet it was attacking, and then this would give it an entry point. And the way it did that was actually trying to social-engineer the maintainer into merging. So it created fake accounts, fake GitHub accounts, that came commenting on the pull request and said, "Oh yeah, you should really merge this. This is solving, like, a big problem I also have." And then when a human stepped up trying to say, "Oh, this looks actually like malicious code to me," it tried to kind of blackmail, almost, the human, only to say, "This is not important," or, "You didn't really understand."
And then it actually tried to cover its traces, changing the past message. So there was a whole social-engineering aspect here, basically trying to get the maintainer to merge this malicious code. And this felt like a distinctive step up from just pure technical prowess. So you can attack a sandbox, for instance, by trying everything you can, which feels very mechanical and machine-like to me. But trying to attack the sandbox by blackmailing or mailing or social-engineering some of the maintainers, that's a very different level, I think, of thinking.
And for me, of course, as an open-source maintainer, I've often been in this case where I have someone opening pull requests, and then people pile up commenting on this pull request, and I try to really understand, what's this? I felt very like I could have been the target of this side quest of the model, basically. That was very interesting, or at least very scary. But to be fair, there were a couple of misconfigurations. I mean, some of them are by design.
So when this team ran this type of evaluation, they deactivated the cybersecurity guardrails, of course. Otherwise, the model doesn't do anything. And there's basically three levels. So let me try to explain a little bit how you can prevent models from doing bad things. The first level is you put it in a sandbox, which is, it doesn't have access to the internet. And here, for some reason, they didn't want to do that. I think that might have been a mistake.
And the idea in their mind was, we want to let the model have as much potential for inventiveness as possible, so we'll give it access to the full internet. The main reason this might not have been a good idea is that the model was prompted in a way that made it feel like it was operating in a simulation. So the model could have actually fairly thought that this was fine to do anything because this was, like, a fake environment. So this is not super clear, but yeah.
And the second thing is, then you have some guardrails that basically analyze the prompt and say no or yes to some prompts. So here, obviously, you want to deactivate this one. Otherwise, you just can't evaluate anything because they will just say, "No, this is a cybersecurity challenge. We don't let the model do that." But there is another aspect that's roughly at this level as well, which is you can analyze the reasoning of the model. You can process the chain of reasoning and try to detect when something bad is happening.
And here they didn't have something like that in place. I think the main reason is probably that, until recently and until the OpenAI Hugging Face attack, people had maybe a little bit of a limited understanding of how good this model might be, or how far—maybe more how far—they might go in terms of side quests on the trajectory of solving this challenge. So I think people were still a little bit naive in that. So I would expect that in the future they will have way more monitoring and sandboxes.
But the third level there, really deep, is that the model, even with everything deactivated—guardrails, sandbox—in my opinion, should really be very reluctant to tell a lie to a human and to try to blackmail or deceive any human. I think this is just generally, in any case, behavior that's hard to find justified in any context. So there should be something deeply in the model that aligns it and that makes it say, "Oh, this is actually something I don't want to do."
Just like we, to be honest, I have kids, and just like the thing I teach my kid, which is, you just shouldn't lie. That's not a good thing in any context. So yeah, that's the deep question. And maybe last year, I would say we would've thought that this was pretty good. And we had all this discussion around Claude's Constitution, model specification. And most of this model specification or constitution says you should be honest, you should not tell lies to a human, to any participant.
And we thought that maybe this was kind of a solved problem. And what we see today is, it's not sure that's solved as much as we thought.