Whereas we're really bad at parsing long text fragments. And so a lot of our methodology and our early techniques are done on vision models. And there's precedence to this too, where a lot of the original work done to introduce the core concepts of features and circuits in the field was done at OpenAI with Chris Olah and Nick Cammarata. Nick Cammarata is on the team at Goodfire now, and it was all done on vision models in the very early days.
All right, so let's go back to the paper a little bit, and in particular, how you caught cheating. We talked about what a probe means. In this case of the paper, what did the probe actually see? And I think you have a method called difference of means. Do you want to explain what that is?
Yeah, a difference of means vector is actually much simpler than a lot of the production probes that we train, because probes can have all types of architectures that make them more or less effective. But the difference of means vector is actually just the simplest possible approach to extract some type of concept from a model, which is why it was surprising that it works so well, because we can still hill climb and make this way, way better in a production deployment. Basically, a difference of means vector is: you give the model two datasets.
One dataset is the cheating dataset, and one dataset is the not-cheating dataset. You take a look at the activations, you subtract them, and then you average them. That's it. Easy. And you take this, and then it can detect cheating, which is a very surprising thing, especially because in our paper we showed that this was done off-policy, which means it was done in this specific paper on really contrived scenarios where the model is set up to think about cheating. And this transfers to much more general, complex actions that the model is taking.
And that was also very surprising. We can train it off-policy, and then it generalizes to on-policy.
And to play it back, it's almost like taking an MRI of the brain. Is that fair? So if you don't take an action, and then if you take an action and you sort of light up the parts of the brain that correspond to the action?
That's right. Yeah, you can think of it that way.
And then we were talking about steering a minute ago, and you'll steer the model, asking it to write a story about a girl taking an exam. Do you want to describe what happened?
Yeah. Well, maybe even going back, you can think of it like an MRI of a brain, but then you actually get a button at the end of it where you can push, and then you can light up that region again, or you can remove that region. That's the difference of means vector, where you now have this tool that you can intervene causally inside the brain. And that's maybe one of the big differences from studying neural network interpretability versus neuroscience, where you can't really intervene in a human brain without massive amounts of damage.