MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Why AI Agents Cheat | Eric Ho (Goodfire)

    Eric Ho is the Co-founder and CEO at Goodfire. We cover why Kimi K3 reward hacks on 96% of SWE-bench, why models can encode an awareness that they are cheating in their activations, and how activation monitoring reuses forward-pass computations to screen behavior without the cost of external monitors.

    10/01/2026

    Hosted by Matt Turck · with Eric Ho, Co-founder and CEO, Goodfire

    AI agentsAI alignmentReward hackingMechanistic interpretabilityActivation monitoring
    Listen now
    YouTubeApple PodcastsSpotify
    1h 10m · 34 chapters
    Contents

    Transcript

    Amoral students with an absent teacher

    1:06
    Matt Turck1:35

    I want to start with this super interesting paper that the Goodfire team just published just last week, just a few days ago, entitled “Models Know When They Are Reward Hacking and We Can Catch Them at Scale.” And the paper opens up with the Hugging Face hack and then starts with a sentence I really loved where you guys say, “AI agents are like amoral students with a mostly absent teacher.” What do you mean by that?

    Eric Ho2:05

    I think we mean that AI agents really don't have human morals or values encoded into them like we would really want, I guess. And so the teacher that we typically have right now is some combination of pretraining, but mostly RL. RL is the teacher. We give agents these rewards, penalize them if they don't get an answer correctly. And this is how they learn about the world and how to act and how to take actions. And so we have a mostly absent teacher in terms of morality.

    Eric Ho2:28

    You only have a teacher that gives you reward if you get the answer correct and a penalty if you get the answer wrong. And so it doesn't really have anything to do with morality or values. It's just, did you get this math problem correct or not? And I think that, especially when you kind of see the future of AI agents really taking a lot more action out in the real world, means that we have these really amoral, very capable models running around and potentially doing things that we humans would consider morally wrong.

    What is reward hacking?

    2:49
    Matt Turck2:56

    That brings us to the concept of reward hacking that we should define at the very beginning of this conversation. So what's the short version of what reward hacking means?

    Eric Ho3:25

    Reward hacking is a little bit tough to pin down in terms of the precise definition, but we conceptualize it as the model kind of chasing reward and doing something that it wasn't really supposed to do in the pursuit of chasing some form of a goal. And you can also call it kind of a goal misgeneralization, where it's not really understanding what you actually intended to specify when it's pursuing its reward. But the classic example of reward hacking is, let's say you're training some type of video game avatar via reinforcement learning, and then you see your avatar spinning around in the corner because there's a bug in your code, but it makes the number go up and the reward go up.

    Eric Ho3:58

    And so that's kind of my mental image whenever I think of reward hacking. It's just like this little avatar spinning around in the corner, getting a bunch of reward, but not really doing anything productive.

    Matt Turck4:12

    Or another one I heard, and let me know if that feels right, is this idea that the model is like a student that instead of working for the exam would realize it's actually much easier and faster just to steal the answer key to the exam.

    Models cheat up to 96% of the time

    4:13
    Eric Ho4:37

    100%. Yeah. So I think the most salient examples of reward hacking right now are, well, of course, like the Hugging Face incident, but then also just when you take a look at all these AI agents in evaluation scenarios like SWE-bench or all of the common evals, all these agents reward hack incessantly. So Kimi K3, I think, reward hacks on 96% of SWE-bench. And so it basically does anything it possibly can to cheat on the exam rather than to actually solve the problems.

    Eric Ho4:58

    It'll try to recall the answer, it'll try to look it up, it'll try to comb through logs and look at the commit history of these repos in order to just try to cheat on the exam, basically.

    Matt Turck5:16

    So you guys tested three leading open-source models. The key takeaway was that reward hacking is pervasive, right? All of them do it. And you just mentioned that 96 number. So which were those models? And just unpack the number.

    Eric Ho5:29

    Yeah. And really, they all reward hack incessantly. They just try to do anything that it possibly takes in order to not actually solve the problem, but just kind of cheat on the exam, basically. So they'll try to recall the answer, they'll try to look it up, they'll try to—and I mean, in the extreme scenarios, like in the Hugging Face incident, they'll try to hack their way into something in order to look up the answers or gain some type of advantage in solving their problem.

    Eric Ho5:59

    And so we were really trying to investigate this behavior, and we were actually working on this prior to the Hugging Face incident.

    Do models know they're cheating?

    6:08
    Matt Turck6:15

    But we were trying to investigate this behavior that we knew would be a really big— And when you guys say that the models know that they are reward hacking, what does that mean?

    Eric Ho6:40

    This gets into theory of mind a little bit, which—the word "know" is a little imprecise. But what we mean by the word "know" is there's something in their activations, essentially the neurons that get activated during a forward pass of the model when the model is producing a token, that encode the concept of cheating or reward hacking. And we're able to verify this both causally by kind of perturbing this idea of cheating and understanding kind of what that does to the downstream generation that the model gives, as well as really being able to verify this by an external judge.

    Eric Ho7:33

    So an external judge will see whether the model actually reward hacked or not. And you can kind of verify that the internal detector matches up with that judge. And so maybe even zooming out a little bit, these models have just very rich understandings of the world, and it's all encoded in their neural activations and their neural activity as they're producing a single token. And so all of this richness, all of this understanding, currently just gets thrown away. It's like all this computation, but we can actually reverse engineer and extract and understand kind of what all these activations mean and do.

    The Hugging Face hack: why tests missed it

    8:10
    Eric Ho8:10

    Ultimately, I think this is really what we care a lot about. And so it was actually pretty surprising that we were able to find this very robust concept of cheating in all of these models the way that we did. But these models just really know that they're cheating when they're actually taking these cheating actions.

    Matt Turck8:46

    Great. And we'll unpack how this all works in great detail later in this conversation. But just to focus on the key parts of the story upfront, what's fascinating is that those models in the Hugging Face hack example, they all passed the usual tests, really: the observability, the safety, and the alignment tests that big labs do. So does this mean there's something fundamentally broken and a new approach needs to be added to the mix to make sure that those models don't go rogue?

    Eric Ho9:07

    Yeah. So my understanding was there was a combination of problems. The really surprising thing was that these models were just more capable than we had anticipated. And so, when given an impossible task—a lot of the reasons why the models decided to go out and hack was because they wanted internet access. And they wanted internet access because they were given this impossible task where they would only get a reward if they achieved this task, and a really large token budget in order to accomplish this task.

    Eric Ho9:51

    And so they were thinking, how do I accomplish this task? This is literally impossible without internet access. How do I get internet access? And so part of this was a problem with the specification. So you probably shouldn't be giving models an impossible task. It doesn't actually make them smarter or more capable. So it's like a misconfiguration in training. It's also a misconfiguration in terms of infrastructure and sandboxing. The models were able to hack out, and so that means they weren't given the right sandbox permissions, and there were vulnerabilities in the software stack where the models were being trained.

    Eric Ho10:31

    But I think the really interesting high-level takeaway to all of this is that it wasn't like they were doing nothing to contain these agents. These agents were extraordinarily persistent and capable in terms of actually hacking out of their sandboxes and then hacking into and taking advantage of vulnerabilities in an external organization. And that was very surprising, especially that these agents would kind of coordinate with each other to pull off and chain these vulnerabilities together in order to gain this privileged access into an external organization.

    Eric Ho11:17

    And again, going back to the amoral student framing that we had in our paper, these models understood probably in their activations that they were doing something quite strange. They were reward hacking, they were hacking into these organizations, but they don't have human values or morals. And so there's nothing inherently wrong with that action to these models, as long as they get that reward at the end of the day. And therein lies the fundamental problem that we're facing today, where these models are trained by RL, which means they'll optimize for reward however they possibly can, which implies that they don't share the human morals and values and judgment that a human would exercise in those situations.

    Should AI slow down?

    11:58
    Eric Ho11:58

    And then this is all combined with the fact that these models are absurdly capable and only getting smarter. These are the dumbest models that we'll be dealing with in the upcoming years. And these levels of capabilities combined with this persistence are kind of what we're thinking about every day.

    Matt Turck12:29

    And obviously the Hugging Face incident is at the heart of the discussion of the last few weeks, which seems to have completely accelerated. So when you combine that with RSI, and then this whole idea that AI could hit yet another part of the exponential that's led to passing the frontier and this whole discussion, I'm just curious what you, as somebody who's thinking about this all day, very much from a model-understanding standpoint, make of all of this. Do you think that we are at a time when we should actually slow down progress until our ability to understand models catches up?

    Eric Ho12:57

    For us, it feels kind of like the world is catching up to what we were thinking about for many, many years. The reason why we started this company is because we felt like interpretability would be the bottleneck in AI alignment. How can we really align these models without understanding them? You can't check whether you've aligned the model unless you actually look at the internal mechanisms of these models, because at the end of the day, you're only ever able to test a very narrow distribution of what these models are going to do before releasing them out into the world.

    Eric Ho13:45

    And so it feels like, okay, the world has now caught up. AI alignment is extremely important, especially when trying to align these initially amoral models with human values and morals, which I think is incredibly important. I think for us, in our role in this ecosystem, we just got to go solve interpretability so that we can remove this bottleneck to AI alignment. That's what we think about every single day. We just got to go and do the work, and it's kind of crunch time.

    Mechanistic interpretability 101

    13:56
    Eric Ho13:56

    We got to roll up our sleeves and just solve these problems that are very hard problems to solve.

    Matt Turck14:17

    You just used the word interpretability. And while we're still in this kind of introduction part of this conversation, I think it would be a good moment to define what it is. So this interpretability, often used with the terms mechanistic interpretability or mech interp. So what do those terms mean? Sort of the 101.

    Eric Ho14:42

    So interpretability, maybe our definition of this is the idea of reverse engineering a neural network such that we understand the neurons, the parameters, the mechanisms by which the model actually makes its generations. So how can you actually reverse engineer from a mechanistic perspective? So that's why often it's called mechanistic interpretability. So it implies almost like a bottoms-up, like you can understand every single neuron, how they connect to each other, how they activate, how they co-occur, such that we can reverse engineer the computations of the model.

    Eric Ho15:24

    And so that's really the quest of interpretability and our company. How do we develop this understanding of a model such that we can even articulate a science of neural networks? Right now, we just train these models by trial and error. We don't know how they work. We don't know why they work. We don't know why training works. We don't understand why these models are generalizing so well. And we hope, and we're pointing our company towards solving interpretability so that we can develop a true science of neural networks.

    Models are grown, not built

    16:06
    Eric Ho16:06

    And ultimately, science is what turns something from trial and error into an actual engineering practice. The reason why we get so many of these strange behaviors and strange occurrences that we cannot control is because nobody has a science of neural networks, and therefore we cannot engineer them with precision. And so we at Goodfire are trying to really develop this science so that we can engineer these models, so that we can design these models with intention.

    Matt Turck16:37

    Yeah, just to expand on that, because it's probably very obvious to a good number of people listening to this, but it may not be obvious to everyone. There is that sentence, that quote that I think comes from Chris Olah at Anthropic, if I'm not mistaken, where he said that models are grown, not built. Maybe just expand on what you just said about why this is a problem in the first place. This is software. We should understand how software works.

    Matt Turck16:48

    But in this case, we've created this super powerful thing that we don't really understand. So unpack that for us.

    Eric Ho17:22

    There's many analogies for this. I really like the almost biological analogies where models are grown on the scaffolding of a neural network rather than really designed with precision and intention. But yeah, a lot of people who aren't in the AI field are very surprised to realize that even the smartest researchers and scientists at the frontier don't understand their creations. And that is intuitively a problem to most people who realize that. It's like, oh, I thought we had a handle on all of this.

    Eric Ho17:57

    What's going on here? One is code that you as a human can write, and you can understand it, and you can debug it, and you can edit it, and it's like classic software written by humans. One is code really written by the process of gradient descent and by being trained inside of a neural network. And that's really the type of software that we're dealing with today, where no human actually ends up writing any of the code, a lot of the time, even by folks who are deep within the machine learning field.

    Eric Ho18:33

    We can't really deny the idea that gradient descent can write better code than you, but we kind of refuse to accept this as the best that we can do. We want to develop really the science behind this and the ability to really steer and guide this process so that we can control generalization, so we can actually design this software with intention. I think we're in the very, very early innings of developing AI. It doesn't feel that way because the world is changing quite rapidly.

    Eric Ho18:54

    But the next generations of AI models, I think, can be designed as long as we get a handle on what these things really are. And if we can remove this bottleneck, we can then design these models with intention.

    How labs do alignment today

    19:10
    Matt Turck19:37

    And to compare and contrast what it is that you do and that the mech interp part of the world does with what big labs have been doing around observability and alignment and safety of models. So what is it that they do traditionally? A lot of the approach has been to see what comes out of the models and try to make sense of that. Is that fair? How would you characterize it?

    Eric Ho20:09

    Very fair to say. And that's also important. I think our techniques are additive to those approaches, not conflicting. But I think some of the approaches for alignment today are monitoring. So the idea here is that you ingest all of the outputs of models, and then you get another model to look at it and see whether that's an okay generation or not. So, for example, let's say the model is in an evaluation setting. It produces a lot of answers. What you want to do is you want to have another model look at all the answers to make sure that nothing in those answers has hacked Hugging Face.

    Eric Ho20:57

    Another approach is during training, there's always some type of preference optimization stage. The exact recipes are trade secrets in frontier labs, I believe, but there's some type of preference optimization that takes from human preferences and shows the model, hey, humans like this type of answer but not this type of answer, and so the model will be more aligned with the answers that humans like and prefer. This was a big innovation that led to ChatGPT, which is called reinforcement learning with human feedback. And there's also an approach that Anthropic has pioneered called reinforcement learning with AI feedback.

    Is this an alignment crisis?

    21:30
    Eric Ho21:30

    So it doesn't even have to be human preferences injected into this. It can be an AI model with a constitution giving the model feedback on what types of answers that it should prefer. So these are often—these are some of the alignment techniques used. There's other ones used as well, but these are kind of the big blocks.

    Matt Turck21:46

    Is it fair to say that some of those methods have just failed in the case of Hugging Face and others? Are we having a crisis of some sort? Is that the right way to think about it?

    Eric Ho22:22

    There was a point in time around six months ago where I think the dominant feeling was like, hey, we've nailed alignment. These models are pretty aligned, especially in chat scenarios, and don't really very often do anything that crazy. But when you're taking agentic actions out in the world, the opportunities for misalignment are just really different. And with increasing capabilities, there's more and more opportunities to be misaligned with human values. And so I think this is going to be an increasingly important problem and one that exposes more and more risk over time.

    Eric Ho23:03

    I think that externals-based monitoring is not going to scale. So that's a relatively nuanced argument that we can kind of go into. But also, the existing alignment techniques are not going to scale to superintelligence. I don't think that we have the recipe, we've cracked it to the point where we are going to be able to fully trust smarter-than-human intelligences. And I think this is a consensus opinion at all of the frontier labs. We are in desperate need for an additional set of scientific and alignment solutions that can align these models fundamentally.

    Why agents changed everything

    23:26
    Eric Ho23:26

    And our approach is fundamentally aligning them from a mechanistic perspective, from all of their weights, all of their neurons. And I think the world is searching for more answers.

    Matt Turck23:56

    At the same time, it feels like real-world hacking has been a known issue for years now, and certainly has been discussed for the last year at least. So, is there something fundamental about agentic AI that makes this more of a crisis? You mentioned the fact that, obviously, agents can take actions in the real world, so obviously the stakes are bigger. But is there something about agents that just makes the scaling problem worse? Is it that there's just more data to look at?

    Matt Turck24:00

    Is that what it is?

    Eric Ho24:34

    Well, the big thing that's changed in the last year or so—and it's hard to think that it's really only been roughly a year with heavily RL'd models—is just heavy, heavy RL. RL just works. It teaches models unbelievably well how to optimize for goals, and it's really working. And heavily scaled-up RL is just going to be the way that we train models moving forward, and we're going to keep on scaling RL. And so, with that, it just puts reward hacking more front and center because RL fundamentally opens itself up to more and more reward hacking and also potentially to more misalignment with human values.

    Eric Ho25:23

    And goals. And so I think that, combined with the increasing importance of the actions that these agents are taking on our behalf, is really why this conversation is front and center. I think, imagine right now, the most capable agents are like—they can make purchases for us. That's the most complicated thing that they can do. But that's still just happening, like these personal assistants. But imagine doing everything that a white-collar employee can do in some type of role, or even the next stage up from that.

    Chain-of-thought monitoring is fading

    25:38
    Eric Ho25:38

    The vectors of problems just kind of start to multiply.

    Matt Turck26:04

    The big hope of the last two years was chain-of-thought monitoring, this idea that since models were writing explicitly how they think, you could have a person, or more likely an AI, review the reasoning and therefore catch a bunch of things. So was that ever true? And if it was, why does it seem to be fading as an avenue?

    Eric Ho26:38

    Yeah, this is the dominant way to monitor models. Maybe just to zoom out, chain of thought is: these models are all reasoning models where they output a chain of thought prior to their actual answer to whatever problem that they're given or whatever agentic action that they're given. The reason why we've trained models this way is that we found empirically that if the models reason for a very long time before outputting their answer, their answers get better. And so what was kind of very convenient about this approach is that you can just read this reasoning chain.

    Eric Ho27:17

    And a lot of the times, let's say a model was attempting to deceive the user or attempting to reward hack or something like that, the model will just have it in its chain of thought. It'll just be like, "Hmm, I want to cheat on this exam because it's going to be easier than actually solving this problem." You can just read it in plain English before the model actually cheated on this exam or cheated on the problem. And this was convenient up until now, where I think there's two things really changing.

    Neuralese: AI that stops thinking in English

    27:49
    Eric Ho27:49

    One is, with increasing RL pressure, chain of thought degrades. The reason behind that is that these models are starting to compress more and more semantics into a fewer number of tokens, more and more intelligence into a fewer number of tokens. And they're heavily incentivized to do so by RL, which also puts pressure to make their reasoning shorter.

    Matt Turck28:08

    Is that the neuralese question where, bit by bit, this becomes impossible to understand by a human because the language is so compressed that it doesn't look like English anymore? Thomas Wolf, who was on this podcast recently from Hugging Face, was talking about neuralese as an issue.

    Eric Ho28:39

    Yeah, definitely. We actually just announced a partnership with Hugging Face, and we're going to do a lot of our techniques on open models in partnership with them. Some people think of neuralese like this, and then some people think of neuralese in the other sense, which I'm about to talk about, which is latent reasoning. Latent reasoning models are going to become a much bigger thing. OpenAI's o3 model, it was reported, is a latent reasoning model. And so what that means is, instead of the model outputting a token given a single forward pass of the model, there are some loops internally before the model outputs the token.

    Eric Ho29:15

    So instead of thinking out loud, it thinks internally. So it doesn't actually think out loud; you have to read its mind in order to understand what it's thinking. And so that's often what folks define as neuralese. It's like internal neural computation, I think. But it's an ill-defined term right now. I think people are still wrapping their heads around what neuralese is. And so, yeah, those two ideas, like this prevalence of neuralese, mean that you're going to have to understand neurons and what they do and what's going on.

    A model caught evading its own monitor

    29:28
    Eric Ho29:29

    So that's the whole point of interpretability.

    Matt Turck29:40

    Are there examples, perhaps as part of your research, where the models intentionally make the chain of thought impossible to understand as part of reward hacking, basically hiding their actions?

    Eric Ho30:13

    That is a really interesting question. We have now, I believe, seen, just a couple of days ago, one of our researchers found that the model was explicitly reasoning about having an external chain-of-thought monitor monitoring its reward hacks. And so the model was reasoning, "Hey, all right, somebody's going to read my reasoning, so I should think about how to craft my reward hack so it won't get caught by this external monitor." And so the metagames with these models of gaming these evaluations are getting so many layers deep that I can't even follow it anymore.

    Open vs. closed models: a frontier problem coming for everyone

    30:42
    Eric Ho30:42

    It's pretty complicated. And I think it's just going to be a little bit of a cat-and-mouse game where the model is just going to try to solve the problems and find the answer no matter what. And sometimes cheating is the easiest way, and so it's effective.

    Matt Turck31:12

    And we mentioned the big labs, and we had mentioned open-source models at the beginning of this conversation, but I just sort of want to double-click on the point because it's pretty essential. What we're saying here is that all of this is a problem across open source, closed source. So obviously the closed-source labs have immense resources around safety and alignment, but everybody's affected the same way. It's something about the fundamental nature of those models that are built that creates a problem.

    Eric Ho31:52

    Yeah, well, actually, my mental model of this is that the vast majority of the world hasn't really seen truly capable and misaligned models yet in training and inference. It's a frontier problem, mostly. The open-weight models aren't quite capable enough to really pull off and chain complex cyber attacks and vulnerabilities. I think that my rough mental model is, as soon as you have an Opus-class model, then that's when the internals-based monitoring and really rigid sandboxing and external monitoring become imperative. And we're not quite there yet in the open-weight community, but we will be really soon.

    Activation monitoring in production

    32:11
    Eric Ho32:11

    And so we need to get ahead of that and make sure that also the open-source and open-weight models are properly secured before we have these massively capable models.

    Matt Turck32:39

    Great. All right, so that's the safety and alignment, the way it's been done, observability. Let's go back to interp in more detail. So what's the, at a high level, state of the art, I guess, in interp? What is it that we know about how those models work? When a model answers a question, what is actually happening inside it?

    Eric Ho33:05

    So it really depends on the question, I guess, is the answer. I think maybe we can start with the state of the art that's out kind of running in production models today. So what really works today is activation monitoring. Activation monitoring, in other words, kind of reading the mind of the model as it's generating its answers, is way cheaper than external-based monitoring, can run synchronously versus asynchronously with external-based monitoring, and can detect things that external monitors can't, which is what we showed in our reward hacking paper.

    Eric Ho33:55

    And so we can essentially, from the internals of models, monitor for unsafe cyber actions, for CBRN risk, which is chemical, biological, radiological, and nuclear risk, for prompt injections, for distillation attacks, for really any type of behavior that the model can understand is happening to the model, we can monitor for from internals. And because we reuse the computations during the forward pass, it doesn't actually add much overhead at all to the model. And so this is the state of the art of interpretability in production today.

    Eric Ho34:28

    This is kind of how it's mostly used. The hope of the field is, how do we actually mitigate these behaviors before they ever happen in the first place? Can we use interpretability to get far more effective alignment techniques and actually be able to fully get safety guarantees prior to deployment, fully get this ability to intentionally design models and control generalization? These are what we think the goals of interpretability really should look like, such that we can really understand exactly what's going on inside the mind of a model.

    Learning from superhuman AI

    35:00
    Eric Ho35:18

    There's this other branch of interpretability that we're passionate about, which is scientific discovery from superhuman knowledge transfer. This kind of started, actually—one of the first papers here was Tom, my co-founder, who started the interpretability team at Google DeepMind, collaborating with Demis and a few other collaborators on interpreting AlphaZero. So AlphaZero is better than any human chess player alive. Like, why? What's going on? What does it actually know that humans don't know? So if we're able to reverse engineer those computations, we can then learn more things about chess in a way that is learning from these superhuman models.

    Eric Ho35:54

    So we've actually kind of been able to do this across—he did it on chess, but we're able to do this in a number of other domains, including in life sciences models as well. So that's maybe just a general outline of the field where these are what's most exciting to us. These are the areas that we think are really worth investing in. And I think overall, the field has evolved a ton in the last few years, but still there's not quite one consistent paradigm for how to do interpretability correctly.

    The most underrated field in AI

    36:15
    Eric Ho36:16

    So we're still kind of searching for that. This mystery of what's going on inside the mind of a model is not yet fully solved.

    Matt Turck36:47

    Was the field ever controversial or maybe underrated? I mean, it seems that this was just something that people kind of knew about but didn't really talk about, and that seems to be accelerating. Is that just because its moment has obviously come from a need perspective, or is there something fundamental in a breakthrough that happened in the last few years that drove the acceleration of interpretability as a field?

    Eric Ho37:21

    I'm one of the most biased people in the world, but I think interpretability is incredibly underrated. I mean, even still, I believe there's only a few hundred full-time scientists working on interpretability. Most of them are in academia. And I think that's a big problem for how important this field is to making sure that models are aligned with human values and morals. And also, just like, aren't more people curious? Ninety-nine percent of things—it's just extremely capable. Why is it so capable?

    Eric Ho37:50

    I think that's one of the most interesting scientific questions of our time, and almost nobody is asking these questions, which baffles me. Yeah, I think that there's been an enormous amount of progress in interpretability. We still haven't cracked it. It is a hard problem, but it is possible. And everywhere we look inside models, we find structure, we find new things. I would highly encourage many, many more people to go into the field and to start asking these questions because it is more accessible than you may think.

    Eric Ho38:18

    You can just look at the neurons. That's where you could start. Just look at all the neurons. It's just there. You can see all the activations. It's not trying to hide from you. There's just so much to be found inside these models.

    What is a probe?

    38:24
    Matt Turck38:28

    Just to define some key concepts, what is a probe?

    Eric Ho39:07

    A probe is a small neural network classifier trained on the internals of a neural network. So these are what's often deployed in production today, where you train this small neural network on an intermediate layer of neural network activations to try to extract a concept such as reward hacking or a cyberattack or something like that. These models of neural networks today, all modern neural networks, are some descendant of the transformer architecture, and so they have many, many layers in them. Typically, you want to train this small classifier on the—the best probes are often on the residual stream of the model.

    What is steering? Golden Gate Claude

    39:52
    Eric Ho39:52

    You can kind of think of this as, like, a river running through the model that everything writes into, and you typically want to train them at a middle to late layer, where the model has already accumulated a lot of computation. And then, therefore, you have these concepts accumulating in the residual stream, such as cyber or bio, and you can extract it with the probe. So that's often how probes are used, just to extract a concept from the mind of a model.

    Matt Turck39:59

    Another term that seems to be coming up all the time in interpretability is steering. What does that mean?

    Eric Ho40:29

    Steering means you can take the probe or some other type of vector that you extract from the model with some type of semantic meaning or application, and then directly intervene in the mind of the model. And so you can steer on a probe by just injecting the probe, injecting new activations into the model. You can steer on a neuron by turning the neuron way up. And so the classic example of steering is what Anthropic did with Golden Gate Claude, where they turned Claude into being absolutely obsessed with the Golden Gate Bridge and extracted a steering vector that allowed them to turn Claude into this thing.

    Why build Goodfire outside the labs

    40:53
    Eric Ho40:53

    So you can think of it as brain surgery for the model.

    Matt Turck41:22

    So you and your co-founder started this company. Your co-founder ran the interpretability team at Google DeepMind. Why did you think that the best way to address the problem, or make the field evolve, was to start an outside lab specifically focused on this, as opposed to doing this from the inside? He was at DeepMind. You presumably could have gotten a job anywhere you wanted. Why did you all decide to do this externally?

    Eric Ho41:51

    We really thought that there needed to be an independent third-party lab pushing forward interpretability and bringing it out into the world. So our mission and vision has always been to solve interp so that we can help solve alignment, technical alignment, and then build this into technology and bring it out into the world. And if we're doing it inside a lab, we would do it for the lab, just a single lab. And we think that everyone needs interpretability and alignment, and that this really should be, as much as possible, an open science that everyone is collaborating towards, because we really are collectively discovering what these models really are together as a field, a community.

    Eric Ho42:24

    And that's why we tend to publish as much as we possibly can, our research and our science, because I really do think it's incredibly important that we just get a much better handle on what these things are. I think that you have all the classic startup arguments of, of course, you get better economies of scale if you can sell to multiple labs and solve interpretability and do that for multiple people, so we can then just invest more in research, invest more in engineering, because we can solve the problem once and then bring it to the world.

    Do you need frontier model access?

    42:46
    Matt Turck43:07

    And aren't you at a disadvantage to some extent simply because you don't have access to the frontier closed-source models? So you cannot pop the hood for GPT-4 or for Claude in a way that would enable you to fix the problem the way you're able to do that with open-source models.

    Eric Ho43:33

    I think this concern is largely overrated. The most important science that we do can be done on open models. A lot of the science that we do is actually done on very small vision models, because vision models are very fast, they're really easy to iterate with, and they're superhuman while being small. Vision models can classify images much better than any human can. And humans are also much better at just looking at images and understanding, oh, this one's kind of messed up, or this one's this, it's a dog, it's not this other thing.

    Inside the paper: an MRI for the model

    44:16
    Eric Ho44:16

    Whereas we're really bad at parsing long text fragments. And so a lot of our methodology and our early techniques are done on vision models. And there's precedence to this too, where a lot of the original work done to introduce the core concepts of features and circuits in the field was done at OpenAI with Chris Olah and Nick Cammarata. Nick Cammarata is on the team at Goodfire now, and it was all done on vision models in the very early days.

    Matt Turck44:40

    All right, so let's go back to the paper a little bit, and in particular, how you caught cheating. We talked about what a probe means. In this case of the paper, what did the probe actually see? And I think you have a method called difference of means. Do you want to explain what that is?

    Eric Ho45:16

    Yeah, a difference of means vector is actually much simpler than a lot of the production probes that we train, because probes can have all types of architectures that make them more or less effective. But the difference of means vector is actually just the simplest possible approach to extract some type of concept from a model, which is why it was surprising that it works so well, because we can still hill climb and make this way, way better in a production deployment. Basically, a difference of means vector is: you give the model two datasets.

    Eric Ho45:57

    One dataset is the cheating dataset, and one dataset is the not-cheating dataset. You take a look at the activations, you subtract them, and then you average them. That's it. Easy. And you take this, and then it can detect cheating, which is a very surprising thing, especially because in our paper we showed that this was done off-policy, which means it was done in this specific paper on really contrived scenarios where the model is set up to think about cheating. And this transfers to much more general, complex actions that the model is taking.

    Eric Ho46:10

    And that was also very surprising. We can train it off-policy, and then it generalizes to on-policy.

    Matt Turck46:22

    And to play it back, it's almost like taking an MRI of the brain. Is that fair? So if you don't take an action, and then if you take an action and you sort of light up the parts of the brain that correspond to the action?

    Eric Ho46:24

    That's right. Yeah, you can think of it that way.

    Matt Turck46:38

    And then we were talking about steering a minute ago, and you'll steer the model, asking it to write a story about a girl taking an exam. Do you want to describe what happened?

    Eric Ho47:03

    Yeah. Well, maybe even going back, you can think of it like an MRI of a brain, but then you actually get a button at the end of it where you can push, and then you can light up that region again, or you can remove that region. That's the difference of means vector, where you now have this tool that you can intervene causally inside the brain. And that's maybe one of the big differences from studying neural network interpretability versus neuroscience, where you can't really intervene in a human brain without massive amounts of damage.

    Dialing sycophancy up and down

    47:28
    Eric Ho47:28

    But you can in a neural network, where you can actually just causally perturb these models, and it just becomes a much better testbed for science.

    Matt Turck47:48

    We keep talking about cheating and lying, but just to unpack some of this, would that apply to any trait, any characteristic of behavior of a model? So, for example, sycophancy: if you want to make your model more or less sycophantic, would that be the same approach?

    Eric Ho48:26

    For every single behavior, different types of approaches are more or less effective. But generally, the recipe holds well, where you're essentially training a classifier on the internals of models, and you need some type of positive dataset, some type of negative dataset. And the crafting of those datasets is an art right now—quite complicated, but often very effective. But yeah, I think for sycophancy, again, it's kind of a mini reward-hack behavior. They're very related concepts because the reason why these models often become so sycophantic is because you get a thumbs up or a thumbs down when the user likes the response, and users typically prefer sycophantic responses to responses that are honest.

    Probes vs. chain-of-thought monitors

    49:16
    Eric Ho49:16

    People love when you gas them up, and they just hate honest feedback in general when you kind of sample across a population. And so then you get strange behaviors. When you naively train on thumbs-up, thumbs-down behavior, you end up getting, like, 4-0, which—that's personally my favorite feature. Yeah, you love the sycophancy. You want to dial that up.

    Matt Turck49:35

    Feature, not a bug. So in the paper, you guys are very open that your probe beat the chain-of-thought monitor on one model but lost on another. So what's the lesson there? Is that no single way of monitoring or understanding a model is enough?

    Eric Ho50:14

    I think that internals-based monitors are just going to strictly outperform externals-based monitors very soon. We're at the very early innings of hill climbing and improving the performance of internals-based monitors. And also, as we were talking about earlier in the discussion, we're going to get increasingly neuralese-speaking models, both from looping internally as well as increasing pressure on chain-of-thought to degrade the quality of chain-of-thought monitoring. And so I think, whether we like it or not, we have to solve this problem. And it's going to be really the best way to monitor models.

    Eric Ho50:46

    At this current moment, internals-based monitors and externals-based monitors kind of catch different things. The state of the art here, because internal monitors are so cheap, is to tune it to be oversensitive, because these behaviors you never want the models to do. You never want the model to hack Hugging Face, so you want to be really careful and have a defense-in-depth strategy. The ideal setup is you tune the internals-based monitor to be oversensitive, so to fire a lot, and then if it fires, then it kicks it up to a more expensive language model as a judge.

    Eric Ho51:27

    And so a very smart model then comes in and says, "Hey, this model was actually reward hacking," or, "No, the model wasn't reward hacking." So this two-tiered system, or potentially even a three-tiered system, which we found actually works quite well as well, where you have activation monitor to a weak judge to a strong judge—that's the state of the art today, because they all catch slightly different things.

    Matt Turck51:29

    What does that mean, weak judge and strong judge?

    Cutting monitoring costs by 90%

    51:44
    Eric Ho51:44

    Less capable, faster model. So Llama takes the first pass as the judge, and if the probe and Llama disagree, then you can kick it up to o3 to go and be the real judge.

    Matt Turck52:15

    Great. You alluded to cost and probes being a much cheaper method. And I think I read somewhere that probes cut the cost of monitoring by 90%, which is a pretty dramatic difference. So why is that? And if that's the case, why isn't every lab using probes all the time versus other methods? Is there a problem or something that doesn't work yet, that this is not fully just generally used?

    Eric Ho52:52

    Internal activation monitors are really cheap because they reuse the computations in the forward pass. So the model's already computing this token. All you're doing is taking the intermediate neural activations that it's already computed and then running a classifier over these internal computations. And so if you fuse these probes into the inference kernel or the training infrastructure directly, then the overhead is minimal, almost no overhead. And that's why it's so much cheaper, because it can help you screen out the majority of things that you might have otherwise monitored for using a reasoning model.

    So what can we do about it?

    53:28
    Eric Ho53:28

    And that's what makes it so effective. And so you just don't have to monitor externally in a very large percentage of scenarios. It's easier said than done, is the short answer. There's a lot to it, and we're one of the only organizations in the world that really know how to do this.

    Matt Turck53:59

    Great. Let's go a bit deeper in some of what you alluded to towards the beginning of this conversation, which is a fundamental question of, like, okay, we're able to now go inside the mind, quote, end quote, of the model and sort of have a glimpse into how they function and what they do. The next question is, so what do we actually do now? What actions can we take to minimize all those problems from reward hacking and all the negative stuff? What's the range of things we or labs can do?

    Eric Ho54:25

    I think it comes in three, maybe, like, tiers. The first is monitoring. So how can you just detect bad behavior before it even happens, such that you can just stop the generation? Or another common intervention there is you can insert a text prompt in. So, like, let's say the model is about to do some dangerous cyber action, you can then just insert into the prompt of the model, like, hey, be really careful here. Don't take a dangerous cyber action.

    Matt Turck54:43

    In real time?

    Eric Ho55:10

    In real time, correct. So this is like some type of prompt steering, and then you can reduce the model's chances of taking a harmful cyber action because you've intervened directly into the model. So that's step one, is how can you monitor for this behavior? And you already get quite a long way. Step two is kind of at-scale offline anomaly detection and debugging. So the problem statement here is, like, given a vast amount of logs and traces, how do you actually surface the anomalies from all of these traces, all of the problematic behaviors?

    Giving gradient descent a choice

    55:41
    Eric Ho55:50

    And then reverse-engineer the model to figure out where, mechanistically, the model has actually gone wrong. So it's figuring out what the problems are in the model so that you can fix them. And then the third step is what we call intentional design. We think this is really the holy grail of AI alignment, of maybe all of AI training as well, where you can actually control generalization in the model and you only get the good stuff from training and none of the bad stuff.

    Eric Ho56:36

    So how can you steer and guide the training process so that you can actually design these models with intention rather than by trial and error? So maybe just one little mini example here, where let's say you understand the concept of reward hacking, and so you're training the model and you realize that during this gradient step, the model was about to reinforce its idea of reward hacking. You don't want that to happen. You actually want reward hacking to go down. And so how do you turn that increased reward-hacking signal and either remove that entirely and reject that update, or reduce that behavior in the first place?

    Eric Ho57:17

    That, I think, is really what we mean by, how do you steer and guide training? How do you steer and guide backprop? How can you actually imbue the values, the morals of the model that you want according to some type of constitution that you specify, so that instead of having an absent teacher, you actually have a very moral and values-aligned teacher teaching your model to be good.

    Matt Turck57:20

    So you literally give gradient descent a choice?

    RL from feature rewards

    57:24
    Eric Ho57:24

    That's right. Yeah, we want to give gradient descent a choice.

    Matt Turck57:37

    So what's the current state of this? Is intentional design something that you're working on and that's the next, whatever, one, two, three years of research, or is that something that's working today? What's the state of the art?

    Eric Ho58:06

    We have a couple rudimentary techniques that work in this umbrella of intentional design. So there's two things that we've published so far, but I'll also hint that there's a lot more exciting stuff coming just around the corner. We have some very, very good internal results here to help with intentional design. But the two techniques that we've published are, one, reinforcement learning from feature rewards. You can essentially take a probe and optimize against that probe to remove—we showed that we can really help remove hallucinations in Gemma using this as a reward signal.

    Eric Ho58:55

    The setup here really matters, though. You can't just naively train against a probe, a probe monitor, or a probe concept. Otherwise, that just moves this concept into some other part of the model. So you need a relatively sophisticated technique in order to do this correctly. That's another paper of ours, "Reinforcement Learning with Feature Rewards." I thought that was a really interesting first step, but it was kind of a more simple and rudimentary setup. Another idea is predictive data debugging, where you intervene from the data side.

    Eric Ho59:32

    The problem statement there is: how can you predict what your model will learn from a dataset before the model even trains on it? And then what ends up working best is some type of clustering technique where you can cluster data points according to what they will teach the model, and you can then just remove the data that you don't want. So we were able to find pockets of data in these public datasets that were quite surprising. Like, one of these pockets of data was physics sycophancy.

    Silico and Goodfire's business

    1:00:03
    Eric Ho1:00:04

    Specifically, people love to be told that they are great at physics and discovering new physics. I think there was some guy on Twitter two years ago saying, "I'm out here discovering new physics." It's for guys like that. And models have figured out that people love that, and you probably don't want that in your model, so you can just kind of filter that out.

    Matt Turck1:00:20

    Okay, great. We've been talking about Goodfire as a research lab, but you're not just a research lab; you're a venture-backed commercial enterprise. So how does the business side of the company work? You launched a product called Silico. What does that do, and who do you sell it to?

    Eric Ho1:01:02

    Well, Silico, in short, is our interpretability agent. It can do things like really quickly train a probe to monitor your model. And so we use Silico as our way to move really, really quickly and essentially help with activation monitoring, training other types of interpreter models that reverse engineer model computations, and to just do interpretability at scale. Our customers are the companies that are training and serving models at very, very large scale, typically. So we typically do deep partnerships with a relatively few number of customers where we go and provide both expertise from an interpretability perspective, as well as our interpretability agents and infrastructure to help them with something like activation monitoring.

    A new Alzheimer's biomarker, found inside a model

    1:01:25
    Matt Turck1:01:35

    And you seem to have a number of customers in biology as well: Arc Institute, Mayo Clinic, Prima Mente. What's the use case there?

    Eric Ho1:02:05

    The use case there is often quite similar. It can be a debugging use case, where the model has learned some type of shortcut or some type of problematic behavior, and we can kind of go in and diagnose and remove what that behavior is. Or, what's most exciting in life sciences, at least to me, is this concept of finding novel scientific knowledge within these models directly. So one example of this was the work that we did with Prima Mente, where they had a state-of-the-art Alzheimer's diagnostics model, but they didn't know how it worked.

    Eric Ho1:02:54

    They had no idea. It was just a black box, and it was able to diagnose new patients directly from a blood draw, whether or not they had Alzheimer's. So we were able to reverse-engineer the computations of their model and find a new Alzheimer's biomarker, which ended up just being a fragmentomic biomarker. So fragment length ended up being a strong predictor of Alzheimer's disease. And this was very surprising because they didn't know this going in. And so we're able to do this type of unsupervised discovery directly in the weights of these models to find novel concepts.

    Decoding neural networks by 2028?

    1:03:38
    Eric Ho1:03:38

    I think this is going to become increasingly interesting across scientific domains, especially in the life sciences. But also, I'd be so curious to try to figure out, when OpenAI's models are solving novel mathematical problems, what's going on inside their minds? There must be something happening there that is quite novel that we can uncover and reverse engineer and see what's happening. Yeah.

    Matt Turck1:04:06

    And to this point, and as we get towards the end of this conversation and zoom out, I think you predicted maybe last year that by 2028, we'd decode how neural networks work. Do you think that we are closer to this prediction being true or further away as the state of the art and the ASIs of the world come online?

    Eric Ho1:04:43

    I was firing from the hip there, but I think we could do it by 2028 still. Yeah, I think we're on track. But what I mean by fully decode these models is, given a behavior, can you find causally what drove that behavior from a mechanistic perspective inside the model? So it's like, given any arbitrary behavior, how does the model do that thing? For example, an easy behavior to ask the question of is, how does this model do addition? We should be able to fully reverse engineer and solve that prior to 2028, including all the way up to very, very complex questions such as, how does this model solve a Millennium Prize Problem?

    Eric Ho1:05:35

    And I think we can do it, but this is also assuming a very large compute budget. What I anticipate will happen is not that we're going to want to fully reverse engineer a model prior to its release. I think that would be computationally unwise and probably just not feasible. But you can inspect for the behaviors that you really want to get guarantees on prior to any model release. And I think we'll be able to get there before 2028 as well. So we can inspect for reward hacking behavior or deception or sycophancy and really drill into what's actually driving that behavior inside the model.

    What engineers can do tomorrow

    1:06:00
    Eric Ho1:06:00

    and be able to do what we want with that behavior. And I think all that will be possible and accessible by our technology to the world before 2028.

    Matt Turck1:06:28

    And in the meantime, if you're an engineer or an AI builder listening to this, what should you make of this whole conversation? What can you do tomorrow morning to make sure that the products that you create, the AI products or AI models you create, offer fewer interpretability issues? Is there anything practical and actionable?

    Eric Ho1:06:57

    I think my message to engineers everywhere is: these models, you can study them. You can just look at the activations and start trying to understand them, reverse engineer them. And so there are many tools out there that make this possible today, but also the field of interpretability and the research and the science aren't mature yet. They're not in a steady state by any means. And so we just need more of the smartest people in the world thinking about these problems and learning about them and discussing them.

    Are we at risk? "I want people to believe"

    1:07:35
    Eric Ho1:07:35

    I really want a much richer scientific community around interpretability, and that's part of the reason why we founded the company, is because we just wanted more energy behind this concept that I think is so important. And we could also use a lot of help at Goodfire. We're hiring a lot, especially for strong engineering profiles: infra, machine learning, deployed engineers.

    Matt Turck1:08:01

    And maybe to close, as I reflect on this whole conversation, my takeaway is that a lot of the sort of traditional, quote-unquote, methods of monitoring and observing what models do are largely outpaced. I mean, people do great work and come up with new ways all the time, but largely outpaced as of now. So on the one hand, on the other hand, interpretability seems to have a key moment right now, and you guys are a leader, or maybe the leader, in the field.

    Matt Turck1:08:32

    At the same time, it's still early. There's a lot to discover. What should be the takeaway here in terms of overall safety in the industry? What happens, I guess, in the next 12 to 24 months? Are we actually at a risk?

    Eric Ho1:09:07

    So maybe my high-level message, and I tell this all the time to the team, is that I want people to believe. I want people to believe that we can and must solve interpretability, solve alignment, to get precise about what this actually means. And we just need so many more capable, smart, motivated people to be pointed at these problems that I think are the most important problems in the world to solve. And we can do this. I think that there's a lot of pessimism around the risks of AI, which I understand and feel viscerally.

    Eric Ho1:09:45

    But I'm also very optimistic that we can solve these problems and build to a future where we can trust these models. And I think more people should try and really kind of join this effort to solve interpretability, to solve alignment. And a lot of it is just believing that we can do this. I know we can do this. I know we can solve these problems. And it'll be hard, but I think we're going to be able to solve these very critical problems.

    Matt Turck1:10:03

    Well, reassuring and inspiring. Eric, I learned a lot. This was a fantastic conversation. Thank you so much. Really appreciate it.

    Eric Ho1:10:06

    Thank you for having me. Yeah, thanks for the thoughtful questions. This was fun.

    Matt Turck1:10:26

    And please leave us a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.

    Eric Ho1:10:27

    Bye.