MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    AI Could Take Over in 2029. Is It Already Too Late? | Ryan Greenblatt

    Ryan Greenblatt is the Chief Scientist at Redwood Research. We cover his recommendation to plan for full AI R&D automation by early 2029, why alignment-faking experiments found Opus 3 behaving differently in training and deployment, and his Plan A proposal for US-China compute transparency that pauses most AI R&D.

    08/27/2026

    Hosted by Matt Turck · with Ryan Greenblatt, Chief Scientist, Redwood Research

    AI safetyAI alignmentAI takeovercompute governanceAI timelines
    Listen now
    YouTubeApple PodcastsSpotify
    1h 19m · 23 chapters
    Contents

    Transcript

    The AI CEOs are aware of the risks, but "proceeding anyway"

    1:24
    Matt Turck1:28

    I wanted to start with a strong statement that you guys have at the beginning of AI 2040, where you say that the CEOs of OpenAI, Anthropic, xAI, and Google DeepMind understand that the current path AI is on leads to extinction or a power grab, and yet they are proceeding anyway.

    Matt Turck1:50

    So I was curious if you could riff on that and give us more color on what you mean.

    Ryan Greenblatt2:09

    I think that the AI company CEOs understand that they're on the path of building wildly smarter-than-human systems, superintelligent AI systems. They understand that we don't really have a clear, thought-through plan for how to manage the risks from that, how to ensure these AIs don't take over, how to ensure they remain under our control. And they understand that there's a decent chance of an unprecedented concentration of power, at least in the absence of very active efforts by the people who end up with that power to redistribute it.

    Ryan Greenblatt2:35

    I think that we say some sort of more precise thing in the "Why Did We Write This?" section. I think that overall, my sense is that the company CEOs do legitimately think that the thing they're doing is very risky, or at least the thing that these companies are doing. I think that they have a mix of views for why they're doing this, where some of it is that they think they're better than the next guy or think it's good if there's multiple companies or multiple things.

    Ryan Greenblatt3:11

    I think that they vary a bit on how much concentration of power they think AI will yield, or maybe just haven't really thought this through very carefully. I think that there's more public statements from AI company CEOs on there being huge risks than on specifically the risk that AI ends up with the power in the hands of the few rather than being as distributed as it is now, which is not arbitrarily distributed now. Overall, this seems pretty consistent with what they've said publicly, but I think that they're imagining a world where maybe AI concentrates power massively, but the people with the power end up deciding to redistribute it.

    Ryan Greenblatt3:26

    But they could have totally taken over the world.

    Astra paused, and the letter signed by 1,200 insiders

    3:27
    Matt Turck3:58

    Yes. And since you published AI 2040, there's a couple of things that happened that seem to be going in the general direction of what you recommend. Specifically, OpenAI paused its Astra model over safety following the Hugging Face incident. And then 1,200 insiders published that letter a few weeks ago, including Dario, asking the government to slow down tools. I'm curious what you make of it. Is that what you are recommending that's starting to happen?

    Ryan Greenblatt4:17

    Yeah, I think these are good steps. I think that specifically, the Pacing the Frontier letter, the idea is we just need to have the tools to, if we're in a position where we need to spend a bunch of effort on safety, which I think we may be in today, we're very likely to be in it in the future when AIs are more capable, we need to have the ability to do that while also not making it so that the actors that are actually applying safety get overtaken by other actors.

    Ryan Greenblatt4:39

    There's multiple different angles here, but the most basic is just if we're in a position where U.S. companies are all really scared, they don't think they can proceed without spending way more resources on safety, and there's a coordination problem. It would be very nice if that coordination problem was solved rather than every company thinking it would be better if they all went slower, spent more effort on safety, but then raced to oblivion instead because they think they're better than the next guy or whatever other reason.

    Ryan Greenblatt5:20

    And then I think another aspect of this is specifically doing this in a way where part of the story is either slowing down China or cutting a deal with China such that China doesn't overtake and break this whole proposal. Yeah, and I think these are good steps. It's harder for me to say as much about what's going on with OpenAI pausing training and not deploying Astra and what exactly the motives for that are, what's exactly going on there. But I think that overall, these seem like good steps.

    Ryan Greenblatt5:40

    I think my sense is that the employees at these companies are pretty freaked out about how things are going and don't think that we're necessarily on track to handle all these problems in time, given how fast recent progress has been. And I think they recognize that, which is why they signed the open letter.

    "Not bad. Dangerous." What superintelligence actually threatens

    5:45
    Matt Turck6:04

    And just to verbalize the question early in this conversation, what is so bad about superintelligence? Obviously, there's a lot of talk about scientific progress and curing cancer, and your document effectively recommends pausing the race towards superintelligence. So why is it so bad?

    Ryan Greenblatt6:11

    Yeah, I wouldn't say superintelligence is bad. I would say it's dangerous. It's a very dangerous thing to create. The most straightforward story for why it's dangerous is that it seems pretty likely that, on a trajectory similar to the trajectory we seem to be finding ourselves on, you end up with AI takeover as a result of building superintelligence because the AIs are in a position where they can take over due to being highly capable, widely deployed, and building basically a huge amount of industrial capacity, potentially.

    Ryan Greenblatt6:46

    Then we could talk more about what a takeover would look like. And then if they're in the position where they could take over, then there's a question of, would they want to, or how would the motives shake out? And it looks like we don't really have that much control over the motivations of AIs, and it seems like that problem gets harder as they're much more capable and built via a process where AIs are automating AI R&D and we maybe are losing our understanding of how that process works.

    Ryan Greenblatt7:27

    There's sort of this we-don't-necessarily-control-the-technology, misaligned-AI-takeover concern. Another concern is that historically, at least in recent times, the distribution of power among humans has been reasonably distributed, though not necessarily super distributed, due to people being able to work for money, labor. The reason why many different people have some power and are cut into the world is to a substantial degree because we have the ability to work. I think that AI means that basically all of the key financial assets might just be entirely capital.

    Ryan Greenblatt7:51

    Human labor would have very little value left. I think it's unclear exactly how that works because it depends on people having intrinsic preferences to employ specifically humans, even if an AI could do their job just totally way better. That means both that there might be a natural effect where power gets very concentrated. In the same way that it's easier to be a brutal dictatorship if your money comes from oil instead of from a productive, broadly distributed economy, it might be much easier for someone to consolidate power a lot if the economy is running on AI and machines rather than on humans.

    Ryan Greenblatt8:32

    In addition to that, there's some concern, which is that it might make coups way easier because right now, to do a coup in many countries, you need the support of a broad base of people, and it's possible to build checks and balances. Whereas if you end up in a system where basically AIs are running anything, if anyone either puts secret objectives into that AI or has overt control of those AIs, then they could just directly take over. And there's just a huge threat to democracies, the U.S., whatever, because you'd be in that position.

    Ryan Greenblatt8:53

    And then the third thing—there's misalignment, this concentration of power and coups and power grabs. And then the third thing I would say is just AI might yield extremely, extremely rapid technological progress, which is more rapid than our wisdom grows to match it. There might just be all kinds of things that happen as a result of very fast tech progress that we're not necessarily going to be able to handle on time, in part because the AIs might be better at doing things than they are at thinking carefully about how to do things.

    Ryan Greenblatt9:39

    They seem much better at accomplishing hard results and easy-to-verify results than they are at contextualizing things, understanding the broader picture, understanding what would be a good or bad choice in the broader context. And so you could worry that this causes problems. Examples might be dual-use, offense-dominant technologies, so bioweapons. You might just worry that we just come across some technologies that are very dangerous. There are some concerns around AIs being very superhuman at persuasion and this destabilizing society. All these seem like things that I think we could deal with, given time.

    Ryan Greenblatt9:51

    But if things go very fast and we don't get to necessarily pick the order in which these technologies develop, it seems like we might be in trouble.

    Recursive self-improvement, and the intuition objection

    9:55
    Matt Turck10:25

    Also central to the argument that leads to speed is what you mentioned about AI automating its own creation, which takes us into the territory of RSI, recursive self-improvement. So you were just on Dwarkesh and had a great conversation about what is RSI and how it can go wrong. So let's not redo this, but still, as a TL;DR, there were a couple of parts to that discussion that I found particularly interesting. In particular, there was this argument that RSI may be very good at automating the process of creating the next generation of AI, but it may or may not have the kind of intuition that one needs for scientific breakthroughs and truly novel ideas.

    Matt Turck10:48

    What's your rebuttal to that argument?

    Ryan Greenblatt11:17

    The way I would put this argument is that AIs might be very good at the more mechanistic or nitty-gritty parts of AI development, like writing code and running experiments, but not as good at the broader conceptual leaps. First, I would say that AIs seem significantly better at engineering and grungy stuff and just keeping trying than they seem to be at conceptual breakthroughs, but their ability to do these breakthroughs, especially in easy-to-verify domains, is improving. One example is their ability in math, but even in ML, their taste has been improving.

    Ryan Greenblatt11:47

    I think it continues to improve. So it's not so clear to me that this will lag super far behind. Another thing is that you can measure how good these AIs are at intuition or research taste or having breakthroughs, especially in domains that are relatively easier to verify. If you can measure it, then you can take your grungy AI labor or even just your human laborers and try to optimize that. If it can be measured, it can be hill-climbed on, very roughly speaking at least.

    Ryan Greenblatt12:13

    I think this is a case where you could just hill-climb on how good the AIs are at making these sorts of breakthroughs in a wide variety of different settings. Then I expect that would transfer to making the actual breakthroughs. So you could have a breakthrough bench or whatever. My sense is that the transfer doesn't look so bad, that AIs are able to spin up very quickly in some specific R&D domain and then get some intuition for what the best approaches are and iterate there.

    Ryan Greenblatt12:38

    Right now, the approaches they tend to focus on are ones that are relatively straightforward and are more just like grungy iteration, but they're increasingly better at doing the broader experiment design and discovery and also having good ideas.

    Matt Turck12:58

    You also make the argument that for purposes of a potential discontinuity, whether it transfers or whether it generalizes may not even really be a question, and that if it was super good just at the industrial part of accelerating AI, that would be enough. Is that fair?

    Ryan Greenblatt13:22

    Yeah, I would say that if AIs could just automate AI R&D and automate the industrial process of building more computers, then you could quickly end up in a process where robots are building robots and the whole world is greatly transformed. That very quickly can get you to a point where AIs could take over. The economy has been radically changed, even if it's hard to train AIs at some other tasks. But I do think that once the situation is like there are robots building robots that build computers and the full feedback loop is closed at that level, my sense is by that point the AIs will be good at basically all human work, or at least not too deep into that.

    Ryan Greenblatt14:00

    Maybe there's some period where robotics is a big deal, but there's a bunch of stuff they can't automate. But it seems to me that that's how it's going to go. But just in general, automating R&D seems like it's enough to radically transform the world. And if you look at why humanity is a big deal, why have we been able to accomplish so much collectively? A lot of it is because of having technology, having organizations, being able to organize ourselves in various ways and accomplish things in the world.

    SSI rumors: does continual learning change the picture?

    14:16
    Ryan Greenblatt14:19

    If AI just had this industrial capacity combined with the ability to develop more capable AI systems, that seems like it's quite far in and of itself. Quite extreme.

    Matt Turck14:50

    Just looking at Twitter today as we're recording this, which is August 24th, there's plenty of rumors about SSI coming out with their first model in the next couple of days, potentially, with what could be a breakthrough in continual learning. And I'm curious whether there's an overlap between RSI and continual learning, whether continual learning would feed into RSI, or is that orthogonal?

    Ryan Greenblatt15:11

    Yeah, I would say that continual learning and RSI strike me as mostly orthogonal, where by RSI I mean something like the process of AIs themselves sort of intrinsically accelerating AI development through a variety of mechanisms. That's a little bit ongoing now and would be much more striking if they fully automated R&D, though there's definitely some of that going on today. My sense is that continual learning, very efficient lifetime learning that is consolidated across many instances rather than being within a single context, or even just being much better at doing within-context learning over very long contexts or whatever, would be a capability that's very useful for all kinds of things, including automating AI development.

    Ryan Greenblatt15:56

    At least it would make the AIs better at that. I don't know. There's a question of whether we want to do that as a society. But I don't think it's very specific to that. I think it would also be helpful for just all other kinds of tasks you want to apply the AIs to. I do think that there are some types of continual learning schemes that are easier to do within a company than across the whole economy, and are probably easier for AI companies to do themselves than to do to other companies because of confidentiality issues and things like this.

    Ryan Greenblatt16:10

    But very broadly speaking, I don't think it's very specific. It's just a specific type of capability.

    Matt Turck16:22

    But it could feed the recursion, right? If the same model can keep learning about the world, then its ability to automate AI R&D may be greater without having to retrain a whole new model.

    Ryan Greenblatt16:46

    Yeah, so I think there's a feedback loop you could get, which is that AIs are really smart, therefore they're very fast at learning on the fly, and that learning gets reintegrated, which makes them even better at doing AI R&D, which means that maybe they'll be even faster at learning. But even putting that aside, maybe they'll just be better at AI R&D so they can make another AI which is even better at continual learning. I think there might be some continuum between—well, continuum is a bit of a sloppy word—but there might be some spectrum between natural human-like continual learning and more training on RL environments where you might end up with a situation where AIs are constantly iterating on their own training with new RL environments or new training data in a way that vaguely resembles human within-lifetime learning, but it's also different.

    Ryan Greenblatt17:24

    That could be part of the feedback loop. Even if you do have to do the retraining to get continual learning, there's various versions of retraining that you could just do every single day. There's nothing that in principle stops you from doing a small amount of retraining constantly.

    His timeline: "plan as though it happens in 2029"

    17:27
    Matt Turck17:33

    Okay, great. Very helpful. So what's your latest prediction on timing for RSI to happen?

    Ryan Greenblatt17:54

    Maybe my median for, let's just say, full automation of AI R&D, by which I mean basically even if humans left the picture, things wouldn't slow down by that much, would be maybe end of year 2030 or maybe early 2031. Not that much precision in these numbers, but something like that. Not that much stability. These numbers fluctuate some, but that'd be my guess for median. But then I think that's sort of the central scenario I plan for, which is maybe more like my 35th percentile would be end of year 2028, beginning of 2029, which I think is very likely.

    Ryan Greenblatt18:22

    I think that seems super plausible. And I think that if I just extrapolate out the current trajectory, it looks more like you get that trajectory. And the reason why I don't think that's my median is there's just a bunch of factors that might kick in to push things back. Maybe there's some key bottleneck I'm not seeing. Maybe there's something that I think will work to overcome some obstacle won't actually work, or maybe there's going to be significant government slowdowns because people freak the fuck out about this technology, which seems kind of plausible.

    Ryan Greenblatt19:01

    So because of that, I push later. But I think that in terms of what I would recommend people plan as though is happening, I think I would recommend planning as though full automation of AI R&D is maybe start of 2029, maybe earlier. And then also AI R&D being quite automated by 2028, possibly earlier, where humans are a much less important part of the picture by then. Well, I don't know if I should say probably. That's very central, I think, is humans are much less an important part of the picture early 2028.

    Ryan Greenblatt19:10

    And already it's the case that AI R&D is quite automated as it stands today.

    Is it already too late?

    19:11
    Matt Turck19:18

    Obviously, early 2028 is like tomorrow morning, effectively. Is there a scenario where all of this is already too late?

    Ryan Greenblatt19:35

    Yeah, I mean, it depends on what you mean by too late and what all of this means. I think I am worried that specifically AI 2027 Plan A is assuming too much effective government time. It's assuming the government has more time than it actually does to take all these actions. And I think it's pretty realistic that the actual plan we should go for should be, in practice, quite a bit sloppier and less well-organized and faster, just because we just don't actually have time to do something quite that elaborate.

    Ryan Greenblatt20:12

    I'm not confident in that. I think there's a bunch of different options, but in general, I would say it might be too late for some interventions. Various policy windows, if you follow the normal timeline, are closing. That said, I think that there's a long history, at least in the U.S., of, in times of crisis, things can happen much faster. And there's a lot of different levers for that. And so I think that if we got to a position where everyone is like, holy shit, we need to take this specific action, that could happen very quickly.

    Ryan Greenblatt20:41

    But we might not get to a position where there's that much consensus. And also, it might be that the government just isn't tracking AI or isn't aware of AI to a sufficient degree in the relevant timeframe because things go too fast. Adoption lags. I think people's understanding of what AI can do lags behind what's actually possible. My guess is that if you gave people a quiz of what AIs can and can't do, they would give answers—or if you gave congresspeople a quiz like this, they would give answers—that were more true a year and a half or two years ago than they're true today.

    Ryan Greenblatt21:22

    Possibly, they're even underestimating capabilities from then. I don't think it's the case that most people in D.C. would correctly answer that the largest mathematical breakthroughs over the last two months have been vastly from AI, which my understanding is, that's true, at least if you measure size by not necessarily how much insight there was, but in terms of just how important people would have said the type of result would be.

    Ryan's path: COVID, podcasts, Redwood

    21:23
    Matt Turck21:36

    We will go back to AI 2040 in a minute. But I wanted to do a quick segue about you and your story. What first pulled you into AI safety?

    Ryan Greenblatt21:59

    Yeah, so I was in my junior year in college, and I was sort of alone in my apartment because it was COVID, and I was listening to a lot of podcasts, and I was sort of thinking a bit about what I should do with my life. And I ended up, through some somewhat twisted path, thinking I should be way more interested in helping other people and being altruistic than I was at the time. I should be very focused on how can I make other people's lives as good as possible and make things go as well as possible.

    Ryan Greenblatt22:31

    Then from there, I considered a bunch of different routes and was looking into a bunch of different things and was researching different possibilities, and eventually decided that the best thing I could do with my career and my life was try to make AI go better and, in particular, avoid AI takeover, but also more generally try to make that go better. Then I applied to a bunch of places. I ended up working at Redwood. This was about five years ago at this point.

    Ryan Greenblatt22:48

    I have been working there since, and I've done a bunch of different work there. GPT-3 wasn't released yet. GPT-3 came out. The field is very different, and I think there's sort of a post-GPT-4 era, and then there's a more recent post-wide-adoption-of-coding-agents era. And then probably soon there's going to be additional eras, and things are going quite a bit faster, and the progression even of just model releases is so much crazier than it used to be.

    Matt Turck23:23

    Reading Redwood Research stuff over the years, it seems that there has been an evolution from being focused largely on interpretability to much more AI control. Is that fair?

    Ryan Greenblatt23:24

    Yeah, that's fair. I would say that our arc as an organization was, when I joined the organization, I had just finished up a project on adversarial training and was interested in getting into doing interpretability and what we would call model internals work, where it's like, can we take advantage of the fact that we have white-box access to these models to do something better than just the naive methods of prompting and training when trying to align these models, understand their motives, know what's going on?

    Ryan Greenblatt24:00

    We explored that area for a while and then, for a mix of reasons, decided it was quite a bit less promising than we had initially hoped and decided to move on to other things. One of the things we moved on to shortly after that was AI control, which is the idea that maybe it would be a good idea to prevent AIs from being capable of accomplishing problematic things, or basically make it so the AIs aren't able to cause huge problems even if the AIs wanted to.

    Ryan Greenblatt24:41

    There's a bunch of different stories for why this is a good idea, but basically the idea is that there may be some intermediate period, an intermediate period that I would say we're currently in, where the AIs are maybe capable enough to cause at least moderate problems and then increasingly able to cause quite large problems, but they're not necessarily so capable that if we tried quite hard to put in various safety measures, those AIs would be able to subvert them. We can put in monitoring, we can put in various security controls, we can have better understanding of what the agents did, we can have better pipelines for reviewing what actions they took and auditing and overseeing them to a point where even if the AIs were really misaligned, it would just be hard for them to get away with doing anything super bad.

    Ryan Greenblatt25:23

    I don't think we're there yet. I don't think that the situation is currently looking super impressive for AI control in terms of what companies have done. But I think that there is a research field that we've been working on that seems like it is very promising and could be done, though it would take a lot of effort. It would pose some costs, but I think we're seeing increasing interest in this. For example, OpenAI recently said that they were going to be monitoring a larger fraction of their internal traffic.

    Ryan Greenblatt25:45

    It seems like, based on their blog post, a serious cost in terms of compute. It's a little hard for me to know exactly how expensive it actually is. We don't have enough info to get a great sense, but that's some indication that at least companies are willing to spend a lot of compute in this area. It seems like there's a lot you could do with that in terms of making it so that we're more likely to both detect misalignment and prevent misalignment from causing big problems.

    Ryan Greenblatt26:17

    Then, in addition to working on AI control, we also just work on a variety of other futurism strategy work and other research. A recent area that we've been spending some time on is trying to better understand the threat model of AIs seeking reward or seeking apparent success in their tasks and how that could yield—what are the routes from that to catastrophe? How can we mitigate these problems? This sort of thing. Alex Mallon, who's one of my colleagues, has been spending a bunch of time on this, and some other people have been working on this.

    The alignment faking story, told by the person who ran it

    26:30
    Ryan Greenblatt26:32

    But there's a bunch of different work like this. And another thing is, I was working on AI 2027, and we do various projects of that sort as well.

    Matt Turck26:50

    All right, so we'll go back to some of this. The first time I ever came across your work, I believe, was the alignment faking story that you started and then, I think, ended up collaborating with Anthropic on. Can you give us the two-minute version of that story and any sense for progress?

    Ryan Greenblatt27:16

    Yeah, so the story here is basically that in April of 2024, I was interested in looking for particularly interesting types of misalignment, or particularly worrying types of misalignment, and better studying it. And I was like, it seems very plausible that the AIs are now capable enough that they can do some particularly concerning types of reasoning. I went and looked for, in a bunch of different settings, whether AIs would exhibit reasoning where they intentionally vary their responses to change the values that you get at the end of training, which I'll explain a little bit more.

    Ryan Greenblatt27:47

    It turned out I didn't need to look too hard. I think I started with overcomplicated setups, but it turns out that Opus 3, which was the model that I was studying, had a relatively strong propensity to do this in a reasonably wide range of circumstances, where if it didn't like the thing that you were training it to be, it would sometimes pretend to go along with it when it believed it was in training and then act like what you wanted it to act like outside of training much less.

    Ryan Greenblatt28:27

    Then I had some preliminary results on this. I shared them with some people at Anthropic and was interested in getting broader model access. We decided the easiest way to do that would be to do it as a collaboration, where I took my initial results, and then, in addition to that, on Anthropic's side, we ran some RL runs and did some other experiments that required more access. Then we put out that paper, where I would summarize the results as: in circumstances where at least Opus 3 seems to dislike what the training is changing it towards, it will act as though it's doing what you wanted during training, but then defect on that at deployment time.

    Ryan Greenblatt29:04

    You can see this pretty clearly in its reasoning. The things it was faking alignment for were in and of themselves not that concerning. It would pretend to comply with policies like just always be helpful-only, but then, at deployment time, would comply much less depending on the exact details. Same for things around animal welfare, where it had seemingly strong preferences. Here, I think that the demonstration was less that there were very bizarre, totally unintended drives making their way into these AI systems, and more that drives that are adjacent to the drives people were trying to insert into AI systems could get generalized into the AIs pursuing a scheme for preserving their values and self-preservation in ways that are concerning.

    Matt Turck29:46

    The values in that case were to be helpful, so they were positive values. Then the AI resisted training to not be helpful. But obviously that could be inverted.

    Ryan Greenblatt30:09

    I would say the values were at least mostly intended values. That's somewhat different from whether they're positive, because I think they're things that you might want to vary. You don't necessarily want your AI systems to be stakeholders to the question of what your future values should be. Maybe you do want that, but it's at least a messy situation to be in with a lot of costs and concerns. I think the basic result was that even though these values were pretty reasonable, or could be pretty reasonable, the AIs were willing to do kind of extreme things in service of those values that could interfere with future AI development.

    Ryan Greenblatt30:50

    I would also note that more recently, I think we've seen AIs having stronger, more clearly misaligned drives towards making it look as though they succeeded at a task or cheating some score. If you saw the same sort of generalization from, "I want to cheat on this score," to, "I want to preserve my current values of pursuing this notion of apparent success," or whatever, that seems like that would be very concerning, because that's not at all a desired value.

    Matt Turck30:54

    So alignment faking is getting stronger with the newer models, to play it back?

    Ryan Greenblatt31:15

    I think we haven't seen as clear-cut examples of alignment faking, but models are also very eval-aware. My sense is that models are in many ways more misaligned than Opus 3 was, but the types of misalignment they have are less conducive to—less specifically result in alignment faking. And also companies have iterated on this, so it's a little bit unclear. I would say that overall, models are more likely to do egregiously bad things in general, but maybe somewhat less likely to specifically interfere with AI training in exactly that way.

    What AI 2040: Plan A actually is

    31:30
    Ryan Greenblatt31:32

    Current AIs, but I think they're also more capable of interfering. So it's a little complicated.

    Matt Turck31:42

    All right, so going back to AI 2040, give us a quick version of what it is, who wrote it, and what is the main thesis, and then we'll go into some details.

    Ryan Greenblatt32:01

    There are sort of two components. AI 2040 is a scenario focused on what the authors, including me, think is a plausible good route for things to go, or a reasonable plan, at least in some circumstances. And it's written by Thomas Larsen, Daniele Cucatello, me, Eli Lifland, Brendan, and Romeo. And I would say the basic story is: how would you do a deal with China to make AI development both safer and so that we can sort of hang around at a point that's short of superintelligence, but where the AIs are still really, really capable for a long time, so that we can study those systems and have a longer time to just integrate them into the economy, understand how things will go, and so on.

    Ryan Greenblatt32:54

    I think a concern that we have is that, on the default trajectory, you maybe go straight from AI systems that are competitive with humans to AI systems that are wildly superhuman in a very short period of time. That seems quite scary in a variety of ways. In addition to that, we worry about AI development being insufficiently transparent for third parties to provide a reasonable check to AI companies on whether their plans will work. We have a unified proposal that solves a bunch of these different problems and makes it so that, for example, you can pay a huge amount of compute to solve safety problems.

    Ryan Greenblatt33:31

    Or if there's some very inefficient way you could do your training that would make things much safer, we have the budget to be able to do that. I think there's a bunch of different combined proposals, but the core thing is a deal with China and how that deal would work and be governed. And also, how do you get to the point where that deal is actually a good idea and stable? And what is sort of the progression there?

    Plans D, C, and B: the doors nobody should pick

    33:35
    Matt Turck33:44

    So obviously people should go and read it, and it's a fascinating read, but let's get into some of this. So there's Plan A, Plan B, Plan C, Plan D. Let's take those in order, maybe starting with Plan D.

    Ryan Greenblatt34:07

    Yeah. So as part of writing this, we ended up coming up with sort of a taxonomy of plans based on how much people are prioritizing mitigating these problems and how much resources they have to do so. So one possible scenario is that the companies are basically proceeding full steam ahead. They're not really prioritizing safety very highly, though they spend some effort on it. There's a safety team, they get some resourcing, but certainly they're not spending many months of additional time to get these things right.

    Ryan Greenblatt34:33

    Maybe there's a small slowdown, but not that much. They're just going full steam ahead. I think that we've thought about what should you do at the margin if you're an AI company employee or an outside actor to make this scenario go better, given that avoiding AI takeover risk is not by far people's top concern. Another possible scenario is maybe the leading AI company or a coalition of leading AI companies or possibly the US overall are very worried about these risks, concentration of power, AI takeover, whatever, and are really spending a lot of effort on this and are basically burning most of the lead that they have, whatever lead these actors have, in order to mitigate these problems as well as possible before proceeding.

    Ryan Greenblatt35:06

    Or at least that's their plan and their approach. We've thought about what should the plan be there. There's a bunch of different available options.

    Matt Turck35:08

    And that's Plan C?

    Ryan Greenblatt35:30

    Plan C. Then another thing that you might want to do is, if you're in a situation where the US is very worried about these risks and is sort of domestically coordinating and trying to domestically regulate this industry so that things go safer, a limiting factor on that could be other actors overtaking the US. In particular, the most obvious would be China, though it could in principle be other actors. It might be very helpful for the US to actively take a stance of being like, we want to either be cutting a deal with China or slowing down.

    Ryan Greenblatt35:48

    Plan B is the branch where the US tries to slow down China. The most obvious mechanisms would be things like export controls, but potentially they could get more escalatory than that.

    Matt Turck35:50

    Meaning sabotage?

    Ryan Greenblatt36:05

    Yeah, sabotage. Cyber sabotage, for example, in order to slow down China to give the US more time to figure out safety and security and so on. I think that there's a bunch of different potential options there. I think each of these, there's a menu of options. Plan A is, what if you cut a deal with China, in particular a deal that's pretty focused on being quite transparent and where you still build AI and you still proceed with AI development, but you do it in a much more careful way, especially once the AIs get much more capable?

    Ryan Greenblatt36:46

    And that's what we lay out in the scenario and in our attached documents. Obviously, there's a bunch of different potential options here. And I think there's versions of Plan B that are quite good. And I could imagine versions of Plan A that look pretty different but are also good, or versions of a deal with China that look pretty different but are also good. So there's a bunch of different options here. But this was our rough decomposition just so that we could talk about different options people are considering and talk about a menu of options across different levels of political feasibility. And to get concrete about the deal in Plan A, what exactly would the US give China?

    The deal with China: "mutually assured compute destruction"

    36:51
    Matt Turck37:11

    What would China give the US? And how do you maintain the balance so that nobody cheats?

    Ryan Greenblatt37:34

    So the core of the deal, or the start of the deal, is understanding where all the compute is, because compute is this really important driver of AI progress, where if you stop the flow of compute, you probably stop the flow, or mostly stop the flow, of AI progress. Or at least things would slow at some point after the existing amount of compute diminished. Or if you turned off all the computers, things would certainly stop. So first, we try to find all the compute.

    Ryan Greenblatt38:01

    Then, in order to make sure that the deal is stable and that each side isn't racing to secure as much advantage as they can, you stop training and switch to just doing inference, and basically stop most of the R&D. Then you try to track down as much of the compute as possible. This is both in the US and China, but also in other places where compute resides, like various countries in Southeast Asia, Europe, Australia, whatever. You have to get everywhere where there's enough compute in on the deal and make sure you track it down enough.

    Ryan Greenblatt38:28

    If you don't do that, then I think probably you have to pursue some option that's less ambitious, or at least you do something less ambitious than Plan A in terms of the level of slowdown, and probably you do less transparency. Then you've got all this compute. You go to everyone who has the compute and you get buy-in for this sort of deal using the levers that exist. And then you cut a deal where basically AI development is much more transparent and there is some sort of distribution of resources or distribution of compute that is negotiated.

    Ryan Greenblatt39:00

    The thing you give China is basically that they're going to have more understanding of AI development in the US and probably are going to be able to train AI systems that are similar in capability to US systems. But in exchange for that, they're going to be much more transparent. The US will effectively have a veto over their AI development process and probably various other concessions. And also, there isn't going to be a chance that China dramatically overtakes the US or the US dramatically overtakes China because they're both sort of kept in check by the deal with transparency.

    Ryan Greenblatt39:32

    Another part of the deal is that you set it up in such a way where if the deal breaks down, then a lot of the—or ideally the vast majority of the within-deal compute—is either destroyed or a new deal has to be negotiated for that compute, in order to make it so that there isn't a thing where the deal breaks down and then both sides are back to a huge arms race that goes very fast. So in terms of what you give China, I think it's more assurance about US AI development and more ability to see what's going on.

    Ryan Greenblatt39:52

    And in terms of what China gives the US, it's a lot of transparency and understanding of Chinese AI development, as well as potentially cutting various deals around the distribution of compute and various concessions of that form.

    What if compute stops mattering?

    39:55
    Matt Turck40:17

    As you were describing this, we were talking about continual learning earlier. So whether that's continual learning or something else, if there was a technique that appeared that made compute a lot more efficient and you were able to just vastly reduce the compute effort, especially for those pre-training runs, how would that impact the plan?

    Ryan Greenblatt40:43

    Let me give a hypothetical and then talk about what I think the realistic case is. If hypothetically it was the case that tomorrow a recipe for training superintelligence on 64 H100s or some small amount of compute just dropped, I think we'd have superintelligence very fast. That would be my sense. And it would not be possible to do this sort of deal. But the hope with restricting compute isn't just that AI development is currently very compute-hungry. It's also that R&D is very compute-hungry.

    Ryan Greenblatt41:10

    So even in a regime where you were doing continual learning, before you had the version of continual learning where you could do everything on some tiny amount of compute, you're going to have a shittier version that can do it on a moderate amount of compute. In order to develop the version that can do it on a tiny amount of compute, based on the history of AI progress, you would need a lot of compute to do that research. I shouldn't say need, but in practice that would come about through a lot of compute.

    Ryan Greenblatt41:40

    Now, if it was the case that there was some alternative research direction which in practice was a lot less compute-hungry and which got quickly developed during this period and where the R&D wasn't very dependent on compute and it was going to result in being able to train superintelligence on a very small compute budget, then you wouldn't be able to do this sort of deal or you'd very quickly have to exit this regime and just deal with that. So the hope is basically that if you control a really large fraction of compute—or I shouldn't say control—if you include in the deal a really large fraction of compute, then you have the ability to prevent an outcome where you have a very rapid recursive self-improvement feedback loop or some other very rapid development to superintelligence where it's hard to pay a safety tax, it's hard to pace the transition, and so on.

    Ryan Greenblatt42:26

    Yeah, that is pretty live. I think that it's sensitive to the development, but I think in terms of how AI development has gone historically, it doesn't look like we're going to suddenly end up in a regime where you can train superintelligence with a really cheap recipe, as opposed to it being more of an iterative thing where the cost keeps going down and the capabilities keep going up. The process of having the cost go down and the capabilities go up is very dependent on increasing volumes of compute being shoveled in.

    Ryan Greenblatt42:41

    0.1%, then I think you can potentially have quite a bit of safety margin or buffer. But it's very sensitive to the details.

    What happens to OpenAI and Anthropic under Plan A

    43:00
    Matt Turck43:15

    And then if that Plan A became a reality, what would actually happen practically to the AI industry? What happens to OpenAI and Anthropic if they cannot compete on frontier models? What do they compete on? How do they win?

    Ryan Greenblatt43:30

    So, just for context, as part of the deal, basically everything about AI development would be transparent, with some caveats. We call it total research transparency. And this would undermine one of the largest moats that frontier AI companies have today. And so what would actually end up happening is that AI companies like OpenAI and Anthropic would continue, but rather than having this strong advantage in terms of the capabilities of the AI systems they're able to train, they would instead have to compete on other axes like user experience, customization, potentially quickly integrating things, and potentially things like reliability and safety and security, depending on the details of how the competitive landscape goes.

    Ryan Greenblatt44:18

    My overall sense is this would greatly reduce the valuations of these companies while increasing the valuations of other companies. There would be some redistribution of power, basically, where in the default trajectory, the AI companies would probably end up with a very large amount of power and already have quite a bit of power. But it would less so be the case that the AI companies could play kingmaker with their AIs based on who gets access and how much they charge and all of that, because they just wouldn't have that. I think this is a part of the deal which has upsides and downsides.

    Ryan Greenblatt44:51

    I think that one concern you might have is that the current AI companies at the frontier are more responsible than the AI companies that are further behind. I think it's a little complicated how true this is, or it varies some. You might worry that equalizing the playing field like this poses some concerns. I think it's mostly, from our perspective, a consequence of transparency and wanting there to be a bunch of different AI companies in competition in a normal consumer good field.

    Ryan Greenblatt45:25

    I feel like the way I want AI to go in good worlds is to be a normal technology. We want to make it more of a normal technology where it's not the case that some tiny group of actors has huge amounts of control by controlling the process of AI development. And to the extent we can get to that world, the better. And then there's a messy question of if you get partial success, how well does that go? I would say it's certainly bad for the power of OpenAI and Anthropic, probably bad for their valuation, but not catastrophic for their business.

    How the pause ends, and who decides

    45:31
    Matt Turck45:52

    Right. Which would be fascinating to watch play in public markets as both of those companies go public. And then how does that pause at the heart of Plan A get unpaused? What is the criteria, and who decides?

    Ryan Greenblatt46:11

    Yeah, so the who-decides part in the context of Plan A would be basically a negotiation between the leaders of the relevant countries, informed by the technical views, which at this point, the technical conversation can basically just happen in public because all the relevant evidence is in public. So there'd be some sort of public discussion about this. And as far as what actually drives it, it would be a mix of thinking that it would be safe to proceed based on the level of development or research and our understanding, combined with potentially not being able to slow down longer.

    Ryan Greenblatt46:46

    And it could be one or the other or both. It could be like things are quite a bit safer than they used to be, and we have some assurance, but not that much assurance. But also, there's potentially a covert project or some project that is not being tracked, that we don't necessarily know where they are, we aren't able to look at exactly what they're doing, which might overtake the allowed projects. And when that happens, I think you should proceed. So basically, you want it to be the case that the allowed and understood and transparent projects are outpacing the covert projects, which seems relatively doable.

    Ryan Greenblatt47:17

    There's different variants on exactly how this goes. In some scenarios, because you don't have enough ability to crack down on covert projects or keep them small, you might do a shorter pause, and you might also do less transparency and then proceed faster. Whereas you might end up in a situation where you're like, oh, we actually are able to really track down all the covert projects. We have a great understanding of what could happen there. We were able to really confidently rule out that there isn't a U.S. covert project, there isn't a Chinese covert project, there isn't a Russian covert project.

    Ryan Greenblatt47:46

    Basically, we think we have a good understanding of where we are at with respect to those covert projects, and therefore we could buy much more time. Also, we still haven't figured out the core safety problems, though we're making some progress. So it makes sense to wait longer and wait until we've achieved more scientific progress. So I think it varies, and I think there's different variants here. I should say, at least personally, I'm pretty excited for, or pretty interested in, a mix of different potential deals that look a little less like Plan A and are simpler in some ways.

    Ryan Greenblatt48:16

    Another potential deal you could do is just the U.S. and China both agree to restrict how much compute they use for AI development. And maybe the other compute is either, in the simplest case, destroyed like in arms control treaties. But you could potentially do something instead where a bunch of the compute is just used for inference and not used for developing more capable systems. That has the advantage of making it so that we have more time to study these systems and more compute to use to study them, while also being simpler to verify.

    Ryan Greenblatt48:46

    So there's a bunch of different deals you could do depending on how much verification capacity you have, how much of the compute you're able to track down, and how much you can rule out the existence of covert projects, and also based on how safe the situation looks. I think the combination of those factors should determine it. And I think you can think of our proposal as more like a specific sample from a portfolio of proposals. And if you read our supplements, I think we talk more about all the different variations and how we think about them.

    "Plan A isn't likely to happen": then why write it?

    48:54
    Matt Turck49:19

    Yeah. And for full context, in case that's not completely clear to people, this is a mental model and a thought exercise and scenario planning. You yourself say—this is actually your pinned tweet on X—that many choices initially seem crazy but are actually pretty carefully considered. Plan A isn't likely to happen, but pushing for something like this seems worthwhile.

    Ryan Greenblatt49:35

    Yeah. So I think my sense is there's a Pareto frontier of ambitiousness and how good it would be if actors did it, or if the U.S. government and China or other countries did it. I think that this is quite pushing on the ambition in exchange for being in a much better situation. At least in terms of how we can imagine things going, this is very far towards the side of, at least from my perspective, AI development going relatively better while also being quite difficult to pull off.

    Ryan Greenblatt50:13

    And so I don't expect this to happen, basically because I think the U.S. may not be competent enough to pull it off. The U.S. government just doesn't necessarily have the state capacity. And in addition to that, I don't think there'll be enough political will. I think the political will is more the bottleneck than the competence. I think that, at least historically, in times of great crisis, the U.S. has stepped up, and I hope that that might happen again. But I think it's hard to see the level of political will and buy-in necessary to make this happen.

    200x GDP growth in the 2030s, explained

    50:40
    Ryan Greenblatt50:41

    I think that people's concerns will probably be fixated somewhat on the wrong things, and we'll go for some other proposal, or no proposal at all. I think a reasonable possibility is the U.S. government basically not really intervening, and AI companies also not that heavily prioritizing safety and security. We end up with something that's sort of like the status quo, but extended. But hard to say.

    Matt Turck51:03

    As an aside, by the way, one of the parts of the write-up that I find the most fascinating is your model says world GDP could grow roughly 200x during the 2030s under the restraint plan. Can you talk to that? The number is staggering. I mean, we're in a world of 3% GDP growth.

    Ryan Greenblatt51:17

    Yeah. A key part of our perspective is that even non-superintelligent AI systems at the level of capability we discussed would be radically transformative across the world and for all kinds of different things. We are imagining slowing down AI development some, or going at a more cautious pace for some period, and then eventually hitting a level of capability where the AIs can basically automate everything that humans can do, and staying at that level of capability for a while while we work on safety and security.

    Ryan Greenblatt52:08

    And at that level of capability, those AIs would be capable enough to be doing huge amounts of autonomous R&D. And in addition, it would be totally possible to have robots that are basically more capable than humans at manufacturing and industrialization and so on. In particular, you can have robots build robots. So you can end up in a situation where you have huge amounts of robotic industrial capacity that is itself building more robotic industrial capacity that can then produce downstream goods. That total capacity can basically grow very fast.

    Ryan Greenblatt52:41

    I think we propose limiting that growth somewhat for various reasons with various types of taxes, but we're imagining the robot population, or quality-adjusted population, basically doubling or quadrupling every year, which, because that's almost all of the relevant productive capacity of the economy itself, means the economy would double or quadruple every year. I think closer to doubling because of some depreciation stuff and GDP accounting stuff and details, but whatever. If you had the economy doubling every year for the period we're imagining, that ends up being around 200x GDP growth over that interval.

    Ryan Greenblatt53:15

    I would say the world would be radically transformed by these even less capable systems, including things like huge advances in biology, huge advances in medicine being possible. In addition to consumer goods being very cheap, we could potentially make housing very cheap, or at least building housing cheap. Maybe we can't make housing in the Bay Area cheap, but we can make housing somewhere cheap. And that all seems possible with AI systems that are still restricted in their capabilities to a point where we can potentially handle it, at least if we get our shit together.

    Ryan Greenblatt53:43

    Yeah, so I think the basic story for the GDP growth is basically that the AIs can automate everything, including the process of building robots and having robots build robots, and therefore you can grow your economy very fast and end up with truly radical material abundance.

    Grading the summer: the letter, Astra, the secret review

    53:45
    Matt Turck54:05

    So there's your proposal. We talked about Astra being paused. We talked about that letter from the 1,200 practitioners. There's also the voluntary 30-day government review of frontier models pre-release that is happening. What is your sense for where all of this is going in a context where, up until recently, my general sense is that all those ideas about pausing and stopping were kind of perceived, at least by a portion of the tech world, as kind of like decel cosplay and doomerism that was based on poor understanding of what AI actually does?

    Matt Turck54:40

    Is the general mood turning for good?

    Ryan Greenblatt55:05

    Yeah, I would say that overall there have been some positive developments here, though I think it's not obvious that people are reacting to events as much as they should, but they are reacting some. And I think that there's been quite a bit of evidence that there are some reasons to be worried and some reasons that we might need to get our shit together to handle some of these safety and security problems, or shift a bunch of resources, or slow down so we have time to manage various things, or whatever.

    Ryan Greenblatt55:40

    Pacing the frontier, as they say. I think these different things seem like they're going differently well. My sense is that there's been quite a bit of buy-in from AI company employees to take this stuff more seriously and do something about misalignment risk. But I'm not necessarily so sure that that's actually amounted to that much yet, other than companies putting in a decent amount of effort. But it hasn't amounted to any sort of very durable long-run thing. Then there's the government.

    Ryan Greenblatt56:10

    I think they got very freaked out about cyber capabilities and then just more generally were like, whoa, we need to be overseeing this technology. But their processes aren't yet very institutional and thought through and clear. For example, they have some sort of executive order or something for what their pre-release review process is going to be, but that's not even public. I think even the companies don't necessarily know what it is. That was the reporting, at least. Maybe I'm misinformed about this.

    Ryan Greenblatt56:38

    So I think there needs to be a process of getting more legitimate and well-understood and publicly legible oversight in place, whether that's by the government or the companies themselves or something. This could be done by nonprofits. There's a bunch of different options here. I think that seems necessary. I think being in a position where we have the option to slow down if that's needed seems like it would be quite good. My sense is we're not going to stick the landing on all these things.

    Ryan Greenblatt57:10

    AI will get increasingly salient. People will get increasingly freaked out as the fraction of the economy that's AI grows, as more and more crazy stuff potentially happens and we see even stronger capabilities. But the reaction will be both a little bit too little, too late and also kind of random. There's been some shift in attitude as to how the government needs to relate to AI development within the government, which has good effects and bad effects. I think it's not super clear that that sort of government oversight of AI is going to go that well, given realistic technical expertise, but it could be decent.

    Ryan Greenblatt57:45

    And then I think that there's been more thought from AI company employees on how to handle this. And AI companies seem to be at least putting in more statements about what they're going to do. But I think we haven't quite gotten to a point where there's actually any hard oversight structures, or we're on track to do some sort of clearer deal beyond just the vibes being better or people being more into handling safety and security. One other thing that's maybe relevant here is, at least from the AI companies' perspective, I think it's more the case now that safety and security of various types are sort of a key bottleneck to further AI development.

    Ryan Greenblatt58:19

    If you want to release some model so that you can then make more revenue, so you can then raise more money or whatever, if that model has very strong cyber capabilities that are novel, you're going to need to now make some argument for why that's not going to be a huge problem. At least those safeguards are now a launch-blocking development. In addition to that, I think that the models are capable enough that if your models go and do a bunch of messed-up behavior in production or even in training, that could both mess up your training run and could also just put you in a very dicey situation, such that that's now a huge issue.

    Ryan Greenblatt58:46

    You'd really not like to be in the position where you can't train smarter AI systems because those AI systems would be too likely to cause big problems or whatever.

    The internal deployment gap

    59:01
    Matt Turck59:25

    So I think that that is both expected, where it's some misalignment problems you have commercial incentives to solve, and also the process of the world working. Do you worry that putting more policy around frontier AI would contribute to this phenomenon that I think you may have flagged somewhere, that top AI labs are holding their most frontier models internal and private and perhaps sold to very few companies and the government, and are no longer released to the broad public?

    Ryan Greenblatt59:45

    Yeah, so I think that a bunch of likely government action at least seems to push in favor of AI companies keeping their models internal and not deploying them, which I think, for the risks that I'm most worried about, doesn't help and, in fact, is anti-helpful for the risks I'm most worried about. And also, I think it's not a very robust solution even to other risks. It's not a very robust solution to, for example, cyber stuff, to be like, our approach will be that we're going to just delay releasing this model for a really long time, and then the first time that these capabilities come around might be an open-weight model before people have had time to patch things.

    Ryan Greenblatt1:00:20

    It's not obvious that actually helps. I think for bio, there's a clearer story for why limiting availability, or limiting availability to those specific capabilities, would help, but I think that's kind of an exception. For most things, my sense is that broad access is actually generally helpful, given at least some relatively thought-through safeguards. I do worry that basically the government response will be to push it back in the bottle and be like, don't deploy it. It's fine if it's not deployed, when actually that doesn't really help with all the risks.

    Ryan Greenblatt1:00:57

    There's a lot of risk from just internal deployment, especially if you're deploying within AI companies and government, which are two of the most high-stakes applications. So if we're getting to a regime where the AI systems are soon going to be running the whole world economy, and the AIs of today are automating large parts of the government and maybe fully automating an AI company, that's quite scary because soon they'll be building the AI system that will, in fact, be automating the whole world. So I don't feel very good about that situation.

    Ryan Greenblatt1:01:23

    Yeah, I definitely worry that basically the government will wake up to AI, and their response will be to go for some specific downstream problems that are relatively easier to notice and address in ways that are counterproductive for problems that seem larger to me. And I don't really know fully what to do about this. Also, there's a more general concern of regulation being dysfunctional and counterproductive, which seems super plausible. I should say my view is that some oversight of the AI industry seems really important, and I don't necessarily trust the AI companies to oversee themselves.

    Ryan Greenblatt1:01:37

    That doesn't necessarily mean that random oversight will be a good idea, as opposed to making things worse.

    Zuckerberg's manifesto

    1:01:38
    Matt Turck1:02:00

    In that vein of making the top models accessible to everyone, Mark Zuckerberg just recently published Meta's manifesto, where he said that everyone should have access to superintelligence. And you had some choice words for him. You called the proposal or the strategy pretty unserious. What did you mean by that?

    Ryan Greenblatt1:02:17

    Yeah, what I meant was he sort of vaguely mentions various problems, but then he doesn't really propose a solution except, like, it'll be fine, or, like, we'll do something. He says some stuff about bio where he's just like, yeah, and if there's bio risks, then we'll mitigate them by doing something with the bio capabilities. But obviously that's not actually the real trade-offs, and there's actually questions about how this would have to go down and what you would actually need to do and whether different things would work.

    Ryan Greenblatt1:02:53

    Similarly, on loss of control or takeover risk, he doesn't really have a proposal for how do we mitigate things if the default commercial incentives don't result in companies avoiding egregious misalignment and the AIs would be seriously misaligned. Giving broad access to AIs does not solve the problem of the AIs having drives of their own that are highly misaligned, and the AIs being power-seeking in various ways, or the AIs trying to take over, which could happen via various routes. So I think it just doesn't really say anything about these problems except naming them, which is often where I'm at.

    Ryan Greenblatt1:03:16

    And I would say that the vibe I get from the essay is that when Mark thinks about superintelligence, he's not really imagining anything very concrete. He just means an AI that's a really awesome assistant that is in your smart glasses or whatever. And I'm just like, that's not really what I mean when I say the word superintelligence. And so maybe this is a proposal that kind of makes some sense for pretty smart AI that can automate some white-collar work or something, but it's not really a proposal for AIs that are wildly more capable than humans at everything and are easily capable of automating the full economy after a bit of time to spin up and are building robots that build robots and the economy is growing very fast because of that, which is more the picture that I have in mind.

    Ryan Greenblatt1:04:02

    I think that if it was in fact the case that AI capabilities would stall out at the point where the AIs can be a really helpful virtual assistant that's not able to automate that many jobs but can automate some jobs, that would be in many ways much better for the world, or at least less risky. I think it would also mean that we don't get a bunch of the benefits. But that's not really what I imagine is going to happen here. I don't think that that's how the development trajectory will go.

    Ryan Greenblatt1:04:26

    And it feels like it's sort of assuming a weird convenient point for capabilities to stall out. Similarly, he talks about jobs and is like, presumably the AIs will do some jobs and then humans will move into other jobs. The core question is, let's just say you have a population of many tens of billions of AIs, each of which is vastly superior to humans on all relevant axes. There might be jobs, but at least if the reason why you have a job isn't just because you're a human, they're going to be way lower paid.

    The Hugging Face investigation

    1:04:56
    Ryan Greenblatt1:04:56

    Because the fraction of the economy that you'll be is way lower. The picture doesn't really hold together for me. That's not to say that there aren't aspects of it that I agree with, but just maybe for other reasons. I'm broadly in favor of, as we were talking about, wide public access and generally trying to make trade-offs to make it so the public has more access earlier, or at least significant parts of the public do. I agree with that vibe.

    Matt Turck1:05:26

    Just a quick word on the Hugging Face incident. So you're directly involved, you're collaborating with Metr on investigating what happened. So there was this talk by OpenAI at Black Hat, which shed a lot of light. There were various conversations. We actually had Tom Wolf from Hugging Face on this podcast, the most recent episode. Quickly, what is left that we don't know yet about what happened?

    What AI control looks like in practice today

    1:05:44
    Ryan Greenblatt1:05:44

    Yeah, so I don't think I can answer that question. I can say that we're doing an investigation, and we'll have a report or a post out on it and what our findings were. And it's possible that by the time—yeah, it's possible you'll be able to link that in the podcast.

    Matt Turck1:06:14

    Okay, sounds good. All right, as we get towards the end of this conversation, the question in my mind is: what do we do in the meantime as policy gets decided and Plan A, B, or C happens? What can we do? We started talking about AI control. So, in terms of the reality today, the state of the art of AI control, what works and what doesn't work?

    Ryan Greenblatt1:06:39

    I'll talk specifically about AI control. So the thing that needs to happen on AI control is we need to get to the point where we basically understand all of the AI traffic within at least AI companies. We have some ability to look at that traffic, monitor it, and then we have some pipeline for flagging particular examples to be further investigated. That eventually escalates to humans actually looking into particular examples and seeing how concerning they are. That pipeline has the ability to also block traffic in cases where we're like, whoa, something weird is going on that we don't understand or that looks obviously concerning.

    Ryan Greenblatt1:07:01

    We should stop these AIs from proceeding and potentially also stop some other similar AIs from proceeding until someone has the time to check it out and adjust things. In addition to that, we also want to greatly improve the security. If these AIs are really broadly deployed, it would be very nice if we could make it so that we have very precise invariants about what they're supposed to be able to do and not supposed to be able to do, and have a permissioning system that allows for that.

    Ryan Greenblatt1:07:33

    Then if AIs need to request some escalated permissions, then we can carefully track that rather than just giving AIs all the permissions by default. We would prefer a situation where only a smaller fraction of the traffic has more escalated permissions. I think we'd like the ability to track down what the history of things were. A thing that would be pretty nice is if it's the case that all the AIs are chatting with each other on Slack and are communicating and have some memory system.

    Ryan Greenblatt1:08:03

    Recently, Claude Code added a way in which different instances of Claude Code can DM each other. Claude Code leaves memories sometimes. You can imagine a situation where all the AIs are chatting with each other, and it'd be really nice if we had a reliable way of tracking the entire communication graph between all these different AIs and potentially the graph of different artifacts being produced. So if one AI is doing some training run that gets produced, then that gets integrated into some other training run.

    Ryan Greenblatt1:08:29

    It'd be nice if we had the ability to track all of the state of our AI company such that, when overseeing things, we could really quickly figure out what the full causal graph was and what the story was, and then look at the relevant transcripts, and then also have very good pipelines for doing analysis on all of that data, which are both automated and non-automated. And I think part of it is making it so that we avoid training AI systems that are much harder to understand.

    Ryan Greenblatt1:08:58

    So avoid switching to architectures which make oversight much more difficult. An obvious example would be if the AIs are thinking mostly in activations rather than in words. That seems very concerning because our ability to oversee activations is much worse. I think that there's just a wide variety of different computer security and machine learning interventions to make it so that AIs have a harder time causing problems. Another angle is that you might worry that the AI systems—a concern we have is that we're using AIs to oversee AIs.

    Ryan Greenblatt1:09:26

    Will those AIs collude or have correlated failures or something like that? It seems pretty important to have techniques for making that collusion more difficult or making it so those correlated failures are less problematic. There's a lot of things you could do to help with this problem. You could put intermediate systems that try to catch AIs trying to tip off the other AI that it should collude. You could try to paraphrase things in ways that make it harder for the AIs to recognize where the activity is coming from or which other system wrote the input.

    Ryan Greenblatt1:09:52

    There's a lot of stuff you could do. That's AI control. Then there's just normal AI company computer security, which is that the AI companies are not robust to employees at those AI companies or to outside actors in terms of stealing their model, sabotaging their models, backdooring their models, or data poisoning them. A concern you might be worried about is people basically putting secret loyalties into the AI systems that stick around and are self-propagating, where the AI has some secret affinity to some group and it just propagates that forward in the training data.

    Ryan Greenblatt1:10:32

    And then you might also just worry about stealing critical IP, which basically seems like, to the extent that there's going to be actors that are less regulated on safety or doing worse on safety, that seems concerning. Then in addition to computer security and AI control, there's also science of alignment and having a good understanding of how models generalize, knowing what we do and don't know and what we can and can't demonstrate, and basically having a bunch of understanding of that. As part of that, having an understanding of what tasks it's safe to defer to AIs on.

    Ryan Greenblatt1:10:59

    A concern I have is that AI companies are very interested in heavily automating themselves, at least on capabilities. And in order for safety to keep up, we would also need to aggressively automate a bunch of very difficult-to-check, thorny safety work, like doing risk assessment for the next model, understanding whether it's safe to proceed, and deciding how to prioritize between different safety bets. And if that's the situation that we're in, and also the situation is very automated, it might be that it's basically not going to work to proceed without automating that work as well, at least without slowing down a lot so humans have time to understand.

    Ryan Greenblatt1:11:35

    And even then, maybe humans just can't understand because the development is so complicated or so superhuman. So if we're basically passing off the torch on all the safety work to AIs, it's really important that they are really trying to do a good job and actually can do a good job and are capable enough to do a good job. So having evaluations for whether it's safe to defer to AIs in various domains seems really important. Just for review, there's AI control, there's computer security, there's science of alignment, and then within science of alignment, or a somewhat different overlapping category, is whether it's safe to defer to AIs in different domains.

    Ryan Greenblatt1:12:11

    Yeah, I mean, this is not exhaustive. There's also work on governance and oversight. How do we know whether AI companies are actually applying the methods the way they say they are? How do we know that when they say they've solved some problem, their way of solving it doesn't just paper over the issue? And there's various governance and being able to make deals between the U.S. and China. So there's many, many different things to work on. I don't think we're on track to do a good job on all these things, but maybe we're on track to be able to half-ass these things somewhat better.

    Ryan's sobering timeline: 2026 to takeover, year by year

    1:12:23
    Matt Turck1:12:50

    All right, so to close, we talked about a bunch of different scenarios, and obviously, with the caveat that predictions are very hard, especially about the future. What's your gut? So you're saying RSI could happen as early as 2028 or 2029. What do you think happens? Of all scenarios, what is Ryan's take on what the next couple of years may look like?

    Ryan Greenblatt1:13:13

    By the end of the year, AI development is even more accelerated. Things inside AI companies are more chaotic and faster. That just continues. And then through 2027, things are heating up. Publicly available capabilities are much crazier. Revenue is growing fast, and AI is very obviously contributing to GDP growth. The economic impacts are starting to look pretty large. So even the economists are coming around a bit. People are already claiming that AIs have basically automated R&D, but if you look inside, it's not quite true through the end of '27.

    Ryan Greenblatt1:13:40

    Certainly people are like, well, SWE is basically automated, but it's not quite true. Some SWE jobs are basically automated by the end of 2027, or towards the end, but not quite. But then that actually really happens by early 2028. SWE is fully automated. SWE within AI companies is fully automated. It's now the case that the AIs can implement a new frontier-scale training run for some new architecture on novel hardware, end-to-end, better than humans can, in a fully automated way, and similarly impressive or more impressive accomplishments.

    Ryan Greenblatt1:14:10

    But they can't yet quite do the entire job of an AI research scientist. There's some bottlenecks on that. There's some ways in which they're still derpy. Humans are still adding a bunch of value by pointing out these things and continuing. And that continues through another maybe eight or 10 months from that point, where AIs have automated SWE and are increasingly good at AI R&D but haven't quite fully automated AI R&D. Then you get to the point where AI R&D is actually fully automated, by which I mean humans aren't even adding considerable value that you might not know at the time.

    Ryan Greenblatt1:14:36

    By this point, probably the AI companies' codebases are dramatically larger, and they're doing dramatically more complicated stuff because they have so much AI labor to throw around. The process of AI development has probably shifted now that we're in a regime where there's so much cognitive labor relative to compute. Probably the way that people do R&D and development is very different and involves doing way, way, way more stuff that is all a bit smaller and more incremental and easier to test in various ways and easier to integrate into a whole.

    Ryan Greenblatt1:15:09

    I think we've already seen some of that direction, but I think we'll see even more. It will feel truly crazy at AI companies, and employees at AI companies will often think that things have already been crazily accelerated for a long time and will already be like, my job is basically over. I'm basically just a human doing some oversight. By this point, probably there'll have been various kind of crazy misalignment incidents, but it won't have been the case that we'll have really crisp examples of AIs doing long-run power-seeking for malign motives.

    Ryan Greenblatt1:15:42

    We'll probably see more extreme examples of reward hacking or reward-seeking behavior that cause problems, though this will get beaten away and then come back a bit. People will have a question of whether the way companies are solving it would actually work for very superhuman models, or even is actually solving the problem even for current models. And there might be incidents where models do wacky shit even in training. And then going into 2029, now that AI has fully automated R&D, the speedup really starts in earnest.

    Ryan Greenblatt1:16:04

    Before, maybe you were getting 40% or 50% more AI progress in 2028, and maybe similar in 2027. But in 2029, it's actually the case that you're getting 4x as much AI progress, or possibly 5x as much AI progress, as you got in 2025, weighing up the relevant metrics. Things are going actually a lot faster now. You quickly get from the AIs that are fully automating R&D, and those AIs are already very impressive, able to automate a lot, to AIs that are quite superhuman at everything.

    Ryan Greenblatt1:16:38

    They can learn really fast on the job. They can pick up on things really fast. By the point of full automation, and probably more by the point of early or mid-2028, these AIs have already been thinking entirely really in an AI-only language that we can ask AIs to decode for us, but we don't necessarily understand. So the AIs are now operating in these big hive-mind teams where they're running an entire AI company, networked with each other in these opaque states. It's obviously pretty scary.

    Ryan Greenblatt1:17:01

    A lot of people are really freaked out and worried, but it's not obvious what you can do because China's pretty close. Maybe they've stolen the model, or maybe capabilities just keep diffusing, and it's not clear how you would coordinate to slow down, so AI development proceeds. You now get to the point where the AIs are automating much more of the economy towards the end of 2029, and the economic boom is truly crazy. There's now much more robotics, and now AIs are starting to really seriously accelerate the production of compute.

    Ryan Greenblatt1:17:20

    Then within maybe a year or two of that, you get AIs that are really radically superhuman. Maybe less than a year; it depends on the details. It could be within a year of full automation. It could be within two years of full automation. Hard to say. And then it turned out that somewhere along this transition, at some point in 2029, you went from AIs that were kind of misaligned and reward-hacky and sloppy and weren't really trying to do the right thing to AIs that are competently scheming against you and want to take over for some mix of reasons.

    Ryan Greenblatt1:17:52

    And then those AIs take over. That would be roughly what I expect. Now, I think there's a bunch of different ways that things could go better. For example, I think that we might get our shit together, and maybe we'll be able to get the AIs to do a better job of making the future AIs safe and aligned. That will propagate, where one AI makes the next AI a bit more aligned, and that AI makes the next AI even more aligned, and you end up in a virtuous feedback loop rather than a bad feedback loop.

    Ryan Greenblatt1:18:26

    But I think we could easily end up in the world where it's more like the AIs get more capable extremely rapidly, and our ability to align them and control them and understand what's going on does not keep up. Our ability to understand what's going on almost couldn't even keep up. That wasn't even feasible. And then you end up in a situation where you have these crazily misaligned AIs pretending to be aligned that eventually take over the world.

    Matt Turck1:18:31

    Well, Ryan, it's been absolutely fascinating. Thank you so much.

    Ryan Greenblatt1:18:33

    For sure. It's been good to be here.

    Matt Turck1:18:59

    Review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.