MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    OpenAI's Dan Roberts: Why AI Can Now Make Discoveries

    Dan Roberts is the Lead, Foundations of Reinforcement Learning at OpenAI. We cover how reinforcement learning teaches models to use test-time compute through language-based thought processes, why OpenAI’s unit-distance result came from exploring a conjecture mathematicians assumed was true, and why AI’s role in science is likely to expand gradually while research taste remains hard to verify.

    06/04/2026

    Hosted by Matt Turck · with Dan Roberts, Lead, Foundations of Reinforcement Learning at OpenAI

    reinforcement learningtest-time computeAI reasoningmathematical discoveryAI for science
    Listen now
    YouTubeApple PodcastsSpotify
    49 min · 26 chapters
    Contents

    Transcript

    What OpenAI's Foundations of RL team does

    1:21
    Matt Turck1:20

    Hey, Dan. Excited to do this. Thanks for taking the time.

    Dan Roberts1:22

    Of course. Very happy to be here.

    Matt Turck1:29

    You are the lead of the Foundations of Reinforcement Learning team at OpenAI. So what does that mean?

    Dan Roberts1:59

    What does that mean? The larger team that we're on is called Foundations, and we think about reinforcement learning. So, very boring: Foundations of Reinforcement Learning. But the team comes from a mandate of thinking about the science of reinforcement learning. And a long time ago, which in AI speak is like six months ago, maybe a year—I guess now two years—so before we released o1 and reasoning models, we were studying this internally. And one of the advantages to being first, or at least being forced into spending a lot of resources on scaling things up, is that you can empower a group of people to not just work on making the thing work, but work on understanding how it works.

    Dan Roberts2:44

    And then beyond that, how do we scale? How should we think about scaling reinforcement learning versus scaling pre-training? So what do scaling laws look like? But then going beyond that, what sort of things does this kind of training teach us? What doesn't it teach us? We're very interested in, at the frontier for exploratory scenarios, how do we either improve or understand better what reinforcement learning is doing? We have all this compute that famously we are in the process of acquiring, and we would like to turn that compute into intelligence.

    Dan's journey: from black holes and quantum gravity to frontier AI

    3:08
    Dan Roberts3:09

    And to do that, we need to make thinking models. Somewhere along the way, we interact with that process, usually at the earlier stage for models, not the next model, but things that are like the next model or the next-next model. Great.

    Matt Turck3:16

    And quickly, what was your path to OpenAI? How did you go from studying physics to being where you are today?

    Dan Roberts3:47

    I did a PhD in theoretical physics from MIT, thinking about the intersection of quantum gravity and quantum information. I thought a lot about black holes and quantum chaos, kind of thinking: what if you throw something into a black hole? What happens to the information? Does it come out? If we think about black holes as computers, how fast are they? I was very interested in this fundamental question in theoretical physics, which is: how do you find a quantum theory of gravity? I also got very interested in this interplay between computation and the laws of physics.

    Dan Roberts4:14

    Any computer exists in the universe, behaves according to physical law. So the sort of computations you can do are bounded by the laws of physics, and there's some sort of interesting relationship there. Black holes are pretty interesting because they sort of saturate some conjectured bounds around processing of information. From there, I did a postdoc at the Institute for Advanced Study, and around that time—I'm pretty old now, at least for this field—so that was about 2016, was when the DQN Atari paper from DeepMind happened in 2015, and then AlphaGo was in 2016.

    Dan Roberts5:00

    I got very excited about the possibility of machine learning, and then deep learning was statistical science that lived in a similar framework to the sort of frameworks that we use to study the rest of the universe. Always this question of, how does everything work? This three-year-old question of, I'm curious about everything. If you look outward and you care enough, you end up in philosophy, maybe. If you're quantitative, you end up in physics. Very crude characterization. An AI that works is very fascinating, or AI systems that work are fascinating because they are simple examples that do things that humans do.

    Dan Roberts5:37

    Then if it lives in the same framework that we use to understand everything else, then it's sort of like, you can draw the parallels between how does the universe work and how do I work, or how does intelligence work. So I got extremely interested in AI and deep learning. Then I went to FAIR, Facebook's AI Research lab, around 2017 to basically try and use the tools from theoretical physics to understand deep learning. Deep learning was supposed to be this really difficult thing that you couldn't understand.

    Dan Roberts5:57

    And I thought maybe the tools of physics could be helpful. This actually culminated in a book that I wrote with a collaborator who is still a collaborator of mine now at OpenAI, working on the same thing, Sho Yaida. But we wrote this book, The Principles of Deep Learning Theory, that was a culmination of these ideas of: can we sort of use the statistical ideas of understanding statistical systems, like the gas in the room? We can characterize them with some simple laws of thermodynamics, like the ideal gas law, and maybe we can make similar progress in understanding deep networks.

    Dan Roberts6:52

    So that was sort of my transition. I also had a startup along the way and spent some time at Sequoia Capital as an entrepreneur in residence. So there's some tension between, am I a scientist and am I an entrepreneur? But about two years ago, after thinking about whether I wanted to start another AI company, I realized that the thing that was most exciting right now was what was happening at the frontier, that there's some amazing scientific progress happening in AI. And to really get at the questions and understand what's going on, you need to be there and you need to participate.

    Are AI systems becoming useful for real science?

    7:04
    Dan Roberts7:04

    And that meant joining a lab. So I joined OpenAI two years ago.

    Matt Turck7:27

    Great. Thank you for that. Where do you think we are in the evolution of AI being increasingly able to solve difficult scientific problems? I mean, certainly something that we've been talking as an industry about for a while now, but it seems to be accelerating, perhaps just like everything else in AI. But where do you think we are?

    Dan Roberts7:54

    I think one of the interesting things is that this process is smooth. There's no sharp point, or I don't think there will be a sharp point, where we'll say that systems weren't able to be useful for the scientific process until they're fully fledged scientists. There'll be sort of a gradual shift. If you had to point to one moment, maybe it would be the release of OpenAI o1 and the paradigm of test-time compute and reasoning. But I'm sure if I tried to make that claim, you could go and look at GPT-4 and see that there's glimpses of that sort of useful behavior for the scientific process.

    The AI math moment: Erdős, OpenAI, DeepMind, and Anthropic

    8:21
    Dan Roberts8:22

    Were already present. As a general point, the models are very good at certain types of things that clearly are amenable to making progress in math. They're not open-loop, fully fledged scientists in any domain, although neither am I. It seems like it's just this really nice gradual process.

    Matt Turck8:51

    So it feels like a particularly fun week to be having this conversation because over the last few days, there were a number of different announcements in the general field of AI and mathematics around the Erdős problems. OpenAI came out first with this progress, but almost within a few hours, Google DeepMind had a claim as well on different problems. Then Anthropic had some claims. However, from what I understand, the OpenAI approach and the DeepMind approach were very different, and that may be very interesting in terms of what that means for AI as a research scientist.

    Why the OpenAI result was an act of exploration

    8:52
    Dan Roberts9:29

    This conjecture everyone assumed was true, but could not prove it. One of the things that ChatGPT was able to do was assume it was false. And when you go against the grain and do something contrarian like that, you really have to have strong conviction in what you're doing in order to persevere down a really long calculation path. Because there's a lot of choices that you can make along the path. And if you get any of those choices wrong, if your ideas don't work, then you find out that you didn't make any progress.

    Dan Roberts10:02

    So you need this really strong persistence, and then you need expertise in this other field, which is algebraic number theory, some sort of generalization of number theory on things that sort of generalize the integers and the real numbers. If you go down that path really far, you can refute this conjecture. So that was the big result. The big result was that this conjecture of this lower bound for the number of pairs that you can make is false. Not only is it false, it was false due to a really interesting connection to another field of mathematics.

    OpenAI vs. DeepMind: informal reasoning vs. formal proof

    10:25
    Dan Roberts10:25

    So you would have to be somebody who is aware of this problem as interesting, which sounds like your expertise is one thing, and then be an expert in something else, and then also be super contrarian and go down this really long path, and then you would have identified the solution.

    Matt Turck10:34

    The OpenAI approach and the DeepMind approach were very different. Do you want to compare and contrast the two approaches?

    Dan Roberts11:03

    One of the approaches that DeepMind takes is to take problems, present them in a formal language called Lean, and then use methods to search for proofs in that language. For problems to be representable, there's this process called autoformalization where you take an English version of the problem and translate it into rigorous formal statements, and then you conduct your proofs there. It's designed so that the proofs can be airtight. No one has to go and check for some hidden assumption or some weird thing.

    Dan Roberts11:25

    I guess it's usually hidden assumptions or definitions that are not airtight. But in that setting, which is a setting that DeepMind has cared a lot about, they were able to formalize some problems and use their system to prove them. So that's one approach. Another approach is to just take the problem in English with mathematical expressions as well, but just the English statement of it, which is informal, and understand what is meant by that and solve that in informal language, presenting a proof much like the way a human mathematician would, or a human mathematician who's not using Lean.

    Dan Roberts11:54

    And then you have to check it. The verification problem is harder because it's not something that auto-checks.

    Matt Turck11:56

    And that second approach was OpenAI?

    RL 101: learning by doing, not just watching

    12:13
    Dan Roberts12:13

    Most of our results that we publicize, as far as I can think, are all in the informal setting. We have language models that we've taught to reason at test time, and one of the applications or benchmarks for that is reasoning in mathematics.

    Matt Turck12:35

    Okay, great. All right, so let's get into reinforcement learning. To make this broadly accessible, let's start from the top. What is the one-, two-, three-sentence definition of reinforcement learning? And perhaps give us a simple, non-technical analogy for people to understand.

    Dan Roberts12:54

    Maybe a simple thing to do would be to give you two examples of how you could try to learn something, you as an individual. And maybe we can take a game, or even, say, a video game, right? I'm old enough that I played the original 8-bit Super Mario Bros. And so here are two ways you could learn how to play. One way you could learn how to play is your dad takes it out and plugs it in, and he boots up the game, and then he plays for a few hours, and then you just watch him play.

    Dan Roberts13:26

    That's all you do. So he's demonstrating how to play, and then at the end of that, he's not very nice, so he doesn't let you play. But then he goes and runs outside and does something else. You sneak into his room, you plug it in, and you try to play. How good are you going to be? Well, all you've done is tried to memorize what he's done. You haven't gotten to push any of the buttons yourself. You haven't gotten to interact with the game yourself.

    Dan Roberts13:48

    This is sometimes called expert demonstrations. You're just trying to memorize what someone else is doing. It's a version of supervised learning, the supervision being like you just watch what he does and accept that that's the true way of doing a thing. Reinforcement learning would be your dad's like, "Here, why don't you play?" Maybe he shows you once, or maybe he doesn't even need to show you because the game is beautifully designed to sort of take you from not knowing anything to being able to play expertly.

    Dan Roberts14:25

    There's something called a curriculum, but you play. Maybe the first thing you do is you run, you hit the first bad guy, and you probably—this example is dated—but you lose a life. But then the second time, you press a button and you jump. And so you're taking actions, there's an environment that's giving you feedback, and there's this close connection between the environment, between actions that you can take, and then the responses that you're getting. And then the final part is there's a reward.

    Dan Roberts14:51

    And the reward can be something that you get pretty often. For instance, every time you do something, there's some score that goes up, or it could be just something that you get at the end. So you play a game of chess, and at the very end, you get a reward, which is you won or you lost. But in the middle, you don't really know how you're doing until the very end. So this is called sparse rewards. But I think this is the basic idea, and there's obviously lots of variants here and ways to quibble with this, but it's this notion that you interact with an environment, you get a reward, and often it's in a way where you get this sort of feedback as opposed to just trying to learn from data that you don't get to interact with.

    Why reinforcement learning works

    15:10
    Matt Turck15:15

    And why does it work? And why is RL so powerful?

    Dan Roberts15:41

    It works because of this ability to get feedback from the environment. You can go and learn, if you're doing it right, you can figure out how to learn the things that you don't know. And I also think it's powerful because of this fact that it's much easier to learn when you're learning at the right level for you, right? So if you want to learn addition, you shouldn't read a calculus textbook. You want to learn by being able to practice and learn at the right level.

    How RL breaks: sparse feedback and long-horizon tasks

    15:58
    Dan Roberts15:58

    I'm actually making the choices and learning from my own choices, whether they work or not, then I'm able to place it in a better context for the set of things that I understand.

    Matt Turck16:04

    Great. And then conversely, what's the catch, and how does RL break?

    Dan Roberts16:27

    The setting where it's very difficult is the setting that I alluded to before, where you don't get much feedback from the environment. You have to take many, many, many, many actions, and then you get maybe, yes, that whole set of actions was good, or no, it was bad. For instance, you're playing a game of chess and you don't know until you make all the moves. That has an opponent, so it's maybe complicated. Maybe it's you're trying to do a homework problem, and it's research-level, or someone gives you a well-defined problem, like we give our language models, and it's a problem that requires days and days of thinking. There's so many choices that you can make along the way.

    RLHF: how human feedback shaped early language models

    17:03
    Dan Roberts17:03

    And at the end, if you don't get any feedback at all, if you're just hidden in the woods by yourself, scribbling in notebooks, it's very hard to make progress that way, because you don't have any sense. If you get a yes at the end or you get a no at the end, you have no sense for which of the actions that you took, which of the things you did, were good or bad.

    Matt Turck17:17

    Okay, great. Now let's talk about how RL has been applied in the context of large language models. So, was the first step historically RLHF?

    Dan Roberts17:42

    Yeah, I think that's probably fair, at least in a broad sense, that the first kind of RL that was done on language models was part of this post-training process to turn a model that just tries to predict the next word on the internet into either something that will follow your instructions, be nice to you, or fit the form of a chatbot.

    Matt Turck17:47

    So, do you want to define for people what RLHF is and how it works quickly?

    Dan Roberts18:19

    The basic idea is that you could collect data from humans. RLHF is reinforcement learning from human feedback. So, you collect data from humans and you train a value function. So, you would show, in the language model setting, say, two different completions from a language model, ask them to say which is better. This sort of comparison could be used to train a value function, and then you can use that as a reward for the reinforcement learning process.

    Matt Turck18:25

    Great. And you do that initially with humans, but then you build that into a reward model?

    Move 37, self-play, and the search for novel strategies

    18:48
    Dan Roberts18:48

    Yeah, so you would train a model for this, and then now, because during the training process you can't just pause your training run to ask some humans for input, right? The feedback would have way too much latency. So instead, you need a proxy for what a human would say. So you train this model based on the human preference data, and then you can optimize against it, or at least a little bit.

    Matt Turck19:02

    One of the famous things in the history of RL is Move 37. How do you train a model to encourage the model to do that kind of thing and come up with brand-new ways while being efficient and exploit known paths?

    Dan Roberts19:27

    Yeah. So the great thing about Go is that you can just train it. It's a zero-sum, two-player game. You can train it in what's called self-play. It plays itself, and it can go from playing randomly to expert play, and it will find whatever the sort of best strategies are. So if that means exploring, great. If that means exploiting—actually, I have a funny story about this. So I met Noam Brown in grad school. He went to a different grad school than me, but he wanted to enter MIT's poker bot competition.

    Dan Roberts20:01

    And he had a poker bot that was the best in the world, but it wasn't something that would compete against humans yet. He just won in this research competition. He collaborated with me and another friend to enter MIT's poker bot competition. This was great, actually, for me because I learned some really exciting work in AI, and I got very excited about this while I was doing physics. We were playing essentially this kind of self-play equilibrium strategy. There's some nuances, but essentially, we could not lose, assuming we did not have any bugs in our code.

    Dan Roberts20:31

    The way this thing worked was that it was a tournament where you would be paired with, say, another person and play them in a hand. Depending on the amount of points you got in some sort of round-robin setup, they would eliminate the bottom half and keep going until you got to the final table, which would just be, say, you versus the other person. And so, there was the award ceremony, and we didn't know what happened, but there was someone else who was—what did everyone's scores over time look like?

    Dan Roberts21:08

    And there was, say, 64—I think there were 32, actually—people playing, so it was around a 32-person tournament. And 30 people, over time, their scores were all very negative and going down. And then there was one person whose score was pretty much straight up. And then there was another that was pretty good, but not with a crazy slope. And so, do you want to guess which one we were?

    Dan Roberts21:37

    So we were the lower slope. And then there was this other guy that had this crazy slope, was just completely crushing all the other players. And then this happened for the round of 16, the round of 8, the round of 4. And then in the round of 2, it's heads-up, us versus this guy who, over the course of this tournament, won way more than us overall, like taken more money from everyone else, and then we crushed him. Because why?

    Dan Roberts22:06

    Because he was exploiting the weaknesses of everybody else, right? It had some theory of mind to try to figure out, oh, this guy does this when he bluffs. And so it was very—I assume it was very good at exploiting everyone else, but we were just playing the best possible thing that you could do. So the criteria was not maximize your amount that you get from anyone else. It was don't lose. So it's the best response to anyone's strategy.

    Explore vs. exploit in scientific discovery

    22:16
    Dan Roberts22:16

    And so, at the end, we had to win, assuming we did it right, and someone else playing the same strategy would tie.

    Matt Turck22:35

    Okay, fascinating. So, just tying this back to the beginning of the conversation about the Erdős problem and solving unsolved math problems, presumably the instinct would be that you need a lot of exploration, not exploitation. So how does that work in the context of novel scientific discoveries?

    Dan Roberts23:01

    I think math research, or scientific research in general, has a lot of versions of both explore and exploit. To give the recent example, the OpenAI unit distance proof, I think, is very much in the explore setting, where the model was happy to be contrarian and try to disprove this thing that everyone believed. It has this huge repository of understanding all of human math, and so it was spending a very long amount of time—I forget how many hours, but I think we published a rewritten version of this chain of thought—hours and hours trying different things.

    Dan Roberts23:43

    So it's clearly in the domain of exploration. A lot of times, though, you can ask these models to compute something that they understand very well, and then that has a different structure and might look a lot like exploit. There's a paper that came out recently after the OpenAI result where an unrelated Erdős problem has something to do with if you have a set and you try to add the set to itself or you try to multiply the set with itself. So, take the elements and add them all together, or take the elements individually and multiply them together, and how many unique sums or products you get.

    Dan Roberts24:22

    There's some conjecture around that, and this one was also disproved, and that was done by humans. The core idea was, it's like a totally different problem, but there was inspiration from the unit distance one. The idea that you can sort of generalize from that: pick a certain type of numbers that had a certain property that the OpenAI model figured out, and they realized that this applies in this setting. So that's very much an exploit thing. And so I think the process clearly is like this, the actual discovery process.

    Why RL may now be "the cake," not the cherry on top

    24:49
    Dan Roberts24:50

    I think normally when you talk about explore-exploit, maybe we're talking about, when training reinforcement learning models, how should we train them? But I think there's this interesting point that in the scientific discovery process, there's really this interplay between exploration and then exploitation in order to totally push the field forward.

    Matt Turck25:16

    Switching to RL in modern LLM systems, there used to be a saying, which I think comes from Yann LeCun, that RL was the cherry on top of the cake. But I think you have argued that things have switched now and that RL is the main part, the cake. Do you want to just walk us through what you were thinking?

    Dan Roberts25:41

    Yeah, I said that about a year and a half ago. I had to give a talk that was public, and I couldn't say much, so I decided to invert this meme with the cake and the cherry. RL is really exciting. That's what I'm here talking about. And I think that when you have a lot of compute, you want to turn that compute into intelligence in a way that's useful. And RL is one way of doing it. And we just started doing it then, and we're going to do a lot more of it.

    Why RL started working with large language models

    25:46
    Matt Turck25:58

    Why did RL start working well? It's not an entirely new concept. It's been tried for many years now. What is different now?

    Dan Roberts26:29

    Yeah, I'm not sure, to be honest, when people say it wasn't working, what that actually means. There was this 2016, 2017, maybe even to 2018, before the transformer period, where DeepMind was all in on RL. And OpenAI had Dota 2 and the Rubik's Cube and some other exciting results as well. But a lot of people were all in on RL, and then there were language models, and the obvious thing to do was scale up the thing that worked, which was pre-training. And I don't know whether or what people tried for RL.

    Dan Roberts27:04

    As you pointed out, RLHF was a central thing that came pretty quickly. Originally, it was developed in the context of game environments, trying to prevent reward hacking by using—I think the original paper was about using human feedback to control a character to walk or something like that. But there's an interesting thing to point out here, though, which is that there's this question of how do you get models to think at test time and reason. There was a reasoning effort at OpenAI that was quite early and spent some time and came up with some algorithms.

    Is RL "sucking supervision through a straw"?

    27:29
    Dan Roberts27:30

    I think maybe the simple thing to say is that if you have a powerful enough pre-trained model, then it can start to do well at RL. It can start to use test-time compute to, for instance, solve math problems that it wouldn't otherwise be able to do.

    Matt Turck27:52

    A viral analysis from earlier this year, February, I think, claims that RL produces less than one bit of useful information per 10,000 tokens. And then Karpathy called it "sucking supervision through a straw." What is your take on this and the overall efficiency of RL?

    Dan Roberts28:21

    If you look at the DeepSeek algorithm, which is a public thing that we can talk about, then you train on sequences that are correct. So whether it's correct or not is maybe one bit of information. So I think you can see where that logic comes from. I think the question is, is this doing a kind of thing that you can't otherwise do? Maybe you would want to give more supervision, but how are you going to do that? I think it's very clear that these methods have led to a bunch of breakthroughs in terms of the explosion of what the models can do, both in coding and in science.

    Why language may be the grounding layer for intelligence

    28:47
    Dan Roberts28:48

    I think broadly it's about getting models to think at test time, to use test-time compute and do reasoning. And there's clearly a lot of the pieces of what the RL process is that's essential to make that work.

    Matt Turck29:23

    What's your overall feeling in terms of how far we can go with that current sort of systems model where we have pre-training and then we have RL on top? Somewhat famously, last year there was a conversation with Rich Sutton on the Dwarkesh Podcast where his claim, in my best attempt to paraphrase it, was that LLMs were not really intelligent and therefore RL was the only way to do it, and pure RL, not LM plus RL. What is your take on this? I mean, obviously you're on an RL team at a company that does both pre-training and RL combined.

    Matt Turck29:32

    So what's your take?

    Dan Roberts29:52

    Let me tell another story. So before I did my PhD, I spent two years in the UK, and I was at Oxford for one of those years. And I was at a pub, as one does, and two of my close friends—one was a cognitive scientist and one was a linguist. And so we had the sort of argument that you do in those situations when you're that age. And so, something like, physics is the most fundamental of all the sciences because it explains how the world works, and everything is in the world.

    Dan Roberts30:29

    I said this earlier: my computer exists in the world, I exist in the world, we all follow the laws of physics. And then the cognitive scientist said something like, yes, but then you have to process it. So there's all sorts of cognitive biases about that, and the way you collect data and learn something. But then the linguist was like, Wittgenstein, everything goes through language. That's the method of communication. That's the way words mean things—the central thing.

    Dan Roberts30:51

    And when we want to talk about the laws of physics, we have to use language. And I sort of feel like he and Wittgenstein, that was correct, right? Or at least the path through AI suggests that that is a correct path. I'm conceding now to Kyle, and if he's listening, he's now a linguistics professor. This whole idea of reinforcement learning that kicked off the previous decades' interest in AI, the sort of grounding I think that was needed to make things really work is through language, because everything goes through language.

    Dan Roberts31:19

    All the internet incorporates the grounding of the real world, all of our scientific knowledge, all of our mathematical knowledge, all of the sum total, basically, of human work is represented on the internet in language. And then so having the model have a prior of language and being able to think in language and then train on top of that, that seems clearly the right thing to do and seems also well-grounded in a way that even before all this, somebody might have argued would make sense.

    A contrarian take on the Bitter Lesson

    31:46
    Dan Roberts31:47

    It's an amazing prior to have to start with for an intelligence because it's very much based on us and our society. I have other disagreements with Rich Sutton, but if you want to poke at that.

    Matt Turck31:51

    Yes, just give us one or two quick ones.

    Dan Roberts32:13

    I have a somewhat contrarian take with the Bitter Lesson that it's not that scale is all you need. You also need to have good ideas to guide the scaling. So there's a deeper interplay than just scaling things up. For instance, if you were just trying to scale pre-training, you wouldn't get anywhere near as far as also trying to scale RL on top of pre-training, which is what we do now. And our models are much more powerful for that very good idea and investing in that good idea.

    Matt Turck32:20

    And the good ideas come from humans.

    What test-time compute actually is

    32:41
    Dan Roberts32:41

    Well, maybe they'll come from AI in the future, but before we had AI, they came from humans. Scaling was also a good idea that came from humans, but there's this interplay: you elicit new phenomena at scale, you try to understand them at that scale, and then that points you at new directions, and then you develop new ideas, and then you try to apply scale on those ideas. So I think it's not just scale, scale, scale.

    Matt Turck33:02

    Since you mentioned test-time compute, I think there's something that still puzzles people, which is the whole chain-of-thought thing, which is so magical from a user perspective, whatever you can see. What actually happens during test-time compute that creates those artifacts?

    Dan Roberts33:30

    Well, I think it does what you see it do. We lightly rewrite it or summarize it, but it just produces tokens, and those tokens are like a running thought process, just like you might have. Or maybe it's more akin to, if you're solving a math problem, the scratchpad, the collection of notes that you have, but it just keeps generating. The cool thing about generating is that it's a forward pass of the model, so we're using a bunch of computation. It's a way of leveraging a lot more computation.

    Dan Roberts33:57

    On a problem than you would before. So my colleague Noam Brown likes to talk about the Riemann hypothesis a lot. And wouldn't you want to have a model that runs for years that can resolve that, prove that? If you present it and you want it to produce an answer, then it only has the number of FLOPs in a single forward pass to produce one token if it's forced to answer right away. But if it gets to answer after a long time, it can reuse its weights, produce a final answer that is a function of a much larger amount of computation.

    Dan Roberts34:29

    And the natural way it thinks is in language. It's a language model. And so that's sort of this key insight, that you can cause it to do better just by producing a thought process in token space, in language. And this was known before RL, the idea that if you gave a model examples of thinking things out, it would do this before it produced a final answer. Or if you just told it that, then it would do this sort of thing.

    How RL gives models the ability to think

    34:50
    Dan Roberts34:51

    Going back to this SFT versus supervised learning versus reinforcement learning analogy that I gave earlier, there's a lot of examples on the internet of people thinking for a long time. And so it's not completely useless. It can channel that a bit, but RL really brings that out.

    Matt Turck35:10

    What happens during test-time compute? Is RL related or created? Because that's effectively what you described earlier when you were defining RL. The model goes in one direction, decides maybe that's not a fruitful one, backtracks, tries something else. Is that correct or not?

    Dan Roberts35:35

    I think maybe the result of the RL process is that the model can then think at test time. And that's why we have these dials, or various companies have reasoning-effort dials, right? So you've now created a model that will produce a bunch of tokens before it outputs a final answer. Causing that to be good is what RL is doing, or one of the things RL is doing. And so the output of doing RL training is the ability to have a model that thinks.

    Verifiable rewards, math, coding, and the messy real world

    35:40
    Matt Turck36:04

    One of the key questions in the field is whether you can expand and generalize the success that LLM systems have had, particularly in coding and now math, to domains where you can sort of verify whether what the model comes up with is correct or not. What is your view on that? And perhaps start by explaining what a verifiable reward is.

    Dan Roberts36:34

    So a verifiable reward is, in principle, a reward that can't be hacked. So if it's a math problem and the answer is an integer, you just string-match the integer, and then you verify that it solved the problem correctly. That abstraction has all sorts of problems with it, but a problem that can't be verified is: Is this a good piece of creative writing? There's not something you can sort of string-match against. That involves questions of taste, and maybe different people answer differently.

    Dan Roberts36:47

    So maybe it's a distributional kind of thing. And so there's clearly a big gap between those two things.

    Matt Turck37:01

    So do you think there is a path for RL to be truly effective at domains without verifiable rewards? Consulting, banking, legal. Clearly there's tremendous progress in those domains, but what is happening?

    Dan Roberts37:09

    I definitely think OpenAI will have amazing products that will be relevant in those domains, and some amount of RL will play a role in there.

    Matt Turck37:23

    Does RL generalize? Meaning that as you train it against more and more domains, it becomes disproportionately good at learning the next domain?

    Dan Roberts37:41

    I mean, we want to make a model that is generally intelligent and push that intelligence as far as possible. And to do that, we want to make everything part of the distribution. And then we also want to make it robust in cases where it encounters things that were not in the distribution. But I think there's a vague sense, as I was trying to say earlier, that there's a lot of things that are very fuzzy. But clearly, the question of generalization in AI is an important central one, and there's a bunch of examples, I think, that support that the processes can do this.

    What physics can teach us about AI

    38:00
    Matt Turck38:32

    So going back to your physics roots, a lot of what we just described about this interplay between pre-training and RL and all the various bits that we described, those are clearly pretty complex systems. You were trained in a discipline that is all about studying complex systems. What can physics teach us about how to understand those AI systems that we're currently building?

    Dan Roberts38:53

    I think there's a lot of angles to answering that question. I think the maybe most interesting one, or the most relevant one to how we work currently, and maybe this is a contrarian take, is that the way to think about scaling and scaling laws is not small to big, but big to small. I'll get to why physics really matters for this in a second. When you have the existence of some really big AI system and some weird things happen and they didn't happen at the small scale, and so we say, oh, this whatever emerged at scale.

    Dan Roberts39:28

    Sometimes people use the word grokking. There's something discontinuous about the scaling sequence, or the scaling law is broken. These are things that people might say, but I think I reject that entirely. I think it means that you didn't understand something about what you were scaling up. Maybe even going back to the reasoning thing, I don't know if this is true. This is a cartoon. I wasn't at OpenAI at the time, but if you imagine trying to get small models to reason, GPT-1, GPT-2, GPT-3, and then GPT-4 did that again, cartoon.

    Dan Roberts40:07

    You might say, oh, this emerged at scale and it doesn't happen for the small models. I reject that. Instead, there's some phenomenon that's really exciting that we discovered, like reasoning, or maybe something bad, like your model blew up and your earlier models didn't blow up. And your job is to then figure out how to restore smoothness to the scaling sequence, go back and make smaller and simpler models or simpler toy examples such that the whole thing is smooth. If you can do that, if you can figure out what to put into the small thing, then you understand the thing and then you can move forward.

    Dan Roberts40:41

    This is exactly what we do in theoretical physics. There's the Standard Model, which is—I have a textbook behind me—the description of all the forces except gravity would take, even in compact notation, the entire page. It's completely gross. There's a lot of different particles. Why? Who knows? Some of them there's reasons for, but they're doing all sorts of different things. Different things cancel, whatever. Or this just happens to be the universe that we live in.

    Dan Roberts41:11

    But you don't need all of that to study pieces of it. To study electromagnetism, you forget about everything else. Or if you want to study the Higgs phenomenon, which gives mass to some particles, you can study a simplified version of that. And so what we do, and I think one of the key moves, at least in my training in physics, is to take really complicated systems. This often gets talked about as physicists just study spherical cows. And I think that kind of misses the point.

    Dan Roberts41:39

    If the spherical cow is sufficient to describe the thing that you care about, then you did a good job. And if not, you did a bad job. You don't try to retreat to a setting that's simple enough where you can calculate something. You try to retreat to the setting that's simple enough that contains the thing that you care about. And then you have no idea whether you can make progress there or not. But once you did, you sort of understand what the problem is.

    Dan Roberts41:58

    And that's a lot of the work in physics. And the same thing is true in AI. You have these crazy huge systems that have all sorts of interesting phenomena. And if you think about it the right way, they don't grok. There's just this nice continuity.

    Is there a thermodynamics of AI?

    42:08
    Matt Turck42:08

    Do you think there could be an equivalent in AI to thermodynamics, meaning a compact theory that predicts behavior without tracking every individual bit?

    Dan Roberts42:30

    Yeah, Kaplan-McCandlish scaling, OpenAI scaling laws work originally, is a version of this where you throw away all you know about the network is how many parameters it has and how much data you've trained it on, and you can predict the final loss. I think the missing piece is going from all the individual weights and biases and how does that add up to the scaling law. I have some very initial work, and there's some other initial work about trying to bridge that connection, but I think that's the missing piece, the statistical mechanics to thermodynamics of how do these things emerge.

    From Erdős problems to Einstein-level AI

    43:08
    Dan Roberts43:08

    But there's definitely a lot of useful effective descriptions of how these systems behave. I think the other part of your question is: is it enough to characterize everything that we care about? There's probably a lot that we care about other than just the final loss function. And so there's more thermodynamics to be worked out, in addition to how the thermodynamics arise from the microscopic description.

    Matt Turck43:26

    So at that conference a year ago, you jokingly predicted nine years to Einstein-level AI. Where do you think, all jokes aside, we are on that spectrum of just AI creating scientific discovery? That's where we started the conversation, and I'm curious about where this is going.

    Dan Roberts43:56

    Maybe it's helpful to deconstruct a joke, as it always is, but the joke was taking the doubling time for the amount of work a system can do autonomously and figuring out how long it would take us to get to a system that can think eight years on its own, because Einstein spent eight years discovering general relativity. And I projected that out, and it was nine years from last year. I hate making predictions, but I'm pretty sure something will break before that. I mean, in general, we're not just going to set up a system and let it think autonomously for eight years.

    Dan Roberts44:26

    If anything, because the systems eight years after will be so much more powerful, it probably doesn't make sense to let a system think for a certain amount. There's an amount of time it takes for the system to improve, and then there's the amount of time it's thinking. Probably when those cross, all these scaling laws are going to break in certain ways. I do think that the kind of thing that I was trying to talk about, about how we as physicists approach problems, that the structure and flavor of that is maybe different than, here's a very well-defined thing, and go and do a calculation, which is what these Erdős problems are.

    Dan Roberts45:11

    I think probably we'll need to have some ideas to bridge from one to the other. It's not obvious whether it has to be a discontinuous thing or a smooth thing, but there's part of the scientific process, I think, that the models haven't been imbued with yet. I'm sure people are thinking about how to do that, like trying to get to what is the right question as opposed to, here's a well-defined thing and go calculate. Some of that involves research taste. That's not an easily verifiable thing.

    Is AI already doing original science?

    45:16
    Matt Turck45:22

    Is that what would convince you that AI is doing genuine original science?

    How far are we from AI automating AI research?

    45:51
    Dan Roberts45:51

    No, I'm convinced. And I think the unit distance problem is a great example. And also just being able to take a position that is contrary and think for an extremely long amount of time, explore lots of different options, and bring to bear the full weight of disparate fields where it's very unlikely to find a human that has the exact set of skills to solve some of these problems. That's a huge thing.

    Matt Turck46:02

    How far do you think we are from AI research actually automating itself? Not just AI researchers using AI, but AI autonomously building AI?

    Dan Roberts46:23

    I think it's again one of these smooth things where it's already doing pieces of it now. It'll do more in the future. I know there's strong versions of this that people like to think about, but I'm not sure that we'll see a really sharp phase transition versus just more and more pieces. Right now, a lot of coding that would take people weeks can be done very efficiently with models. So some of these math discovery problems, there's also versions of this where, for engineering, the models are playing a more central role.

    Dan Roberts46:54

    And so I think there'll just be more of that. I think that there's a kind of scientific thinking that humans still seem to be very useful for doing. And I don't want to make specific predictions about when or how. I can imagine you don't want to be caught on record saying the models won't be good at something because you'll definitely be wrong. Or maybe I should say that, and then the models will be good at that immediately. And so I should pick the things that I want the models to do and say that they'll never do that.

    Dan Roberts47:22

    I think it's also just hard to make predictions because I think the way in which people made predictions before, like the actual ways things shook out, often are not in that direction. And so it's another sort of credit assignment thing. Like, if you have this long chain of things that has to happen for whatever to happen, then anything that breaks that chain means your prediction is just way off. I can make a very long-distance prediction for the next six months.

    Why Dan is excited about the future of science

    47:41
    Dan Roberts47:46

    Like, I think we'll see more of these sorts of math and science breakthroughs, and obviously we'll turn this sort of thing on AI itself, and the models will get a lot more powerful, and that'll be fun. You could think about that. You could do science of AI and have it feel like doing physics, and that's true. Another really exciting thing is that I entered physics thinking that I would—when you first start learning a field and maybe you want to commit to it, at least the perspective I had was that, oh, by the time I get to the end, I'll know all the answers.

    Dan Roberts48:23

    All the fundamental questions—obviously, this is a journey, and at the end of the journey, it'll resolve. And then, I don't know, maybe it was in grad school or maybe when I switched to AI, I realized, oh, some of these questions will stay open maybe forever. Maybe I'll never get to learn the answers. Watching older colleagues as well start to retire and realize that they might not get to learn the answers. But I feel really excited that we will get to really answer a lot of fundamental questions in the fields of science that we care about, with the aid—or maybe the models being the driving force.

    Dan Roberts48:35

    And so that's just really thrilling.

    Matt Turck48:43

    Well, that feels like a wonderful place to leave it, Dan. You gave us plenty to ponder. Really appreciate you spending time with us today. Thank you.

    Dan Roberts48:45

    Thanks for inviting me. It was a pleasure.

    Matt Turck49:06

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.