MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek

    Jerry Tworek is the VP of Research at OpenAI. We cover why chain of thought lets models turn hard questions into step-by-step computations, why o1 was mostly a puzzle-solving demonstration while o3 became useful through tools and persistence, and why reinforcement learning needs pre-training but corrects limitations pre-training cannot resolve.

    10/16/2025

    Hosted by Matt Turck · with Jerry Tworek, VP of Research, OpenAI

    reasoning modelsreinforcement learningchain of thoughtOpenAI researchAI agents
    Listen now
    YouTubeApple PodcastsSpotify
    1h 16m · 25 chapters
    Contents

    Transcript

    What Reasoning Actually Means in AI

    1:01
    Matt Turck0:58

    Hey, Jerry, welcome.

    Jerry Tworek1:00

    Hello. Very happy to be here.

    Matt Turck1:17

    We are going to talk about reasoning a lot in this conversation. At a high level, what does reasoning actually mean? When we talk to ChatGPT and ChatGPT says it's thinking, what actually is happening behind the scenes?

    Jerry Tworek1:44

    I think that the thinking process is at least a good analogy, as we, in the early days of AI, always had this goal, dream of trying to teach models to reason. We were thinking about it spending more time to get better results. If a human is posed with a very hard problem in front of them, very rarely do they have answers straight away. Sometimes they need to find that answer. Sometimes they need to perform certain computations. Sometimes they need to look up some information.

    Jerry Tworek2:11

    Sometimes they need to teach themselves something. And the process is reasoning, is getting to an answer that you don't yet know. In some way, it can be called search, but it's not really a very naive search. Search is a loaded word, but reasoning is the process of getting to an answer and the work that you need to do that is longer than what usually is considered answering a question. I think the difference is here: answering a question usually means you already know the answer and you just elicit the answer you know.

    Chain of Thought: Models Thinking in Words

    2:32
    Jerry Tworek2:32

    And the process of reasoning is getting to the answer that you don't know. And usually, the longer you spend on getting to this answer, for whatever you need to do to get there, the better it gets.

    Matt Turck2:59

    And we've all become familiar, since you guys released o1, I guess a little over a year ago now, in September 2024, with the concept of chain of thought, which is, in layman's terms, the little messages that you see when you query ChatGPT and it tells you, it shows its work, it tells you what it does. What does that actually do? Is that a logical tree and it eliminates option after option? What actually happens?

    Jerry Tworek3:34

    What language models do on their own fundamental level is they are often called next-token prediction machines. And that's not completely accurate in the age of reinforcement learning, but they still operate mostly on tokens that are mostly text. The language models, again, are these days also multimodal, and they operate on mostly text. But to simplify a little bit for a second, language models generate text. And what chain of thought is, is their thinking process verbalized using human words and human concepts. So the magic that we are seeing, why this is all possible, is that while you are training on all of the internet, on a lot of human knowledge and human thinking process, the model starts learning, in some ways, to think how humans do and, in some ways, get to the answers how humans do from seeing humans do it a lot in the text that was pre-generated and that was based in training data from humans.

    Jerry Tworek4:23

    And then chain of thought is basically eliciting that capability in language models of thinking and getting to an answer like humans do. A lot of what early chain-of-thought work was doing was solving math puzzles. And the first most famous prompt to elicit chain of thought in language models was so-called, "Let's solve it step by step." There is this very classical result in language models that if you ask them what is some either mathematical expression or some puzzle, then they will try to give you an answer.

    Jerry Tworek4:54

    They will try to predict the next token, but they fail. It's a hard thing. They can't compute it in a one-token jump. But if you ask them, "Please do it step by step," they will start thinking, okay, I don't know the answer, but the first step of getting to the answer is this. And then they write a chain of thought, which is a series of text, a series of tokens doing the first part of the computation, the second part of the computation, the last part of the computation.

    Jerry Tworek5:24

    Then they connect those things, and then they can get to the answer. So the chain of thought is basically a process of thinking encoded in words, how humans would solve a problem on a piece of paper, going step by step from start to the end.

    How Models Decide Thinking Time

    5:25
    Matt Turck5:46

    And since time, and by that I mean the time spent thinking, is so important to that concept of reasoning, how does the model decide how long to think when we're in GPT-5 and we're in Auto mode and it says that it's going to decide automatically how long to think? What happens there?

    Jerry Tworek6:14

    It's basically part of our optimization process, partially for the happiness of the users and what they want to expect. Because when you have a thinking process, you need to balance two things, which is the quality of the result. As we said, and there have been those pretty great scaling laws that we demonstrated with the release of o1, the longer the model thinks, the better result you get. But also, people don't like waiting. Waiting is time lost that you could do something with.

    Jerry Tworek6:42

    Everyone wants to get results as quickly as possible. And there is this saying: you can get cheap, fast, or good, and you can take two. And that applies to language models as well. There is a trade-off, and it's delicate. That's why we also expose some of that trade-off to the users, where you can have a high-reasoning model and low-reasoning models. And this is, in the end, the same model. We just tweak the parameter which says we want you to think longer or shorter.

    Jerry Tworek7:08

    We try to encode some heuristics of what we think the users will want, when thinking on an answer a little bit longer and getting to a better answer is worth waiting for versus not. But it's a bit of trying to guess the anticipation of the users. What's the right amount of thinking for them in this particular situation?

    Matt Turck7:14

    Fascinating. So it's more user-driven. So it's more like a user experience kind of thing.

    Jerry Tworek7:21

    In the end, it is, because the question is, like, how long do you want to wait for an answer? You can always wait longer and get an even better answer.

    Evolution from O1 to O3 to GPT-5

    7:24
    Matt Turck7:48

    It's been a little over a year since the release of the world's first reasoning model, o1, which is an effort that you led. What has been the journey since? So there was o1, then there was o3, then most recently GPT-5. How would you characterize the evolution of reasoning specifically across those three models in the last year?

    Jerry Tworek8:16

    In some way, how I characterize our reasoning or scaling-up reinforcement learning research program is we do a series of scale-up runs that are progressively more and more ambitious. Every one, we try to do something more, something larger scale, something that should result in a better-trained model than the last one. And obviously, we don't release all the models that we train. Some we release, some we think need to wait a little bit longer for the moment where they will have their time to shine in the hands of the users.

    Jerry Tworek8:55

    But o1 was the first model we decided to release as kind of a demonstration to the world that there are those models. And o1, to be perfectly honest, it was really mostly good at solving puzzles and maybe a few kind of thinking problems here and there, but it wasn't yet a very useful model. It was almost more like a technology demonstration than actually a really polished product. But we were thinking, we have something cool, and we wanted to share it with the world as OpenAI.

    Jerry Tworek9:31

    o3, I think, changed that pretty significantly. In some way, it is a model that is meaningfully useful. A little bit self-serving, but it was the moment when I started using ChatGPT quite a bit. And I'm basically a user completely hooked on reasoning models. In ChatGPT right now, I use basically exclusively reasoning models because those are the only models whose output and results I trust. And I think o3, its ability to use tools and get to an answer, leveraging a lot of contextual information from various sources and persevering towards getting to that, has really been something.

    Jerry Tworek10:01

    I think there was a little bit of a tectonic shift in the trajectory of AI. And I think we did something really, really great there. o1, it's a little bit of an iteration of the same thing and the same concept. And what I and my team are after right now is something next that would be the next pretty significant jump of how we interact with models that are even more capable, thinking even longer and interacting with even more systems and sources of information on their own journey.

    Jerry Tworek10:51

    But separately, in the meantime, we continue to build a lot of things on top of o3 technology, like Codex, which I think coding agents are, at the moment, the first pretty successful agentic products built on top of AI. There are things like computer-using agent, it's called ChatGPT agent right now, I think, and Deep Research, and a few other things that we will keep on building on o3-generation technology.

    Before OpenAI: Growing up in Poland, Dropping out of School, Trading

    11:00
    Matt Turck11:22

    Great. All right. So we're going to go into all of this in much greater detail in a minute. But before we do that, let's talk about your journey. I think it's a super fascinating topic for all of us. You guys are changing the world. So I think I'm curious, and we're all curious, I think, about the people, the human aspect of who those people are that are just having such an impact. So, starting from the beginning, you grew up in Poland, I believe, right?

    Jerry Tworek11:29

    Yes, I grew up in Poland.

    Matt Turck11:35

    Walk us through your formative years and how you got started in this field.

    Jerry Tworek12:01

    Yeah, happy to do that. An interesting fact: it's almost like a crystal starts from something, and you put a little bit of something in the beginning. There's, I think, one part that was important and part of the starting point of my journey where I didn't know where it came from, because it was there with me from the very beginning of my life, in a moment that I don't really know when it started. It just was always there with me.

    Jerry Tworek12:27

    I always thought that being a scientist and doing science is the highest calling a human can have. And I don't really know where it came from. My parents maybe were singing the right lullabies to me when I was one or something like that. But basically, since I can remember, I wanted to be a scientist. In the early years, I also discovered I have talent for those things. I was going to school, and I saw I got things slightly faster than people around me, at least in a regular school in the middle of Poland, which made me kind of like doing those things, like studying maths and science, a little bit more because it felt good in a way.

    Matt Turck12:49

    Mm-hmm.

    Jerry Tworek13:13

    It felt like this is something that naturally fit me. And I grew up as a very regular kid, just being a slightly nerdy guy and trying to balance my side of being interested in science, programming, maths, and having some social life. And I definitely had some kind of party arc in my life. But I think that the most important part and moment was when I actually went to university, college. University of Warsaw is where I went.

    Jerry Tworek13:44

    And I decided in the end to study mathematics. At around that time of being 18, my idea of life was to be a mathematician with a pencil, sitting in a room with a piece of paper and solving equations. This is kind of my 18-year-old dream of how life should be lived and what I want to do in my life. And my personality is built, again, in a way of really appreciating solid science, pursuit of truth, great engineering, and all those aspects.

    Jerry Tworek14:19

    But I definitely have also a little bit of a misfit, kind of rebellious tinge to it. And that resulted, after a few years of studying mathematics, in what I realized about myself and about the world: I really like maths and I'm quite good at it, but I didn't like academia that much. And I realized I don't want to stay in academia. I don't want to stay in university, and that this would not be an environment where I thought I would be long-term very happy and fit.

    Jerry Tworek14:54

    It felt a little bit too rigid, a little bit too structured in a way that I kind of didn't know if I would feel good. And in some way, for young me, I was around 21 years old at that moment, that was a pretty big crisis of faith for me. I had a moment of lost purpose in life. So I just did a very simple first-principles thinking. I am graduating with a degree in mathematics.

    Jerry Tworek15:28

    I need to get a job to get food. And what job can I do to use mathematics in that job? Looking at the job market, that moment was 2011, I think, or 2010, somewhere around that. I decided to become a trader and trade for a living as the one way where I can do what I like, which is mathematics, and get a career. I got a quick internship at J.P. Morgan investment bank, on the trading floor in the equity derivatives group, spent six months there learning a little bit how trading works and what it looks like.

    Jerry Tworek16:05

    Finished my degree. I got a message from the boss of my boss at J.P. Morgan saying, "Hey, Jerry, you were one of our best interns ever that we had. We really, really liked you working with us. And we are leaving the bank and starting a new hedge fund. Would you want to come with us?" And for 20-, 21-, or 22-year-old Jerry, that sounded like kind of a cool adventure-type story that I was interested in going there and doing.

    Jerry Tworek16:46

    It had enough interesting problems to be solved. And at the same time, it had this kind of trying something new, trying something ambitious kind of bet that I generally like. So I was in London. That company didn't really work out, unfortunately, but it was hard and ambitious, and not everything works out. I did try that again, starting another hedge fund from scratch with a few other people in Amsterdam. I worked there for a few more years, and eventually, eventually I got bored.

    Jerry Tworek17:23

    Generally, working in trading is an interesting and exciting problem. Market is very hard. The depth of what you can go into trying to understand and model is very deep. And I worked with pretty smart people overall, but I stopped feeling I was growing after a few years of doing that. And at the same time, together with a friend I was working with, we just started chatting about AI and about this artificial intelligence. And what really drew me to artificial intelligence was reinforcement learning, and specifically the DQN agents trained by people at DeepMind in 2013. But I think it was a few years later that I actually learned about those results.

    Jerry Tworek18:09

    From my perspective, and again, this is just how my brain works, the 2012 ImageNet results weren't that significant. During my university years, I learned a bunch about how classical AI—the neural networks weren't very fashionable back then—but I still learned about what they are. I learned about SVMs and all kinds of methods, how you train classifiers. And for me, it was kind of obvious and natural. If you have enough parameters and tweak it hard enough, you will fit a classifier to whatever you want.

    Jerry Tworek18:45

    It was kind of obvious. What was not obvious to me is I never considered classifiers a smart thing. Classifiers: you learn a function on some set of inputs to have some set of outputs. And you can keep training it to approximate better and better. What was something that I missed back then is that when you can fit any function better and better, you can start shaping behaviors and strategies. And when I really saw that was in the DQN results, where they applied the same things that worked in ImageNet.

    Jerry Tworek19:29

    Neural networks—and they weren't particularly big or impressive neural networks—with a classical field of reinforcement learning to solve simple computer games. And it turns out those simple neural networks with a simple learning algorithm started learning pretty complex computer games and exhibiting very interesting behaviors. I saw those behaviors, I saw those results, and I was like, "This is what I want to do for the rest of my life," which is not a very long horizon, what are 20-something things about? But I was like, "This is what I want to do."

    Jerry Tworek19:49

    Where do I do that? Google search: where are places where you can do reinforcement learning in this world? Google DeepMind and OpenAI came up with this kind of, at that moment, pretty small and somewhat known, but they were—

    Matt Turck19:53

    Yeah. You joined OpenAI in 2019, right? So very much—

    Jerry Tworek19:53

    Yes.

    Matt Turck20:01

    Very much in the early days still, very much in the kind of nonprofit era of OpenAI. So how did you connect with them?

    Jerry Tworek20:24

    I just applied through the website: OpenAI.com/jobs. Apply, send resume, and hope they respond. And luckily enough, they did. I don't know how many resumes OpenAI was getting at that time. I think it was definitely much less than today, but I came there and I was like, it doesn't matter what I do as long as it's reinforcement learning.

    Working on Robotics and Rubik's Cube Solving

    20:32
    Matt Turck20:59

    So you joined in 2019 with a passion for reinforcement learning. Was that around the Dota 2 moment? Because OpenAI, interestingly, in those early days of 2019, did a lot of reinforcement learning-focused work, right? And then there was a whole unsupervised learning GPT moment that happened afterwards, but it started from roots in reinforcement learning. So did you work on that project specifically, or was it too advanced by the time you showed up?

    Jerry Tworek21:27

    So the project that I worked on was the robotics project at OpenAI, which shared the same code and same methods as the Dota project. On the one hand, the Dota project was OpenAI's way to demonstrate to the world what scaling up reinforcement learning can do. And in some ways, it was taking the 2013 DQN agents and just doing all the hard work of making it bigger and bigger and solving harder and harder problems. And OpenAI generally, from the very beginning, was aware—and really, it was a simple but genius insight—that you need to have a large-scale system to learn really interesting, complex behaviors.

    Jerry Tworek22:11

    And that was one way of what Dota was trying to show: that by scaling up reinforcement learning, we can solve pretty complex environments. And then there was another project. There were, I think, three reinforcement learning projects at OpenAI at that time. And the second one was robotics, which was applying the same methods that we now knew, or were proving, could solve pretty complex computer games. Could they solve all the practical problems? OpenAI was always optimistic and ambitious and trying to see: if we can scale RL to solve Dota, can it load my dishwasher?

    Jerry Tworek22:46

    Can it fold my clothes? Can it build a house? And this is what we were doing. The project I was working on was focused on dexterous manipulation, which was back then, and still continues to be, an elusive challenge for trained policies. And we got to a showcase of demonstrating that a hand controlled by a neural network was able to solve a Rubik's Cube, which is a pretty delicate and complex task.

    A Day in the Life: Talking to Researchers

    23:02
    Matt Turck23:12

    So fast forward to today, still in the same vein of the behind the scenes of you all at OpenAI and sort of life there. What's a day in the life of Jerry? What does somebody like you do? You read papers, you train models, you manage teams. What's a day?

    Jerry Tworek23:42

    Yeah, my days are surprisingly uniform, which is I come to the office early in the day after driving my kids to school. Then what do I do all day? I basically talk to other researchers. I talk to other researchers all day, every day. And this is basically exclusively what I do. I take ideas from people, bounce with them, brainstorm with one partner, then move to another one and do the same thing over and over, and iterate. And in that way, keep refining our research program.

    Jerry Tworek24:03

    Sometimes those are group meetings, and group meetings are there as well and have their own team dynamics. But that is basically exclusively what I do. The only thing that changes is the topics of research from meeting to meeting and from person to person.

    How Research Priorities Are Determined

    24:06
    Matt Turck24:19

    How are priorities in research determined? Is that top-down? Is that bottom-up? Do people suggest ideas and others vet them? How does that work?

    Jerry Tworek24:48

    Yeah, the art of structuring, organizing, and leading a research project is something that I generally learned to appreciate very quickly in OpenAI's journey and in my career. If there is something we are good at, it's structuring research projects. And I think it's a unique mix. You can't say it's top-down. You cannot say it's bottom-up. It's a mix of those two, which is balancing all the important aspects. One thing OpenAI embodies and demonstrates is we all work on very few projects total.

    Jerry Tworek25:19

    There are not that many projects. OpenAI is not trying to do everything. We are not trying to have a portfolio. We aren't trying to have multiple different bets. Always, the idea is we do a few core things really, really well and put a lot of effort there, which means there need to be a lot of people working together on the same large-scale, large-ambition project. And we have a few of those, a small number, probably three or four, depending on how you call it.

    Jerry Tworek25:51

    And that's it. And from that perspective, people don't have ultimate freedom. It's not that people come to OpenAI and say, hey, I want to do this, and they just do this, because you need to do something towards the goal of one of those four projects. And then within those projects, we try to be relatively bottom-up in a way, as long as it feeds, again, into those goals. And the most important part of the research lead is to keep making sure all the researchers are working towards this one shared goal and that they don't fracture in their own ways of thinking and doing things.

    Jerry Tworek26:34

    And so it's an incredibly hard thing. It's a very, very hard job, and it's not always easy to see how delicate it is. But that's a lot of what it is. I don't think top-down structuring of research works in research organizations. I really don't believe in it, because you're not hiring some of the smartest people in the world, and OpenAI has incredibly, incredibly smart people, to tell them what to do. They need to figure out what to do, but they cannot figure out, in a whole space of things, what are cool things to do.

    Jerry Tworek26:52

    They need to figure out, from within the space of what the project needs and what could advance the research goals of OpenAI the most.

    Collaboration vs IP Protection at OpenAI

    26:53
    Matt Turck27:22

    And to which you're saying, is there collaboration between the teams working on those 3 or 4 projects at the same time? Because I would imagine, putting myself in your shoes, OpenAI's shoes, there is probably a tension between wanting to be collaborative in general, but equally, I mean, this is probably the most important IP in the world. So you probably want to make sure that not everybody knows everything about everything. Well, perhaps not. I'm speculating here. How do you think about that collaboration versus some protection of IP?

    Jerry Tworek27:55

    You'd be surprised. But the truth is, in research at OpenAI, which is around slightly less than 600 people at the moment, everyone knows everything. Really, they do. And we have always been fully transparent. And it is like, in some way, you are a little bit shooting yourself in the foot if there is a researcher that doesn't at least have the chance to learn about everything, because they don't have the best information to do their job in the best way. And there is some risk of losing IP, but I think the risk of not doing the right thing and of people not being informed about research and not being able to do the best research is much higher, in my opinion.

    Jerry Tworek28:39

    In my personal opinion and how I approach those things. So we are extremely internally transparent within research. And that is one of our operating principles, as the goal is to do the best research we can and train the best models we can consequently. And the culture generally is very collaborative. It is always the case when you have 600 people, when you have groups of people, there's always, like, one person doesn't like the other person because they looked at them weirdly. Or one person thinks the other person smells bad, or just doesn't like their ideas.

    Jerry Tworek29:09

    That does happen. Those are humans and those are human things. But generally, at large scale, I think we really have this belief that we are together in this goal that's larger than every one of us. It is a very positive-sum game because AI seems to be getting only more and more significant, and the success of OpenAI is far from guaranteed. It depends on us doing great work every day. So there's a lot of feeling of shared fate and the fact that we all need to rely on each other to do our job to achieve this shared mission.

    Jerry Tworek29:28

    So I generally think, with all the caveats of human nature getting in the way sometimes, I think on a large scale OpenAI is very collaborative.

    Shipping Fast While Doing Deep Research

    29:32
    Matt Turck29:52

    How do you all manage to keep that pace of releases? It seems to me from the outside that there's a tension, again, between research, which in some ways kind of feels like it could be a long-term kind of thing, and shipping. GPT-3 to GPT-5 in a year. How do you balance all of that? How are you guys able to ship so quickly?

    Jerry Tworek30:25

    I think that the fundamental reason for it is, in general, OpenAI, at least in my worldview, is a generational company, in a way that we have incredible momentum behind us. We know that we were doing pretty great in the past, and we need to continue that. We have incredibly smart people. Literally, the most talented people in the world are coming and want to work at OpenAI right now, which means every output per single person is incredibly high, and every single person does a whole lot.

    Jerry Tworek31:14

    So we have momentum that carries us forward. We have really great people that work together. We have a good operating way of structuring research and can borrow a lot from Silicon Valley on how to get things done quickly. And people are generally very excited about work. Everyone feels the weight and potential of what we are doing, what we are trying to do. And because of that, people at OpenAI have a tendency to work pretty hard. And having great people excited about what they are doing, all working together reasonably well, results in doing a lot of things.

    Jerry Tworek31:35

    We understand that there's only one time in history where AI is being built and deployed and developed. And people want to do it in the best way that is possible.

    Using OpenAI's Own Tools Daily

    31:52
    Matt Turck31:58

    Do you all use a lot of your own tools? I think Fidji Simo was tweeting the other day that, I think, in the latest, what you all announced at Dev Day today, a lot of it was written by Codex. Is that part of the daily experience? Do you use a model to come up with new ideas for models? Do you use Codex to write the code? How does that work?

    Jerry Tworek32:23

    Yeah, we definitely use Codex a lot for coding, and this is only getting better. As I said, I use ChatGPT a lot, although surprisingly not that much for actually coming up with ideas, but for a lot of questions that I have. I think I am a pretty heavy user of ChatGPT right now, happily paying like $200 a month for it. And I think I'm getting—

    Matt Turck32:24

    Oh, they make you pay?

    Jerry Tworek32:36

    For what it's worth, they are making you pay. And I'm pretty okay with this, because then you get pretty generous usage limits and are not really bottlenecked on it.

    Pre-Training Plus RL: The Modern AI Stack

    32:43
    Matt Turck33:05

    Thank you for all of that. Let's switch tacks and go back to how all of this works. So, is the right way to think about modern AI systems at OpenAI—by modern, I mean as of October 2025 versus the old days of nine months ago—is as a combination of pre-training and RL? First of all, is that the right way to think about it? And second, if so, just at a high level, how does the articulation between both of those work?

    Matt Turck33:22

    And then after that, I'd love to do a little bit of a deep dive on RL to make this very educational for folks.

    Jerry Tworek33:51

    Today's language models are basically trained in two stages: first, they are pre-trained, then you do reinforcement learning on them. The reinforcement learning would not work without pre-training. And I think, in a similar way, pre-trained models have a lot of limitations that are very hard to resolve without doing something that looks like reinforcement learning. So I think both of those bits are here to stay. I think the way they are combined and done may and probably will evolve in the future.

    Jerry Tworek34:14

    Nothing should be treated as dogmatic and fixed. And we need to keep generally figuring out the way to train better models. And this is what we are trying to do. The interesting thing—and I can credit that to Ilya Sutskever, how much foresight he had—but whenever I started at OpenAI in early 2019, I remember there was a research all-hands or something like that where Ilya came on stage and talked about what OpenAI's research program was, what we were trying to pursue.

    Jerry Tworek34:58

    And what he said at the beginning of 2019 was to train a large generative model on all the data we can and then do reinforcement learning on it. That was the OpenAI research plan at the beginning of 2019. And this is exactly what we are doing today. The algorithms change, architectures change. I don't think he was even thinking about transformers at that moment. There was some GPT, but it was like a toy example that someone was playing with. But the goal of training a large generative model on all the data in the world and then doing reinforcement learning with it was already there in the core DNA of OpenAI.

    Reinforcement Learning 101: Training Dogs

    35:10
    Jerry Tworek35:13

    And that's what is happening right now.

    Matt Turck35:32

    So let's do, if you will, a little bit of reinforcement learning 101 to make this really interesting to a broader group of people listening to this. So, in very simple terms, explain it to me like I'm 10: What is reinforcement learning?

    Jerry Tworek35:56

    Yeah. Usually, the metaphor and the analogy I have for reinforcement learning is like training a dog. It's very close. And I used to have a dog when I was a teenager. And even what I remember my parents did—I didn't know anything about raising a dog—they kind of got advice through some friend of a friend, a fireman, who I think was working with service dogs. And he came to me and he basically told me a little bit about how do you train your dog.

    Jerry Tworek36:30

    And what most dog owners that are ambitious about training their dogs know, it is always extremely important to have a bag of treats in your pocket. That's what you always do. And whenever you see your dog behave well, what you should be doing, you should smile and you should give your dog a treat. Whenever you see your dog do something wrong, bad, you basically give your attention away, turn away, and become sad. And through the years of breeding, the dogs discover it's a bad reward and bad behavior.

    Jerry Tworek37:03

    And this is exactly doing that, but with models. We elicit a lot of different behaviors in the models, put them in challenging situations, and then we give them a cookie if they do something we want, if they do a good thing, and give them some kind of punishment and negative reward if they do something that we don't want and that we don't like. In a good way, the good way to do RL is if you balance those things. So if you kind of give cookies half of the time and punish the other half of the time, but this is almost like a mathematical kind of aspect of it.

    Jerry Tworek37:41

    But that's the most important part, which is: elicit behaviors, reward the good ones, and then going forward, the model will be more likely to do what you want and less likely to do what you don't want. And through that, it improves. It is the way how to train models to, like I said, actual behaviors. That is not versus next-token prediction. If you pre-train a model, you literally train the model to predict the next token. RL is like a completely different gradient on a completely different set of what we want to get out of the model.

    Matt Turck38:11

    And getting the model to do what you want, just for some vocabulary and semantics, you hear sometimes the term policy. So in RL, you hear terms like agent, environment, action, reward, and policy. So I think a lot of those are sort of self-explaining. But policy is what? That's a strategy. That's behavior of the model.

    Jerry Tworek38:31

    Yeah, policy is the behavior of the model, as the model weights represent what it does when put in a different setting. Model, in the end, is a mathematical object, and you can define it. And policy is a mathematical function that maps observations to actions: what you see and then what you do with what you see.

    Matt Turck38:45

    Yeah. So agent is a model, action is what the model does, reward is how you say whether that's good or bad. Environment—you hear a lot of things these days about designing the right environment for RL. What does that mean?

    Jerry Tworek39:19

    In some way, it is everything that the model sees. But the interesting thing about the difference between RL environments and most other types of what you can call supervised learning or unsupervised learning is that reinforcement learning environments, you want them to be interactive. You want them to evolve as the model does things. Similarly, if you want to learn how to play guitar, you kind of take a guitar and you strum it. And what happens is you hear the sound of that, and then you hear it, and then you can learn to play with actual feedback of what is happening with the guitar.

    Jerry Tworek39:55

    And in a similar way, the environment is like, how does the world react to your actions? And a lot of what drives your actions is what is happening in your environment and in your world. And that's kind of the only way how to really teach agents to learn to react to changes in the environment, is through reinforcement learning.

    Matt Turck40:09

    Can you give us a little bit of a bird's-eye view of the evolution of RL over the years? Mostly, how does modern RL differ from historical RL?

    The Evolution of Deep Reinforcement Learning

    40:17
    Jerry Tworek40:35

    Yeah. Again, there was sort of historical RL. It's not even that old, but the main tectonic shift was combining neural networks with reinforcement learning. Reinforcement learning predates neural networks as a general mathematical method of optimizing behaviors in mathematically defined environments and as a method of study.

    Matt Turck40:37

    That's what is known as deep reinforcement learning. Is that right?

    Jerry Tworek41:12

    Yes. And then deep reinforcement learning, that was basically DeepMind's invention of combining neural networks with reinforcement learning, the DQN moment I talked to you about. And then from there, there was a moment where reinforcement learning on games was a pretty active area of research. Even when I started in 2019, reinforcement learning was kind of fashionable at that moment, although not very successful. But reinforcement learning was able to solve a lot of games. The bottleneck was that the models were not pre-trained in any way.

    Jerry Tworek41:48

    We were training a lot of behaviors playing games. We even got the AlphaGo moment out of that, which a lot of people got very excited about. But it was still learning behaviors without models that were meaningfully smart about those behaviors. There was still a lot of, you don't want to call it caveman intelligence, but something in that regard, about models not being really smart even though being pretty heavily reinforced. And there was a long research history in that, and a lot of cool results and theoretical understanding of RL comes from those days because people were researching RL actively.

    When GPT-4 Seemed Underwhelming at First

    42:09
    Jerry Tworek42:28

    But it was, in some way, a dead end of doing RL without pre-training. And then, when I finished working on robotics, I started working on teaching language models to code. But having pre-trained models was a really big deal. And then the GPT era of scaling and of large-scale ingesting lots of data to really train great models enabled us, already at that moment, to start RL. And that was one of the first things I did almost immediately.

    Jerry Tworek42:52

    Whenever GPT-3 was trained, I tried to do RL on it. And there always were bottlenecks. The systems were kind of clunky. It was hard to figure out what are the right algorithms, what are the right problems to work on, and what is the right algorithm to train it on. And what OpenAI did at that moment, and what was kind of how research goes, we kind of cargo-culted a lot of things that were used for games and almost the same things as for robotics.

    Jerry Tworek43:36

    And the first RL I was doing on large language models was kind of the same PPO we used for everything. And it gave some results, but those early results weren't completely mind-blowing in RL. And there was a long time where we kept on investing in it. And personally, I always believed there would be a really, really big moment for RL and language models. But definitely, the early trials and errors weren't super successful. The moment when we trained GPT-4, there was an interesting moment.

    Jerry Tworek44:09

    Where we trained GPT-4, and everyone today thinks, oh, GPT-4 is such a great model. But when we trained GPT-4, we were pretty underwhelmed internally. And there were a lot of moments of, oh, we trained this model, we spent a lot of money on it, and it's pretty dumb. At least we have GPT-3. GPT-3 already does all that stuff, and GPT-4 doesn't really seem to be that much better. And we had this question. It seemed smart on evals that were one token long.

    Jerry Tworek44:40

    It seemed to be able to give pretty detailed answers to complex questions when it was one token. But if you actually let it speak for longer, it wasn't very coherent or really gave a very long answer. And we needed to answer this question: how do we actually make the language model that seems to have some smartness in its way actually sound smart and actually be good in talking to it? And that was the moment where a technique that was developed already a few years earlier really shone, which was called RLHF, which is basically doing PPO on large language models with the reward given from human preferences of seeing two parts of that text.

    Matt Turck45:01

    Thumbs up and thumbs down.

    Jerry Tworek45:23

    Yeah, thumbs up, thumbs down, whatever human preferences is. And that's a very good reward because there are a lot of ways the model can generate bad text, and early GPT-4 was generating bad text in a lot of ways. And RLHF was able to catch those things and correct it and reinforce good behaviors, reinforce generating good text, and then punish bad text. And in the end, GPT-4 plus RLHF together as a package delivered that ChatGPT moment to the world that everyone sees.

    Jerry Tworek45:38

    And as much as it is a big success of pre-training, it actually also was a pretty big success of RL in the form of RLHF. Amazing.

    How RLHF Made GPT-4 Actually Useful

    45:39
    Matt Turck45:54

    And just to double-click on that, the RLHF. So we're all familiar as users with, I mentioned thumbs up and thumbs down, that's on the interface, but the actual RLHF happened post-training. Is that right?

    Jerry Tworek45:54

    Yes.

    Matt Turck46:05

    And what did that look like as an effort? Did you have, like, a bunch of humans sitting down in front of the model, industry specialists maybe, and giving it feedback?

    Jerry Tworek46:36

    So actually, RLHF was a research program that was already happening in the background for a while. I think we did RLHF, at least I remember GPT-2 being RLHF'd for quite a bit. That was already there and already happening. It's gathering data even for RLHF in its own research domain, basically, and always thinking: what is the right data to train the model, what is the right data to train your rewards, and how to shape your rewards. It's research that we've been doing, and it's very open-ended and very deep in many different ways.

    Jerry Tworek47:08

    I think there are papers written on what RLHF is, but there's a lot of depth to it. But long story short, you have what we call AI trainers these days, and they look at outputs of the models and they give them scores, and then you learn, basically, a model of those scores and use that for training. And that's part of, for people who may be curious, the entire data labeling industry.

    Matt Turck47:15

    So Scale AI and a bunch of others, that's what they do, right?

    Jerry Tworek47:29

    Yes. I think, in a way, it's getting more and more to be a thing of the past as the models are getting smarter and smarter. This is becoming less of a thing. But I think a few years back, and especially in GPT-4 days, this was the thing. The interesting bit about the data labeling industry, I'm not sure how much we want to go on that tangent, is that it has to constantly reinvent itself because the AIs are getting smarter at some point.

    Jerry Tworek47:52

    Certain things you don't want to label with humans if AI already can do it. So you move the frontier and you change the type of data you are labeling as you already RLHF'd the previous part.

    Unsupervised vs Supervised Learning

    48:02
    Matt Turck48:17

    We've been talking about RL, but the first phase of all of this is the creation, the pre-training of the models. It is unsupervised learning, right? Do you want to maybe, again, make this broadly interesting to people, define unsupervised versus supervised? And in what way was the pre-training unsupervised versus self-supervised, or whatever nuance?

    Jerry Tworek48:42

    Yeah, I think those are nuances, and I don't think they are as stark and as sharp as some people like to determine them. But pre-training is called unsupervised because, in some definition of it, you don't need any extra labels to the data that you feed into the model. You just feed the text as is. In some way, you may argue that the data is already labeled because it is self-labeled. If you give the model the text and say, "Predict the next part of text," in some way it is a label, but it is self-supervised because we don't clearly tell the model what is right or what is wrong or what we want from it or what we do not want.

    Jerry Tworek49:24

    We want it to just predict the other part of data. You can do the same thing with images. You can mask part of an image and tell the model, "Predict the next bit of image." But whenever there is this classic machine learning notion of targets and labels, I guess we were talking about classifiers. So supervised learning was: you have some notion of targets, what your targets are, and some notion of labels. Supervised learning was like, predict those labels from targets.

    Jerry Tworek49:58

    And this is some type of mapping. But actually, what's interesting is that there are many more bits usually in the targets than in the labels, and studying the structure of targets itself yields much more learning and much more intelligence than learning the mapping itself. So spending the whole compute on just learning the data itself without the labels is the right thing to do, and what is often called representation learning and studying the data and its properties.

    GRPO and How DeepSeek Accelerated US Research

    49:59
    Matt Turck50:14

    Okay, great. All right, so going back to RL, you tweeted the other day, the GRPO release has, in a large way, accelerated the reinforcement learning program of most U.S. research labs. So what is GRPO?

    Jerry Tworek50:33

    Yeah, it was a little bit of a tongue-in-cheek moment. I am extrapolating here a little bit what exactly happened because I haven't been in most U.S. research labs, but I have some mental model of what happened and how. And long story short, GRPO was the open-source release from DeepSeek. And everyone who is terminally online and follows AI discourse knows the DeepSeek moment of whatever it was, when the Chinese company that seems to be doing really, really great work released a new model.

    Jerry Tworek51:16

    And it was also a pre-trained model, a reasoning model. They open-sourced the algorithm. They open-sourced a lot of things they did. Overall, really, really great and technically excellent release. And there was a lot of discourse about how they pre-trained their model particularly cheaply. And that was part of the discussion about that DeepSeek moment. But the other part of the discussion was that they kind of released their reasoning process. That release mostly caught a lot of U.S. labs by surprise.

    Jerry Tworek51:47

    They didn't have a similarly advanced RL research program, to my knowledge—basically no one. And I think the only company in the world that did that, as far as I am aware—there's probably a lot of things I didn't know, but you talk to people sometimes, you hear rumors. So this is my version of the world—is that if you look at the older papers of DeepSeek, that company was doing pretty similar, in some ways, RL research to what we are doing. And I think I have to clarify: what OpenAI is doing is not exactly GRPO.

    Jerry Tworek52:08

    It is slightly different in many different ways, but some parts are definitely similar. And what's most important, those are both large-scale policy gradient algorithms, and DeepSeek was doing research in a slightly adjacent area. They were not very far. And whenever we released O-1 and we told the world that you can get pretty great results with scaling up reinforcement learning on language models, I think it was not a very big hop for DeepSeek to realize, okay, we are not very far from getting similarly good results.

    Jerry Tworek52:48

    And they did it. They trained their reasoning model, and they released it, and they told the world how, pretty much not much later than we released O-1. And I think for a lot of U.S. research labs that didn't yet have a research program for how to train reasoning models, they looked, oh, there's this Chinese company. They released how to do it. It helped us kickstart and train reasoning models much faster than we would have otherwise if we had to find all those bits ourselves.

    What It Takes to Scale Reinforcement Learning

    53:05
    Matt Turck53:30

    What does it take to scale RL? So if there was a phase where OpenAI was very focused on pre-training, and then, if I understand correctly, the last 12, 18 months, or whatever time period, where the emphasis has been on the sort of second part of the plan, which is scaling RL, is that a question of just giving RL more compute, more data, more labeling, as we were saying? What does it take?

    Jerry Tworek54:02

    The first thing that is important to know and understand: RL is hard. Conceptually, if you think about it—and there's still a lot of depth to it—but very conceptually, mathematically speaking, pre-training is dead simple. It is kind of the simplest thing you can do. And there has been a lot of thought and a lot of optimization already put toward that for a few years, of even optimizing and doing very well at very large scale, a very simple mathematical operation. In terms of RL, it is much, much more complex.

    Jerry Tworek54:24

    There's many more things going on in a reinforcement learning run. There's many more things that can go wrong in doing it, especially as you scale up, and many more types of bottlenecks, failures. It's a much more delicate thing, and there is much more room for error. In some way, I don't want to go too deeply into that parallel because it's a little bit overblown, but in some way, just to give some intuition, you can have a steel factory which makes steel, and the process is relatively standardized, and you make blocks of steel, and they are uniform and nice and well-defined, what it is, versus building semiconductors, which there are very, very few companies in the world that can do because there are so many things that can go wrong, and you have to put a lot of attention to details to make great semiconductors, and it's very complex internally.

    Jerry Tworek55:27

    In many ways, this is kind of like—I don't want to diminish, because there is a lot of very hard technical difficulty to do pre-training well at the large scale—but there's just many more moving pieces and many more elements of the reinforcement learning stack that need to be gotten right to get a large-scale run successful.

    Agentic AI and Long-Horizon Thinking

    55:36
    Matt Turck55:46

    You mentioned working on a ChatGPT agent, like agentic AI. Where does that all fit? The tool use, the whole agentic autonomy versus reasoning, RL. Help us reconcile what does what and what impacts what.

    Jerry Tworek56:04

    I think what is important is I believe, and I think, that there can be a lot of positive impact of AI on our world and on our lives through automation, through problem-solving, and through AI doing good things for us, the things that we want. And for a long time—and again, it's not that long, but the last two years or so, or maybe approaching three—we've been living in this world where we kind of ask questions to AI and it gave us an answer.

    Jerry Tworek56:47

    At the beginning, instantly. Now it can think for like a minute or two, which feels long. But in many ways, what can you do for two minutes if you think of how many problems humans solve? And AI is probably a little bit faster in the things it can solve, but it's still a limit of what it can do. There are still a lot of tasks that would take AI much, much longer to do. When I prompt Codex, it works for a while—again, a few minutes.

    Jerry Tworek57:17

    There are a lot of things we have internally and we are doing that allow the models to work for much longer. We still didn't figure out the right product to deploy them, but the models can think for 30 minutes, an hour, two hours these days on certain types of tasks and problems, like even longer than that. They generally are capable of doing so. And we need to figure out how to make that process more useful and more able to actually come to various problems in real life, whatever it is: coding or booking travel or making plans or even designing houses or new electronic devices or whatever else we would like models to do, we would like them eventually to be able to do for us.

    Jerry Tworek57:56

    And a lot of this comes through the models thinking independently for longer periods of time and considering more alternatives and just sometimes going through a slog of very long lists of tasks.

    Matt Turck58:15

    So the agentic part is powered by fundamental reasoning. Is there a concept of, I guess, online RL, where, as the agent does something and learns from the real world, the RL happens in real time?

    Jerry Tworek58:41

    So in general, all of RL is happening. Most of RL that you hear talked about with language models is online, but it's done online in a way that is still a training run. It's still being trained separately from the user. There have been a few models in the world, and I've learned recently that I think Cursor is trying to train some models online with their users in the loop. And it is theoretically possible to train models in ChatGPT or every other product, just responding to the users and reinforced through whatever rewards you get in there.

    Alignment as an RL Problem

    59:19
    Jerry Tworek59:19

    But this is not what I'm aware of, at least not what OpenAI is doing at the moment. And it can be great, but it can also be dangerous because you're not really controlling what you're reinforcing in that loop and what could happen. So, at least until we have really good safeguards, I don't think we should try to do that in anything as complex and large-scale as ChatGPT.

    Matt Turck59:34

    Yeah, interesting. And very much on that note, talking about alignment for a minute, is alignment an RL thing? I mean, do you create alignment in the model by teaching it what is right and wrong?

    Jerry Tworek1:00:01

    It's a little bit of yes and a little bit of no. In a way, alignment is about steering the models to certain behaviors, and that is definitely an RL thing and an RL problem. But also, you want the models to know what is right and what is wrong and understand the world. And not all of them are reasoning and RL problems. Those are very often just AI problems. And, in a way, to be aligned, the model needs to know right or wrong to choose right.

    Jerry Tworek1:00:31

    I don't think you can just tell the model, show it a few good things to do, and it'll do them. The model needs to deeply understand its actions and consequences to really be able to choose the right thing. And I think it's a never-ending pursuit, because even for humans, it is not super easy to define what we consider aligned. And I think as our civilization evolves, the notion of alignment and the goals of humanity will keep evolving, and we'll need to keep nudging the model towards those things and keep explaining to it the things that we want from it.

    Jerry Tworek1:00:56

    But it's definitely a very important and central part of any—should be of any AI research program.

    Winning ICPC World Finals Without Specific Training

    1:01:11
    Matt Turck1:01:26

    Yeah, which brings a whole next series of questions of where RL is efficient versus less. So it seems that it's been particularly good for math and coding. And then the next obvious question is, what about the rest of the world? But taking a quick sort of going down the rabbit hole a little bit about math, so just in September, just a few weeks ago, you guys did something unbelievable with the ICPC World Finals. Do you want to talk about what that was and what went on from a model technical perspective behind the scenes?

    Jerry Tworek1:02:00

    There happened surprisingly little from our perspective, from the model perspective. We just have pretty smart models, and then when we ask them to solve programming problems, they are correct. What's a little bit of a backstory in it is that I think we used specifically programming puzzles for a while as a very nice research test bed of our ideas. Those are nice problems to experiment on, and they weren't ever considered part of the product, but those are pretty complex problems, and they require a whole bunch of thinking, are very nice to give rewards to.

    Jerry Tworek1:02:46

    Do it. So a lot of researchers just liked working on those problems as a way of trying out their RL ideas. You always need a dataset. I'll take a dataset of programming puzzles and try it. And I think because of that, a little bit, our models just were always very, very good at competitive programming as a kind of byproduct. We never tried to be good at it, but researchers were trying their ideas on it. And because of it, every training run, whatever we are doing, just ends up being very, very good at those types of puzzles.

    Jerry Tworek1:03:09

    And then it was a little bit of a formality for us to go and submit that to a competition. It's largely about demonstrating to the world what is the level of capability in those models. It is important and true to acknowledge that in not all domains, at least comparing to human baseline at this moment, we can be as nice and as good as in programming competition problems in many ways because those were tried for a long time by many, many researchers, and researchers don't always spend as much time as they could, and I would like them to, on very practical problems that people go to ChatGPT or our models with.

    Matt Turck1:04:13

    Great. So it sort of came out of the box. There was no specific training for it. And just to remind people, so what I'm referring to, what we were discussing is the ICPC World Finals. That was just in September 2025, which is International Collegiate Programming Contest, ICPC, that happened in Baku, Azerbaijan, where OpenAI solved 12 complex algorithmic problems within a five-hour time limit, basically making it take the equivalent of first place in front of human teams. So just for context.

    Jerry Tworek1:04:47

    We did a little bit of a round tour of various competitions. We did ICPC. We did also IOI, International Olympiad of Informatics, earlier this year. AtCoder Heuristics Competition as well, where we went second behind a single human that is also a Polish person that used to be employed by OpenAI some time ago. Funny coincidence. But I think we were looking for a moment in time where our models are smart enough to be able to compete in those competitions with some incredibly smart and talented humans.

    Jerry Tworek1:05:19

    But it was never our particular goal and focus. It's kind of like we think if we are doing good research for training smart models, they should be smart enough to do those things. And we kind of ticked that milestone. And now we keep on moving forward. And I hope we will see, and I think we are already seeing, more and more practical and tangible things coming out every week or every other week on Twitter. I do see what I think are credible reports of actual scientists using some of our reasoning models to help them perform calculations, solve hard technical problems with our models.

    Jerry Tworek1:05:51

    And I think this is where we want to be. Solving competitions is cool, but people solve competitions to prove that they can go to the actual frontier-level jobs and solve new technical problems. And this is kind of what we want from our models as well.

    Applying RL Beyond Math and Coding

    1:05:53
    Matt Turck1:06:19

    We alluded to this a second ago. So mentally, at least for somebody like me, I understand how RL could be used very effectively to train against math problems or coding problems. I think one of the big questions right now is: how do you do that for the rest of the world, in contexts and disciplines where the answer is not right or wrong, maybe a little more murky? And the rest of the economy. You guys as an organization came up with GDPval the other day, which is a way of evaluating performance against different industries.

    Matt Turck1:06:40

    What is your thinking in terms of generalization of RL as a path to success for the rest of the world?

    Jerry Tworek1:07:04

    I think the short and quick answer is because somehow humans can learn all those things. And as long as there is any way to evaluate performance and figure out if something is going right or wrong, and you can compute that feedback, you need to be able to somehow calculate how well something did, then you can optimize it, and then you can do reinforcement learning with it. I think there can be an argument that if there is no notion of what is right or what is wrong, then also humans are not able to improve and learn, because there needs to be a learning signal that's coming somewhere.

    Jerry Tworek1:07:50

    There is mostly a question of how convenient and how easy it is to get that feedback. And everyone doing reinforcement learning should strive and try to be able to train on more and more complex and interesting training signals. Very often there comes a notion of what is often called reward hacking. And what happens when doing reinforcement learning—it happens a lot, and it's an important problem—you shape your reward in some way to reward certain behaviors. But sometimes it is the case that what you reward is not what you actually want.

    Jerry Tworek1:08:19

    There is one thing: you need to train the model to do the behaviors you reward, but there's also a natural mismatch between the reward that you give the model and what you actually want. And there are sometimes moments where the model does what you reward, but it's not in the spirit of what you had wanted, and you need to fix it. It's almost like a parenting challenge. And in some way, you can say it's a limitation of reinforcement learning, but when I was thinking about it, I realized a lot of that happens in human systems as well.

    Jerry Tworek1:09:01

    There are a lot of incentive systems and reward systems, and it even happens in workplaces, in all kinds of human groups, that humans have rewards that are not always optimized for the ultimate goals of the system, and they hack rewards constantly in many different ways. And there is a constant whack-a-mole game between setting the right rewards and seeing if the system does it. And it's a huge issue in almost any policymaking and any incentive programs. And it's the same kind of whack-a-mole game in reinforcement learning research, trying to make sure your rewards are better and better representing what you actually care about the model doing.

    The Path from Here to AGI

    1:09:15
    Matt Turck1:09:46

    All right, so maybe to zoom out to close this conversation, you said the other day, you tweeted, "We all collectively believe that AGI should have been built yesterday, and the fact that it hasn't yet is mostly because of a simple mistake that needs to be fixed," which is super awesome as a tweet. Do you think that the combination of pre-training and scaled RL takes us to AGI?

    Jerry Tworek1:10:14

    There's always an interesting question of what do we consider something that is not pre-training and RL, and where's the limit? I generally think something that we are doing like pre-training today is necessary. I think something that we are doing with RL today is necessary. And there will surely be a few things more. And we have a lot of very ambitious research programs on some of those things. And I think the question of distance in research space is hard to say.

    Jerry Tworek1:10:47

    For some people, what we want to do and what we are planning to build is not very far from those things. Some will say, "Oh, it's completely different, and it's very much not that." So I don't want to go into debates whether it's the same or not, but we are and want to be constantly changing the way how we train the models to more represent what we think the right form of intelligence is and the most useful form of learning is, and constantly are researching various things.

    Jerry Tworek1:11:12

    And then the distance from, like, what is the distance from AGI, is also a very complex question. I really like—someone said it to me, but I think it is right—that if you talk to someone from 10 years ago and show them ChatGPT from today, they would probably call it AGI. But we're not today because it still has a lot of limitations, and we are all very aware of those limitations, and we are pretty sure we can resolve those limitations.

    Jerry Tworek1:11:51

    There will probably be some further limitations of future models that will need to be fixed. There is an ultimate question which is very hard to answer: when is the moment that the model can improve itself without that much external output and without humans working on it and fixing it? And I think it is a very hard question. It is a serious question that we need to try to answer. Humanity needs to try to answer, because that mind will still largely depend on our infrastructure and our systems, but will be able to start fixing itself without us having to fix it.

    Jerry Tworek1:12:19

    And that's like, the predictions of what really AI will be able to do and will be able to solve at that moment start becoming a little bit murkier than what we can do right now, which I think we still can do pretty well.

    Pure RL vs Language Models

    1:12:23
    Matt Turck1:13:00

    But philosophically, you may have heard Richard Sutton the other day on the Dwarkesh podcast, which is a wonderful episode that people should really listen to, effectively saying that the only path to AGI was going to be pure RL, and that fundamentally LLMs—and maybe I hope I'm characterizing what he said appropriately—but that LLMs were a flawed premise because effectively that was imitation of reality, whereas RL was enforcement of reality. Do you have any thoughts philosophically on that question?

    Jerry Tworek1:13:28

    Yeah, I haven't had a chance to fully listen to that episode yet, so I also don't get the details of that thought. But what I can say is that we are doing quite serious RL on language models these days. And in terms of pure RL, I don't think really pure RL makes sense. RL needs pre-training to be successful. And I think pre-training, as I said before, needs RL to be successful as well. I don't think without RL, the research program we are doing would make sense.

    Jerry Tworek1:14:02

    But OpenAI is, and I'm pretty sure all other AI labs as well, very serious about doing a lot of reinforcement learning on our models. And I think a lot of people are saying that whether LLMs are an on-ramp or off-ramp of AGI, very often they do mean pre-training. But it's also clear that the current way how we are doing things is not yet enough and is not yet everything. And there will need to be further changes to the setup.

    Jerry Tworek1:14:30

    But sometimes people say, oh, if you are doing RL, it's not an LLM, it's something else. Sometimes they say, oh, if you can write a program in your rollout and it's a chain of thought, it's not a neural network only, it's a neuro-symbolic system. So it's easy to get—some people consider something an LLM and the other thing not. But personally, my view is what we have is a pretty good foundation for the next step. We did have transformers first trained for translation, then we were pre-training them on large-scale data, then we were doing RLHF on them.

    Jerry Tworek1:15:20

    Now we are doing large-scale reinforcement learning. If we do a few more and more complex things, there is a chance somewhere along the line, the architecture will start changing more or less significantly. And I personally think we are on the right path, and it will feel less like completely turning around and more like keeping on adding more things and maybe dissolving some old elements that carried us to that particular level of intelligence and were not needed anymore.

    Matt Turck1:15:40

    Well, that feels like a wonderful place to leave it. You've been very generous with your time and thoughts and giving us a glimpse into OpenAI, what you work on, what it looks like behind the scenes, and the key aspects of pre-training and scaling reinforcement learning. So it's been a wonderful conversation. Jerry, thank you so much. Really appreciate it.

    Jerry Tworek1:15:43

    Thank you very much. I enjoyed being here a lot too.

    Matt Turck1:16:04

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.