Kind of mathematically look very similar.
All right, so you evolved from physics to AI, and then what was the next step?
Had an interim where I started a small startup that went through Y Combinator, and it was basically doing machine learning prediction for inventory management. Really felt like I had to get on the frontier of AI research. I felt that was going to be where a lot of scientific impact accumulates. And so I ended up joining UC Berkeley as a postdoc, where I worked in this lab called the Pieter Abbeel Lab, which is one of these great labs for reinforcement learning and what was called unsupervised learning research, which is basically—large language models and diffusion models are the biggest outputs of unsupervised learning.
I didn't realize that at the time, but that was, in some sense, a—I wouldn't say miracle year, but it was a special year in that lab, given the people who were in it. A large portion of that lab ended up going to start impactful companies or being scientists doing very impactful work in large labs. To give you a sense, the first people I worked with were Arvind Srinivas, who now runs Perplexity, and Dennis Yarats, his co-founder. We worked together on some papers there.
Jonathan Ho, who is one of the inventors of diffusion models, was there and invented diffusion—the big paper that made them break through there. And this guy named Ajay Jain, who was on that paper, and they started Ideogram and Genmo, which are two startups in the video gen and image gen space. Deepak Pathak, who is the founder of a company called Skild AI, which is one of the premier robotics companies, was there.
What year was this? What rough time period?
Yeah, it was—now when I look back at it, it was a pretty incredible group of people, like Aditya Grover, who started Inception, which is a company that does diffusion models for coding. I was just thinking, I saw this guy's tweet, his name is Kevin Lu, the other day, who was in the lab, and he was an undergrad then, and most recently was leading a lot of the work for the mini models at OpenAI.
So just an incredible group of people at that time. Not obvious at all then. I mean, the research people were doing was very interesting, but I would never have predicted that so many companies would have come out of there.
And the next step after that was DeepMind.
And then I went to DeepMind. At the time, I was really interested in Toronto and New York. So I joined the group of this researcher named Volodymyr Mnih, who was largely credited with starting the field of deep RL. He was the first author of the Deep Q-Networks paper, which was the paper that got neural networks to play Atari. And his first set of papers actually largely defined deep reinforcement learning as a field then, and were basically DeepMind's claim to fame for a very long time.
And so I joined his group. The team we built together was called the General Agents team. And so the whole point was to do research for us to figure out: how do we build general agents? I think it was much more opaque then than now. And the big problem we were trying to solve is what people called, and still do, unsupervised reinforcement learning, which is really: how do you train reinforcement learning systems that are capable of assigning their own rewards?
If you don't have rewards without supervision, in the same way that you can give things some rewards, but kids and animals, when you look at them, they learn a lot in an unsupervised way. They interact with their environments without—no one is telling them to. And so we were thinking about how do we teach reinforcement learning systems in this way? I actually think the subject is coming back in vogue in the age of language models, with the question of how do you do pre-training-scale reinforcement learning?
How do you generate a lot of synthetic data if you don't have explicit rewards? I think it's actually a really interesting question now again, but that was the general agenda of what I joined to study. And the short of what ended up happening was that language models started working. And once they started working, that really changed, I think, my entire perspective on what problems matter and what didn't. Because a lot of the problems that we thought were fundamental problems were solved in this brute-force way for us.
Gemini 1.5, and then obviously 2 and so forth. And I joined with my co-founder. My co-founder, Ioannis Antonoglou, was leading the reinforcement learning team, the RLHF team. I joined his team and led a lot of the work for training reward models for Gemini and implementing the algorithms and so forth. And it was a very exciting time when I'd say a group of 10 to 20 people were—that was basically—all the people doing the RLHF work there.
And when was the decision to leave and start a company? And what was the thinking?
After Gemini 1.5, we realized that language models crossed this threshold of utility where they're no longer research objects. They're going to be very useful. This was early 2024. We realized that the ingredients were in place to build a superintelligence. We felt that everything was there. There was one more piece to solve of going from RLHF to making reinforcement learning work. And that basically happened over the last year with reasoning models. So we felt that that would happen. Then the question was: superintelligence for what?
That was basically—we felt that you can't answer this question in the abstract by being a researcher that's really far away from product and customers. You really had to go in and define what that means from a product vision and what problem you're trying to solve perspective. It's like, we're not interested in building a superintelligence that will be superintelligence in mathematical Olympiads. And the difference between this era of reinforcement learning and the previous era of pre-training is that when you did pre-training, you made the models generally better at everything.
Reinforcement learning is much more jagged, right? It makes them good at what you want them to be good at. So just because you made them good at competitive code, that improves the general coding capabilities. But that does not mean that you'll have not even a superintelligence, but even just a useful intelligence for software engineering code. And I think an example of that is, I think Anthropic has done a really good job of building models that are meant for users of their products rather than benchmarks.
When I look at academic benchmarks, the Claude models are consistently worse than whatever else is out there, oftentimes not even close. They're consistently worse. And yet, from a user perspective, they're consistently better. Something has to explain that. And I think the explanation is that when you train large language models with reinforcement learning, they become jagged in the sense that they become good at what you wanted them to be good at.
And there are some generalization capabilities, but they're much weaker than people think, which is a little counter to the narrative that you hear a lot, which is that generalization is always going to win. And I guess, I don't know if that's true to it or a bastardization thereof, but like Rich Sutton's Bitter Lesson. And so what you're saying is sort of not the opposite, but that the solution is a combination of generalization and specialization. Is that fair?
Well, I think the Bitter Lesson actually doesn't say anything about generalization. The Bitter Lesson says that the systems that we should be thinking of and building are ones that scale well with search and compute. That's kind of it. And so what he's saying is, if models are limited today—this was actually, basically, the lesson was for researchers, but I think it's for product builders as well—if you're building your product with the assumption that these models are going to stay at their current intelligence level, and you make a bunch of hacks around your product to overcome those things, then in the next iteration of models, a lot of the hacks that you put in place will probably be Bitter-Lessoned.
And that was scientific researchers were seeing these kinds of limitations of models and plugging them in with kind of temporary hacks that would make them better at certain benchmarks. So the same lesson translates there. But the lesson is more to build systems that are good at soaking up compute and scale well with search.
Yeah, we'll put that in the show notes. I think that's probably one of the most often misquoted blog posts in the history of AI.
Well, I think what's happening when we think of generalization, what's really happening is that if your training distribution is everything, then your test distribution just falls in your training distribution, and you have generalization. Maybe one point of view that I don't think that many people share, but I do think there will be a general superintelligence. But I think that it won't be one lab that has built it, but it'll be kind of the plurality, like the collection of all intelligences, will be a general superintelligence.
Because if you think of pre-training as kind of the soil or the substrate from which you can now grow superintelligence in various categories—medical superintelligence, organizational superintelligence, superintelligence for scientists and math—these are all different types of superintelligence. Some of them you can merge into one model. But I think that there'll be, if you think of different superintelligent plants growing from the substrate, yes, like a frontier lab will be able to capture some of them. But I think that there'll be new frontier labs built that build out, sort of grow other plants, and that the collection of this garden is going to be a general superintelligence rather than one company going in and growing all the plants.
I think from a research and compute perspective, that's possible. But from a product perspective, if you want to build organizational superintelligence, you have to go and integrate with all these customers, and you have to have solution engineers that support them, and you have to have salespeople that support them. And in this kind of rosy picture of a researcher who just trains models and hopes the model is a superintelligence, I just don't think that the world will play out that way, because it's meaningless if it's not coupled to a product and its deployment.