So I just think people need to realize when to stop that. I don't work with these types of synthetic environments as much, just because I don't work on embodied AI. So I don't know if we're at that yet.
Okay, great. All right. So, going back to pre-training, mid-training, post-training, let's talk about mid-training. It might be something that people have heard about a bit less. The term comes up a bit less. What is it, and why is it important?
Mid-training. It's just this idea of something that's between pre-training and, as you might realize from the name, the post-training part of the pipeline. And really, the idea is if you have high-quality data that is more representative of what you really want in your final model, you should overtrain on that data. So, taking a step back here, pre-training, what is it? Pre-training is basically trying to learn everything from the world by learning everything from the internet at a high level.
The problem is that most things on the internet are not really useful. If you think, for example, about Wikipedia or GitHub, which is coding data, it just seems like there's way more information in there than some random forums that maybe don't have that much information. For example, ads. There's also lots of ads on the internet. You probably don't want to train too much on that. But in pre-training, we train on everything. And in mid-training, we basically overweight this type of high-quality data that we think is more useful for training the final model.
And this is something—I can't talk about what's happening at OpenAI—but it's something that is definitely happening in all the academic community right now, and all the open-source models have this stage of mid-training.
Great. Post-training, let's start at a high level by defining what that is. So, there's reinforcement learning, but that's not the only part of post-training. What else is there?
It kind of depends how you define the term and where you put the boundaries. In my mind, post-training—I'll take it from a very broad sense—includes all the reinforcement learning and the training for our reasoning models. It's just the idea of having something that knows everything about the world and making something that is useful to people. So, pre-training, I think about it—the metaphor that I like giving is you go in the library and you have a lot of books about everything.
And in theory, you can find all the information that you want in the library. But it's much more useful to talk to an expert who has learned these books and who you can ask questions to, and they can answer and understand what you're actually looking for. So this is kind of the goal of post-training at a very high level: making something that is useful to users and is easier to interact with. So there are multiple stages. I'll talk only about things that are happening outside of OpenAI and kind of the usual stages.
There's usually some SFT that is happening.
Which is supervised fine-tuning.
Supervised fine-tuning. And that's actually what, early on, most of the models that were post-trained were only doing: supervised fine-tuning. The idea is that if you have humans that can give you the desired final answer, if you have humans that can give you the gold answer, you can basically clone the behavior of the human. So this is what we call behavior cloning. The problem with this is that you will never get better than what your ground truth gives you. And humans are actually pretty limited in many senses.
So you will never overcome the human labelers that you're working with. The reinforcement learning stage goes from behavior cloning to really optimizing rewards. So the idea is, I don't know what the ground truth is. I don't know what the perfect answer is, but here's how I would say whether the answer is correct or not. And here are the things that I want in the answer. And what you do is you start optimizing. You start having a model that tries to get more reward, basically optimize this reward function. That's what we call it. And it goes beyond what you currently have, what humans can do, or at least what the humans that you're working with can do.
So this, I would say, is the two big stages. Then in reinforcement learning, that depends on which models are being trained. At least in the open-source community, it seems that there are different ways of doing that. Reinforcement learning when you have verifiable rewards, so reinforcement learning where it's really easy to say whether something is correct or not, and you can really kind of have a binary reward for this.
And that goes back to how we talked about o1 and o1-preview in the past. And then you have reinforcement learning without verifiable rewards, where maybe I could do pairwise comparisons. I can say this answer is better than this other one, but I don't really know. I cannot quite say this is the perfect answer. So, of course, it's a continuum and there's everything in between, but I would say these are the three high-level things to think about when you think about post-training in general and how people are usually doing it in the open-source world, is that they take SFT, they clone the behavior that you can collect online or from humans.
And then once it's already at a pretty good level, they just do this reinforcement learning to go beyond what we currently have. Because if you just started from reinforcement learning, it would be very inefficient. Because the problem with reinforcement learning is that you have to stumble across the right answer, basically. Because how reinforcement learning works is you sample many times, essentially, from the model that you're training, and you say, this one is correct, this one is not. And you say, do more of the one that is correct.