Yeah, so you would train a model for this, and then now, because during the training process you can't just pause your training run to ask some humans for input, right? The feedback would have way too much latency. So instead, you need a proxy for what a human would say. So you train this model based on the human preference data, and then you can optimize against it, or at least a little bit.
One of the famous things in the history of RL is Move 37. How do you train a model to encourage the model to do that kind of thing and come up with brand-new ways while being efficient and exploit known paths?
Yeah. So the great thing about Go is that you can just train it. It's a zero-sum, two-player game. You can train it in what's called self-play. It plays itself, and it can go from playing randomly to expert play, and it will find whatever the sort of best strategies are. So if that means exploring, great. If that means exploiting—actually, I have a funny story about this. So I met Noam Brown in grad school. He went to a different grad school than me, but he wanted to enter MIT's poker bot competition.
And he had a poker bot that was the best in the world, but it wasn't something that would compete against humans yet. He just won in this research competition. He collaborated with me and another friend to enter MIT's poker bot competition. This was great, actually, for me because I learned some really exciting work in AI, and I got very excited about this while I was doing physics. We were playing essentially this kind of self-play equilibrium strategy. There's some nuances, but essentially, we could not lose, assuming we did not have any bugs in our code.
The way this thing worked was that it was a tournament where you would be paired with, say, another person and play them in a hand. Depending on the amount of points you got in some sort of round-robin setup, they would eliminate the bottom half and keep going until you got to the final table, which would just be, say, you versus the other person. And so, there was the award ceremony, and we didn't know what happened, but there was someone else who was—what did everyone's scores over time look like?
And there was, say, 64—I think there were 32, actually—people playing, so it was around a 32-person tournament. And 30 people, over time, their scores were all very negative and going down. And then there was one person whose score was pretty much straight up. And then there was another that was pretty good, but not with a crazy slope. And so, do you want to guess which one we were?
So we were the lower slope. And then there was this other guy that had this crazy slope, was just completely crushing all the other players. And then this happened for the round of 16, the round of 8, the round of 4. And then in the round of 2, it's heads-up, us versus this guy who, over the course of this tournament, won way more than us overall, like taken more money from everyone else, and then we crushed him. Because why?
Because he was exploiting the weaknesses of everybody else, right? It had some theory of mind to try to figure out, oh, this guy does this when he bluffs. And so it was very—I assume it was very good at exploiting everyone else, but we were just playing the best possible thing that you could do. So the criteria was not maximize your amount that you get from anyone else. It was don't lose. So it's the best response to anyone's strategy.