And it's all not very well established. I'd say it's a new domain, and I think we are still figuring out how to best do it.
Let's talk about your new paper that literally came out today. So first of all, congratulations. And it talks about epiplexity.
And that's a new word, right? That's a new term entirely. Among other things, you invented a new word in the dictionary. So congratulations on that. So walk us through the whole idea at a high level.
I guess I want to quickly give a shout-out to my collaborators on this work. The lead author is Mark Finzi. Mark is actually currently at OpenAI working on synthetic data there, but we were doing our PhD together, and our PhD advisor is also on the paper, Andrew Wilson. But then also there is Shikai and Yi Ding, who are other students on the paper. And then Zico Kolter, who's a professor at CMU, he's on the board at OpenAI. He's also on the paper.
The core idea is to think about how the data can look different for an observer depending on how much compute the observer has. You can imagine that there is some complicated process that generates the data, and a very smart observer that has a lot of compute can fully understand what that data is, understand every aspect of it. But for a weaker observer that cannot fully model the data, some parts of the data will look like noise. And so the amount of structure that you will see in the data will depend on how much compute you as an observer have.
And actually, in some cases, you can see more structure if you have less compute, which is kind of interesting.
So just to play it back, given a certain amount of compute, you could be feeding the model noisy data—tons of data, but not a lot of interesting stuff in it. Or you could be feeding the model data that has patterns in it and therefore is more interesting to the model because the model can learn more from it.
It's more that even with the same data, it can appear noisy or structured depending on the model. Like, a very big model can extract patterns that a small model cannot.
And that's, from my limited understanding, in opposition to entropy, which is like the amount of noise, I guess, in the data.
In that case, maybe the more relevant comparison is that we are kind of in opposition to Shannon information and Kolmogorov complexity, which are both measures of the information content of the data. They're different, but they share some properties that we think are maybe leading people to have some wrong intuitions, potentially, about synthetic data, for example. So, for example, there is this idea that if you apply any deterministic transformation to any data, you cannot create more information by doing that. You kind of start with some amount of information and then you transform it deterministically.
It will have the same amount of information, both according to Shannon information and roughly according to Kolmogorov complexity. And that kind of leads people to believe that, for training language models, applying transformations or deterministic changes to the data doesn't necessarily lead to more data, doesn't effectively increase the amount of data.
The amount of data or the efficiency of the data?
It doesn't lead to having more information in the data that the model can extract. But we argue that that's just not correct. It would be true if the model has infinite compute. So if the model can fully understand what a deterministic transformation is, then it's not going to be able to extract more information from the transformed data than it used to extract from the original data. But with a limit on the compute, it's actually very possible to apply deterministic transformations to the data.
And create information through that. So we have the example of AlphaGo, actually, or AlphaZero in the paper. AlphaZero doesn't use any human data. From the perspective of Kolmogorov complexity or Shannon information theory, it doesn't create information. So it's unclear what is actually learned by the model because it's trained on no data. It can only learn the rules of the game, and that's the only thing. But from this perspective, because the model is computationally bounded, it cannot do the full rollout of all the possible games of Go or chess and figure out what's the best move in every possible position.
There is structure that is produced through this deterministic process, and the model is able to learn that structure. And so we are trying to reconcile these different observations and come up with a notion of structural information that is dependent on the amount of compute that the observer has.
And the term itself, epiplexity, is that a measure of that?
Yes, it's a kind of novel measure of information content of the data.
And it's going to be a number on a scale. How does that manifest?
Yeah, it is a number. We can measure it. And it's not easy to measure. So it's kind of a theoretical definition that we prove some things in the paper about the properties of this measure. But yeah, we also do measure it. So, for example, we can approximate it from the scaling laws, and we can, for example, say that text data has more structural information according to this measure than image data at the same amount of tokens.
And as this whole line of research develops, what is the likely impact on industry? Does that mean that we may need comparatively less compute because we know what data to use? What may happen?
In my mind, the main impact is conceptual. For example, for me, I'm now very interested in completely synthetic data, just data that's generated by some computation. You can define some arbitrary programs, use them to produce infinite data. And as we run out of the internet, maybe we eventually want to do something like this, but then we need to figure out what are the programs that we should be using there, which programs are useful, which are not, and why. I think that's going to be very interesting.