Like, hey, is it truly low awareness? Do you also agree that it's a really, really important unbeaten benchmark? I think I walked away from that conversation believing the answer was yes to both those questions. I had just seen Nat Friedman run this Vesuvius Challenge the year prior, and it was really, really successful. It was one of these competitions to focus a lot of attention on a problem that most people just never heard about or were not aware of. I felt like we could emulate that, and ultimately that was kind of how we came to launch the initial version of ARC Prize together.
And let's get into the specifics of how that actually works. So, starting with 2024, but I guess the structure is the same today. So there's a private version, there's a public version, there is a paper prize. How does that all work? I think some of the goals that we have for the ARC Prize are to—one initially was to raise awareness of the fact that there was an unbeaten, really important AI benchmark out there that wasn't saturating. It resisted this 50,000x scale-up in pure language model systems.
This was to counter and provide some public education of the fact that pre-training alone is not enough for these pure language models. This was in contrast to, again, all of the hype, all of the dogma that I think you saw online. Just to put a really clear point on this, if you don't really remember this era, maybe remember last summer in California, there was that big SB 1047 bill, that AI bill that was primarily regulated and introduced on the thesis that this scaling was going to continue and lead to really, really bad situations, and we have to impose regulation right now because if we don't right now, it's going to be too late.
Yeah, it's like back when GPT-4 was released, the universal narrative, like everywhere in the tech industry, was, okay, GPT-4 is just a bigger version of GPT-3 trained on more data. So it's the same architecture, it's the same principles behind it. It's just scaled up, and look at this incredible step-function improvement. And GPT-3 itself was just a bigger version of GPT-2. And so the idea was that, well, GPT-5 is just going to be the same, but it's going to be 100x bigger.
And we're going to need these massive data centers to train it. And it's going to be truly AGI. AGI is just going to spontaneously emerge from scaling up this thing. And basically all the evidence was pointing in this direction. If you looked at benchmarks that were based on just brute-force memorization of knowledge and skills, if you didn't look at ARC, this was the takeaway you had. And so everybody at the time believed that. And maybe some more sophisticated researchers, like maybe some at OpenAI, I don't know, maybe some at DeepMind, had different beliefs, but that was really the mainstream narrative.
And, well, interestingly, one year later, like about one year and a half later, this narrative was gone. People were completely aware that no, just scaling up pre-training is not all you need, that you actually need entirely different ideas and maybe, in particular, test-time adaptation to achieve actual fluid intelligence and AGI.
So the first big goal is to raise this awareness of this fact. And honestly, it was to re-inspire folks. I think I'd spent a lot of time with researchers and students over the year prior to that. And one of the interesting sentiments I heard was just a lot of dejection. Like, hey, isn't all this stuff figured out? There's really not that many interesting things to do in AI. Maybe I should go work at the application layer on language models instead of the research layer.
I felt like this was a really not-good thing. We don't have AGI yet. We don't have the ideas for it yet. We still are idea-constrained, even literally today. If that's true, which I think ARC-AGI-1 and ARC-AGI-2 both show it is, you want to design the strongest innovation ecosystem you possibly can, which is going to be one that's very open. There's a lot of sharing, there's a lot of diversity of approach. What you don't want is a very closed-up environment with no sharing, where there's dogma or a monocultural viewpoint of how to do it.
And so we really wanted to try and inspire folks with our prize last year to work on new ideas. And to this point, one of the versions of the private version, correct me if I'm wrong, you have to publish exactly what you do. There's a limit on the amount of compute. Is that fair?
Yeah, I mean, the basic idea is that there's one track for self-contained approaches with no internet access that are very efficient. So they have a very limited compute budget, and authors must open-source them at the end of the competition. So this is really designed to incentivize open sharing and get as many ideas as possible, with a big focus on efficiency. We believe efficiency is not just a good feature to have in your system. It's actually at the heart of intelligence.
And the other track is to provide continuous benchmarking of frontier models to be able to track, okay, if you look at the best currently available commercial frontier models, things like right now, for instance, o1 Pro, soon in the future it's just going to be o3 and so on, Gemini 3, whatever. How much fluid intelligence do these models actually have? Mostly independently from the efficiency consideration. We still do have this notion that we want to monitor efficiency. So we're going to be reporting results on a 2D plot.
So we're not just looking at the score as a scalar, we're looking at the score associated with the cost per task that was required to achieve this score. And of course, a model that gets you the same score, but at much lower cost per task, is a smarter model.
One other thing we've got too, Matt, on the contest that's worth adding, just on the structure of it, is we've got these score tracks that François was just talking about. We also introduced something last year, which we're doing again for 2025, which is this paper award track. One of the critiques of things like Kaggle contests is you get a lot of overfitting to the dataset or to the benchmark. People do dataset mining, lots of very narrow approaches that maybe don't generalize because they're optimizing just to win the contest, win the top score.
Certainly, we actually saw some of that. There were some legitimate good conceptual breakthroughs as well, but certainly there's mixed in a lot of just hacking the benchmark. We had to introduce this paper award alongside it in order to basically incentivize conceptual progress. I actually think some of the best research that came out from ARC Prize 2024 was actually on the paper track last year. Things like these test-time training, test-time adaptation approaches, we actually got full, not only open-source reproducible code, but the theory that backed it too.
I think one thing that potentially ARC Prize 2024 will be remembered for in retrospect is going to be marking a moment in time where AI realized, ah, yes, we do need test-time adaptation methods to be able to solve things like ARC. And that's certainly what o3 shows as well. We spoke with François a little bit about the o3 aspect. But separately, the top result was—was it MindsAI? There was Jack Cole, that group.