Congratulations. Best form of flattery. All right, switching tacks, and in an effort to make those conversations educational for a broad group of people: one of the key aspects of the release is the thinking model. Could you remind folks what a thinking model actually is versus other forms of models or prior generations of models?
A lot of people have heard about inference-time scaling, which makes sense. If you spend more compute at inference time, you get a better answer. A thinking model is really a way to train the model to exploit that a lot. So you spend a lot of tokens, which are usually hidden from the user as a long chain of thought, and the model therefore has this step change where it's way better at math tasks, coding tasks, agentic tasks. I think our future plans are adding more tool use to the models.
We're not talking a lot about agentic search or agentic code execution on the fly and stuff for this model. But building thinking models is the gateway to doing a lot more interesting things, like Claude Code. Maybe we'll have OLMo Code next year and all these things that we want to do. The thinking model has just been the thing in 2025 that uses a lot more compute per answer. The model gets way better. I don't like thinking models, but it's fine. No, they're good.
Thinking models are really like work mode, and regular instruct models are usually more fun to build. They can be more quirky. But yeah, I think they're like 90% of the cases, especially user-facing cases. Folks are okay spending time waiting for this model to craft a better answer. There's still a space for models that can respond faster. You see stats that Google released about adoption of Gemini Flash, and that's where non-thinking models that can at least approximate, have a good approximate first answer, are really useful.
They're also more fun to build. But yeah, thinking models are where the future is, especially when it comes to agents integration. Before we go into the pipeline very specifically of the OLMo family, because, as you alluded to, that's one of the amazing things about open source, is that we can, in a discussion like this, truly understand how the model works versus other conversations with commercial players. So before we go into the pipeline, I'd love to talk a little bit about you guys, your backgrounds, and AI2, which is a very important player in the ecosystem that people may or may not have heard about.
So who wants to go first? I sort of stumbled into this role by just picking problems that are interesting. So my background: originally from Italy, moved to the U.S. for a PhD. My PhD is in information retrieval. How do you build a search engine, to simplify it a lot? I slowly got into more and more natural language. After grad school, I joined Amazon. I was working on Alexa, at the beginning working on the search part of Alexa.
And then I got, wait, the actual part where the users talk to Alexa? The interesting part. So, slowly moving towards that. Initially, I joined AI2 working on a project called Semantic Scholar. It's still active. It's a search engine for academic papers. And there, the interesting bits were actually interacting with users and less so the actual text of the papers that you were searching on. And then the way I got into LLMs and building language models is really intertwined with how AI2 got into building language models.
It all started around, what is it, November of 2022. This is around the same time ChatGPT got released. A bunch of researchers at AI2—this is individual contributors, it was not a direction from the top—a very grassroots initiative at AI2. A bunch of researchers got really interested in building a model that would be fully open. AI2 had already built sort of proto-language models around 2017, 2018. So a lot of the interest was in recapturing, expanding that line of work.
So a bunch of us got together, sort of started planning, got in touch with a few companies who might be interested in supporting these initiatives. We got initial grants. AMD, at the time, there was about 2 million GPU hours. And so we sort of had the idea, had the researchers interested, we had the compute. So we went to leadership at the time and sort of told them, hey, we're going to go do this thing. I hope you're okay with it.
And one of the nice things about AI2 is, at heart, we are a research lab. So everyone was like, sure, you figure everything out. Just have fun. Great. All right, Nathan, how about you? So you're a man of many talents. You do AI research, you write this very interesting blog/newsletter called Interconnects, you do podcasts, you do a bunch of different things. So tell us about your journey.
Yeah, I say I wear many hats to try to get the things that I want to do done. I showed up to Berkeley as an EE PhD admit in 2017, and then I saw that AI was happening and I decided that I wanted to try to do this, which started by going to all the names that people know, like Sergey Levine and Pieter Abbeel, and asking to be in their group. And then they respectfully said no. And then starts the long process of learning how to actually do it without being directly embedded in these elite groups, which was a mix of robotics and reinforcement learning and finding my way there.
My PhD was mostly in model-based reinforcement learning, and then my one research job was to go join Hugging Face when they said they were going to make an open-source version of DeepMind to do a bunch of research. Realistically, my job was not that impactful or useful at Hugging Face until ChatGPT came out, and then I was like, oh, I should maybe just learn about RLHF. That got very immediate traction as somebody trying to work in public with the team there.
So Lewis Tunstall and other people at Hugging Face are still doing a great job on this, and we worked together for a while. And then mostly I was just getting burnt out on remote work and met Luca in Hawaii at a fun conference and was like, wow, I could have real-life friends. And I joined AI2 to work in person and tried to do the same thing, which kind of takes an evolution of the OLMo story, which was just like I had a lot of motivation in trying to figure out these—what was mostly reinforcement learning from human feedback at the time.
And make versions of these post-training techniques public. Then that evolved through both OLMo, and we have our post-training methods named Tulu, which is like, we spent a long time trying to replicate what we thought was close to Llama 3 post-training with multiple stages and optimizers, which is the project that came up with the name Reinforcement Learning with Verifiable Rewards with a bunch of people. So it's kind of this evolving journey at AI2 in search of impact, which is what we think people are actually doing.
And then largely, the opportunity that Luca and I and others at AI2 fill is that there's so much money in AI, and it only becomes increasingly so, that the amount of people that can talk about these things in public and educate and get more people involved by spreading knowledge is ever smaller. So I describe my career journey as a lot of it is filling that vacuum and thinking about what's impactful. So it kind of pulls you when there's such a void. It has a sort of gravity to make it clear what you should be doing.
You anticipated my question, which is sort of obvious. In a world where we see hundreds-of-million-dollar, billion-dollar packages offered by some commercial AI labs for people just like you, I was curious about your interesting motivation to join AI2, which is a nonprofit. But impact is the short answer, correct?
Yeah. I mean, I've been here for two years, and I wasn't famous when I joined. So let that be told to people looking for new jobs: you want to find a job that you can grow into. And I think AI2 has been a really, really good place for that for many people because you have independence and are encouraged to go forth and do things and not be a cog in a broader grind-out-language-model machine, which is important, but it's harder to get visibility.
So we alluded to some of it, but maybe a few words about AI2. So AI2 was started by Paul Allen, right? AI2 stands for Allen Institute for Artificial Intelligence. You mentioned some grants, Luca, and I think earlier in the conversation we talked about a recent grant as well. I saw that it was $152 million from NSF and NVIDIA. So what is AI2? How did it start? Who founded it at a high level? AI2 was founded around 2014 by the late Paul Allen.
Initial AI2 was very focused on building machines that can do science, can understand science, solve science problems. That's when Semantic Scholar started as a repository of science papers. Slowly, one of the initiatives that started forming was more fundamental research around how language models work, how, at the time, what was called natural language processing was working. You had teams like AllenNLP doing great work. From very early on, we always had this idea of not just releasing artifacts or research, but releasing the tool.
Back in the day, we had this very widely used library called AllenNLP that would allow you to build and customize these models.
I'm going to jump in. It's cool because it's the namesake of our team name and has been for a long time at AI2. And it was the thing, it was the main competitor to Hugging Face Transformers, and they ultimately outcompeted AI2 as the thing that people use for that because they had a very different model and amount of support. Luca can keep going. Luca knows a lot more.
But we have been at it, open-sourcing for a while. I think it's something that folks here understood really early—this is before my time—that it was important both in pushing science and also unlocking commercial use cases that, as a nonprofit, maybe we didn't anticipate. You release a tool, people pick it up and do amazing things with it. Yeah. And we moved on in language modeling more and more recently.
Right now, AI2 has maybe like three main projects. One of them is the OLMo model family. And there are variants of OLMo. There are some that focus on the full pipeline, some that are robotics, some that focus more on processing images and video and audio. Is Molmo part of one of those variants? Yeah, you have Molmo. It's one of our projects that work on multimodal inputs. Recently, we released another one called MolmoAct that's more focused towards robotics, receives multimodal input, and then can act in space.
And then we had the model that was able to do automatic speech recognition. Another OLMo model focused more on document processing could do OCR. So it's a nifty little family of models. We have a working group on agents for scientific tasks, harking back to our roots. This is the Asta family of initiatives. This is agents to help scientists do their work. And that just came out, right? Like August of this year? Yep. The team has been cooking since the middle of last year, but finally we had our first release this year.
There's actually two releases. There was the main Asta release, and then recently we announced a partnership with CAIA, the Cancer AI Alliance, using some of the components in Asta to help researchers make progress on cancer research. And then there is a third branch on AI for the environment, building models that can see and understand so that they can model Earth and can work with different signals to do prediction around the environment and so on. I'm being a little bit vague on this one because I don't know if it has been announced yet.