Yeah, exactly. So, having said that, why do you think that was so successful? Was there, like, just—I think, Zach, you mentioned right place, right time about your journey into the company—but was that true of GPT4All as well? It just hit the right kind of crowd at the right time, or was there any lesson that one could derive from the success? Yeah, I think for me, the thing that stood out to me, and then I think after the fact made a lot of sense, was that at the time, it was really hard to run these models if you didn't have the compute to.
So one of the things that Andrej really advocated for while we were training the model was like, okay, we should get this quantized and be able to run on computers that most people have, or a lot of people have. I was originally against it. I was like, this isn't that interesting. I don't care. But it turns out he was right. I think the access part of it was the bigger part versus the actual model itself. The model itself, at the end of the day, somebody trained a better model the next week or the next month or something like that.
There's been many iterations of the model, but it was more so the access of it. Having models on your computer, where people who care about privacy, I think, really helped shine through there. Great. And maybe as a level set to make this interesting for everyone, so what was GPT4All? What was the concept of the product? What did it do?
Or the model? Yeah, so it started as basically just a reaction to GPT-4 becoming increasingly closed source. We, as a company full of people that basically met through open source and have been open-source contributors through our careers, were pretty frustrated with the fact that this open lab was like, yeah, we're not going to tell you anything about the data. We're not going to tell you anything about the training. And we'll get into this later when we're talking about benchmarking closed versus open models.
But it was born out of that frustration, I think. The original inspiration was like, let's just make a sick open-source model that everyone can use and show people how to do science. But really what it grew into now is sort of this incredible open-source ecosystem where you have all these companies around the world now shipping models in the open source. Mistral is doing really awesome work there. Hugging Face is doing really awesome work there. Replit is doing really awesome work there.
And we're now part of this kind of weird, almost emergent assemblage of hackers and hobbyists that are taking these models and running around and quantizing them, or actually training or slurping new models in the open source. And we sit at this junction now where we help to ensure that those models can run on a wide variety of different hardware so that no matter what computer you have access to, you can run the models. And we make sure that you can actually do retrieval-augmented generation, which we'll talk more about later.
On these models in a very easy way, so in sort of like this drag-and-drop interface. And so really just doubling down on the accessibility portions of this that I think made the GPT4All project such a success in the first place.
And maybe for the dumb VCs in the room, do you want to explain what quantizing means? Is that compression, effectively? Yeah, it's effectively taking the weights of a large model, reducing the precision on them, basically taking the weights, making these numbers smaller so that you can actually load the model into your computer. Just because if you didn't do that, your computer would effectively blow up. I remember trying to run this when we first started doing it on my computer. I had an older MacBook, and at the time it was just barely fitting into RAM and it was on fire.
But now with the newer MacBooks, it's a lot easier to run these with the new Apple Silicon. So to drive this home, GPT4All is now a family of models that could be Falcon, that could be other things. And the common point is that you've quantized all of them. So all of them are now small enough to be run locally. That's the common characteristic of what you did across all the models.
That's right. So, I think the most popular model right now is Mistral 7B, quantized to 4-bit precision, fine-tuned on the OpenOrca dataset. And one of the two key factors that I really want to drive home is the contributions here. One is making sure that those models run as fast as possible across a wide variety of pieces of hardware. So it's not just NVIDIA GPUs. Like, if you have an AMD GPU, if you're running Metal, if you only have a CPU machine, you should be able to run as fast as possible in all of those situations.
And also making retrieval-augmented generation easy. So if you want to personalize these models on your machine, that should be as easy as dragging the files that you want the model to have access to into a folder. And so those, I think, are the two really, really big value adds of GPT4All right now.
I love what you said a minute ago about this collective of hackers, or I forget the exact word you used, but that was great. The open-source AI world right now seems to be just completely exploding and super exciting. Is that, for you guys who are deeply into the space, the same impression as the rest of us? What's your overall sense of the health and vibrancy of the open-source AI ecosystem?
It's all-time highs, I would say, right now. We've got a ton of VC money flowing into companies that are shipping things open source. I love that that's happening. I think on a podcast right after GPT4All came out, I think it was the Weights & Biases podcast with Lukas, I said something like, the biggest challenge for open source is gonna be getting the monetary resources to do a 7-billion-parameter model or these bigger models. And so, it's amazing to see companies like Mistral actually going and doing what they say they're going to do and releasing these amazing models.
And Meta as well, with the release of Llama, I think, has played a big part here. And I think also one of the reasons it's kind of popping off right now is we really are at this kind of new frontier in terms of the discovery of what these things can do. And so it's totally possible that some random person in the middle of nowhere that's just somewhat interested in this stuff spends enough time poking at it. There's so much new stuff to find that they can discover things that are really amazing.
And so you see some of these really incredible techniques coming out of not maybe what you might call the royal science or the universities, but some random person will be like, oh, hey, I did this new type of interpolation between these two model weights and it turns out it gets way better, or things like this. Or, oh, I happen to have gone and curated this dataset that is now incredibly useful for fine-tuning towards XYZ task. These things are things that a single person can do on commodity hardware that really move the needle now.
And so I think now that there's attention on it, and because it's such an inflection point in terms of what is possible with the technology, people are really able to do a lot without needing access to a ton of resources.
Well, there you go. VC money actually helping. Who knew? So you all have a relationship with llama.cpp. Do you want to talk to that? Yeah.
So this has been, as Zach was talking about earlier, one of the biggest parts I think that was integral to the success of the original GPT4All is that it was quantized. We used llama.cpp for that. And I think one of the things that we've been really proud of here at Nomic is the fact that we've been able to actually bring staff on full-time to work on things like improving llama.cpp. And so right now we have Nomic staff working back and forth with Georgi on getting new compute backends merged in to improve the support for these models across a wide variety of hardware.
And I think the lesson here for the broader open-source ecosystem is, as economic pressures start to apply and the VC money starts to run a little bit drier, staying collaborative and making sure that, as a whole, the open-source ecosystem itself continues to be able to move forward. And if someone else's repo is one that's useful in your project, being open to collaborating with them and being flexible about where those boundaries lie is really, really important. And so that, I think, is a lesson that I really hope all of the people building in open source right now take to heart.