MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Nomic: Truly open AI

    Brandon Duderstadt is the CEO and Co-Founder at Nomic AI. We cover why curating training data lets GPT4All remove refusal and one-word-answer behaviors, how Atlas surfaces model failures as localized patches in a dataset map, and why releasing model weights without training data limits reproducibility and auditability.

    02/01/2024

    Hosted by Matt Turck · with Brandon Duderstadt, CEO and Co-Founder, Nomic AI

    Open-source AIGPT4AllData qualityLLM benchmarkingEmbeddings
    Listen now
    YouTubeApple PodcastsSpotify
    42 min · 9 chapters
    Contents

    Transcript

    What is Nomic AI & how it got started

    0:46
    Matt Turck1:12

    Hey guys, welcome. So today is an important day at Nomic because you are releasing a new product, which I'm excited to talk about in a little bit in the second half of this conversation. And I'm also grateful that you chose The MAD Podcast to talk about the offering on the day of the launch. I think thanks to you, we've officially graduated to being a news organization. So thank you for that. But I'd love to start from the top and talk about you guys.

    Matt Turck1:20

    What are your personal journeys and your journey to Nomic AI?

    Brandon Duderstadt1:26

    Yeah, so hey everyone, my name is Brandon. I am the co-founder and CEO here at Nomic AI.

    Matt Turck1:29

    I'm Zach. I'm an ML engineer at Nomic.

    Brandon Duderstadt1:57

    And I really got started with Nomic just based on some experiences that I had at several previous roles. I had been a machine learning engineer in several different verticals, be it robotics, defense, medicine, finance. And I had started seeing the same problems over and over again, a lot of the issues revolving around data quality and how that affects the downstream properties of the models. And when you're working in these domains like defense and medicine and finance, where it's super high impact and you really have a vested interest in your models being right, you start to learn very quickly that a lot of the utility that these models can provide is unlocked if you spend a lot of time really curating your datasets.

    Brandon Duderstadt2:46

    And so, the reason that I ended up deciding to start Nomic with my co-founder, who actually worked with me at the previous role at Rad AI, which is a generative medical startup that was shipping Transformers on radiology text back in 2018, was because we saw how impactful really upping the data quality in these products could be. And so we started that journey together and then were fortunate enough to run into Zach through some open-source work. And I'll let him tell that story.

    Matt Turck3:13

    Yeah, so I came across Nomic and Brandon and Andrej at this VC fund's, I guess, open office. So I worked out of there. I was working at a biotech startup and I first came across Atlas. And I was super intrigued by it because I had been trying to train some models on basically DNA sequences, and I didn't have a great way of evaluating it. And I just wanted to see, like, okay, what does my data look like? And PCA wasn't working just because it would take too long.

    Matt Turck3:42

    So that was my first introduction. We started hanging out every now and then, talking about ML, bouncing ideas off of each other. And on the side, I'd been doing a little bit of open-source work in some Discords with Eleuther and Carper. And one weekend, Andrej comes to me, he's like, "Oh, I think there's a really unique opportunity here. We'd love to train this model."

    Brandon Duderstadt3:46

    We have all this data for it, but I don't have the time.

    Matt Turck3:51

    Being the co-founder of a three-person company, your time's incredibly limited.

    Brandon Duderstadt3:52

    I was a little bit bored.

    Matt Turck4:02

    I think my girlfriend was out of town that weekend and I had some extra time. So I was just like, "Oh yeah, I'll spend some time and train this model if you guys pay for the compute." So it worked out well.

    Brandon Duderstadt4:05

    And I think I joined Nomic, I want to say, like, three or four weeks later.

    Matt Turck4:23

    But yeah, it was kind of right place, right time. Great. Excellent. And Brandon, if I remember correctly, you were at NYU, right? You were doing a PhD at the time you were starting Nomic. Is that right? Did you drop out of it, or what was the story there?

    Brandon Duderstadt4:35

    This is my co-founder, Andrej. Yeah, yeah. I actually, funny enough, also dropped out of a PhD to join the radiology startup where we met. So, we like to joke that, yeah.

    Matt Turck4:56

    I think that's a really interesting pattern that I've seen across many founders, to be PhD dropouts. So, meaning having the fundamental intellectual caliber and intellectual curiosity of doing a PhD, but not the patience of lasting the whole way until completion of the PhD. It's a very interesting heuristic.

    Brandon Duderstadt5:05

    Well, that's exactly it. The two big factors for me, and this is going to be a tangent now, but the two big factors for me were velocity and impact, really. When I ended up leaving Hopkins to join Rad AI, the reason that happened is one of the co-founders of Rad AI sat me down and was like, "Look, I'm sure your thesis is great and all four people that know what a latent structure model of a random dot-product graph is are going to love it."

    Brandon Duderstadt5:47

    But if you want to actually come here and ship code that is going to help radiologists and be deployed in clinical practice a week later, now is the time and we have the GPUs. And that's pretty compelling. And even nowadays, it's increasingly hard to get those GPUs in academic positions. And so it's making increasing amounts of sense for, I think, the flight that we're seeing from academia to industry.

    Matt Turck5:52

    So for a while, that was you and Andrej.

    Brandon Duderstadt5:53

    Andrej.

    Building GPT4ALL

    5:57
    Matt Turck6:15

    Yeah, Andrej. And so I remember you guys burst onto the scene with GPT4All, which I remember felt, from the perspective of an outside person looking at the project, like an overnight success and just crazy open-source traction. And that was a weekend project, right? Wasn't it? What was the story there?

    Brandon Duderstadt6:34

    I mean, that's the model that Zach was talking about training. And it's also funny how everything can look like an overnight success from the outside. But I think a key part of what made GPT4All so good is we had spent the last year building Atlas, and it was the first time we really applied Atlas to the task of cleaning training data for a model. And so within a weekend, we could go in and, if you look at the original GPT4All paper, we show, oh, here's how you identify the sections of the training data where the model's refusing to respond.

    Brandon Duderstadt7:10

    And, oh, here's where you identify the parts of the training data where the model's giving one-word answers. And obviously you want to remove those so the model that you're training doesn't learn those behaviors. And so it's one of those things which I think is so true in startups, where the effort-payoff curve is super nonlinear. It's very much like a series of step functions where you put in all this effort, and then only when you pass some threshold, your payoff shoots up.

    Matt Turck7:15

    So you're saying that it was not as easy as it looks? That's a surprise.

    Brandon Duderstadt7:17

    Fortune favors the prepared, is what I'm saying.

    Running LLMs on a personal computer

    7:23
    Matt Turck7:39

    Yeah, exactly. So, having said that, why do you think that was so successful? Was there, like, just—I think, Zach, you mentioned right place, right time about your journey into the company—but was that true of GPT4All as well? It just hit the right kind of crowd at the right time, or was there any lesson that one could derive from the success? Yeah, I think for me, the thing that stood out to me, and then I think after the fact made a lot of sense, was that at the time, it was really hard to run these models if you didn't have the compute to.

    Matt Turck8:18

    So one of the things that Andrej really advocated for while we were training the model was like, okay, we should get this quantized and be able to run on computers that most people have, or a lot of people have. I was originally against it. I was like, this isn't that interesting. I don't care. But it turns out he was right. I think the access part of it was the bigger part versus the actual model itself. The model itself, at the end of the day, somebody trained a better model the next week or the next month or something like that.

    Matt Turck8:48

    There's been many iterations of the model, but it was more so the access of it. Having models on your computer, where people who care about privacy, I think, really helped shine through there. Great. And maybe as a level set to make this interesting for everyone, so what was GPT4All? What was the concept of the product? What did it do?

    Brandon Duderstadt9:13

    Or the model? Yeah, so it started as basically just a reaction to GPT-4 becoming increasingly closed source. We, as a company full of people that basically met through open source and have been open-source contributors through our careers, were pretty frustrated with the fact that this open lab was like, yeah, we're not going to tell you anything about the data. We're not going to tell you anything about the training. And we'll get into this later when we're talking about benchmarking closed versus open models.

    Brandon Duderstadt9:46

    But it was born out of that frustration, I think. The original inspiration was like, let's just make a sick open-source model that everyone can use and show people how to do science. But really what it grew into now is sort of this incredible open-source ecosystem where you have all these companies around the world now shipping models in the open source. Mistral is doing really awesome work there. Hugging Face is doing really awesome work there. Replit is doing really awesome work there.

    Brandon Duderstadt10:15

    And we're now part of this kind of weird, almost emergent assemblage of hackers and hobbyists that are taking these models and running around and quantizing them, or actually training or slurping new models in the open source. And we sit at this junction now where we help to ensure that those models can run on a wide variety of different hardware so that no matter what computer you have access to, you can run the models. And we make sure that you can actually do retrieval-augmented generation, which we'll talk more about later.

    Brandon Duderstadt10:36

    On these models in a very easy way, so in sort of like this drag-and-drop interface. And so really just doubling down on the accessibility portions of this that I think made the GPT4All project such a success in the first place.

    Matt Turck11:06

    And maybe for the dumb VCs in the room, do you want to explain what quantizing means? Is that compression, effectively? Yeah, it's effectively taking the weights of a large model, reducing the precision on them, basically taking the weights, making these numbers smaller so that you can actually load the model into your computer. Just because if you didn't do that, your computer would effectively blow up. I remember trying to run this when we first started doing it on my computer. I had an older MacBook, and at the time it was just barely fitting into RAM and it was on fire.

    Matt Turck11:37

    But now with the newer MacBooks, it's a lot easier to run these with the new Apple Silicon. So to drive this home, GPT4All is now a family of models that could be Falcon, that could be other things. And the common point is that you've quantized all of them. So all of them are now small enough to be run locally. That's the common characteristic of what you did across all the models.

    Brandon Duderstadt12:00

    That's right. So, I think the most popular model right now is Mistral 7B, quantized to 4-bit precision, fine-tuned on the OpenOrca dataset. And one of the two key factors that I really want to drive home is the contributions here. One is making sure that those models run as fast as possible across a wide variety of pieces of hardware. So it's not just NVIDIA GPUs. Like, if you have an AMD GPU, if you're running Metal, if you only have a CPU machine, you should be able to run as fast as possible in all of those situations.

    Brandon Duderstadt12:27

    And also making retrieval-augmented generation easy. So if you want to personalize these models on your machine, that should be as easy as dragging the files that you want the model to have access to into a folder. And so those, I think, are the two really, really big value adds of GPT4All right now.

    Matt Turck12:54

    I love what you said a minute ago about this collective of hackers, or I forget the exact word you used, but that was great. The open-source AI world right now seems to be just completely exploding and super exciting. Is that, for you guys who are deeply into the space, the same impression as the rest of us? What's your overall sense of the health and vibrancy of the open-source AI ecosystem?

    Brandon Duderstadt13:18

    It's all-time highs, I would say, right now. We've got a ton of VC money flowing into companies that are shipping things open source. I love that that's happening. I think on a podcast right after GPT4All came out, I think it was the Weights & Biases podcast with Lukas, I said something like, the biggest challenge for open source is gonna be getting the monetary resources to do a 7-billion-parameter model or these bigger models. And so, it's amazing to see companies like Mistral actually going and doing what they say they're going to do and releasing these amazing models.

    Brandon Duderstadt13:41

    And Meta as well, with the release of Llama, I think, has played a big part here. And I think also one of the reasons it's kind of popping off right now is we really are at this kind of new frontier in terms of the discovery of what these things can do. And so it's totally possible that some random person in the middle of nowhere that's just somewhat interested in this stuff spends enough time poking at it. There's so much new stuff to find that they can discover things that are really amazing.

    Brandon Duderstadt14:19

    And so you see some of these really incredible techniques coming out of not maybe what you might call the royal science or the universities, but some random person will be like, oh, hey, I did this new type of interpolation between these two model weights and it turns out it gets way better, or things like this. Or, oh, I happen to have gone and curated this dataset that is now incredibly useful for fine-tuning towards XYZ task. These things are things that a single person can do on commodity hardware that really move the needle now.

    Brandon Duderstadt14:38

    And so I think now that there's attention on it, and because it's such an inflection point in terms of what is possible with the technology, people are really able to do a lot without needing access to a ton of resources.

    Matt Turck14:49

    Well, there you go. VC money actually helping. Who knew? So you all have a relationship with llama.cpp. Do you want to talk to that? Yeah.

    Brandon Duderstadt15:13

    So this has been, as Zach was talking about earlier, one of the biggest parts I think that was integral to the success of the original GPT4All is that it was quantized. We used llama.cpp for that. And I think one of the things that we've been really proud of here at Nomic is the fact that we've been able to actually bring staff on full-time to work on things like improving llama.cpp. And so right now we have Nomic staff working back and forth with Georgi on getting new compute backends merged in to improve the support for these models across a wide variety of hardware.

    Brandon Duderstadt15:53

    And I think the lesson here for the broader open-source ecosystem is, as economic pressures start to apply and the VC money starts to run a little bit drier, staying collaborative and making sure that, as a whole, the open-source ecosystem itself continues to be able to move forward. And if someone else's repo is one that's useful in your project, being open to collaborating with them and being flexible about where those boundaries lie is really, really important. And so that, I think, is a lesson that I really hope all of the people building in open source right now take to heart.

    Nomic Atlas

    16:00
    Matt Turck16:10

    Okay, excellent. So that's GPT4All. But as you mentioned, there was actually a product before GPT4All called Atlas. What does that do?

    Brandon Duderstadt16:17

    Yeah, so Atlas is a tool for exploring and interacting with massive unstructured datasets using only a web browser. It was born out of a lot of the work that my co-founder, Andrej, and I were doing at this radiology generative AI startup, Rad AI, where a lot of the work on the models involved exploring these massive datasets, finding places where maybe the doctors had made mistakes or the model had made mistakes, looking for patterns in those errors, attributing that back to the training data, looking for distribution shifts in the actual production data.

    Brandon Duderstadt17:05

    When we were at Rad AI, there was this super interesting global health event called COVID, you may have heard of it, that happened. That was this massive distribution shift in our query stream. At the time, we didn't have the tooling like Atlas to deal with that in a very flexible way, and it caused a lot of headaches, certainly. And so, effectively, the interface that we converged on is what would be ideal for this: this massive scatterplot that you can run in your web browser, where every piece of data in your dataset is a point, and two points are close together if the pieces of data are similar.

    Brandon Duderstadt17:53

    And what we found is this sort of layout is particularly useful for debugging AI models because their errors tend to pop up in localized regions of it. So we're actually doing a collaboration with Hugging Face recently on their open-source IDEFICS model, looking at where that model performs systematically well and systematically poorly. All of the places in the latent space, in this representation space that we're showing people, where the model was performing poorly were all spatially localized. They popped up as little patches of land in the data.

    Brandon Duderstadt18:29

    Similarly, all the places where maybe the model was even memorizing data popped up as these little patches. That was incredibly informative for us in terms of detecting bad data, detecting not-safe-for-work data that should have been removed, detecting data that was, like, misextracted, detecting things the model was memorizing, or detecting categories of data that the model was biased towards reproducing, things it did particularly well on. So that is sort of a broad motivation from the ML engineering perspective of how Atlas got started.

    Matt Turck18:56

    Great. And bearing in mind that you're a young startup and in a super fast-moving industry, how does that all work? Like, those two pieces, how do they work together, including from a go-to-market perspective? Because GPT4All presumably is more of a developer audience. Is Atlas more of an enterprise audience? Is Hugging Face an enterprise customer for you?

    Brandon Duderstadt19:11

    Atlas is definitely a B2B enterprise tool. We have very generous limits for individuals, like power users, academics as well, especially, to make it so that they can leverage it. But it is kind of the core growth engine behind the monetization of Nomic. And I think that puts us in a really, really interesting position relative to a lot of other open-source companies because it means that we can keep GPT4All as this very pure love letter to the community, open-source sort of project.

    Brandon Duderstadt19:53

    And we won't be pressured to eventually find a way to squeeze it for dollars. But I think the two interact also very well from a funnel standpoint. So the kind of person that is interested in tooling to help them understand and curate massive unstructured datasets is the kind of person that's either, one, using that to train models, or two, generating a lot of that kind of data. And that's exactly the kind of person that is interested in tools like GPT4All. And so a lot of the first enterprise sales that we had with Atlas were inbound in the GPT4All Discord.

    Brandon Duderstadt20:09

    Maybe this is just what building a company in 2024 is, but our sales funnel is literally like Twitter to Discord to enterprise sale. It's insane.

    Matt Turck20:25

    Any lessons learned there from a go-to-market perspective? Is that just generally comparable as selling developer tools, meaning you need to be authentic and you need to be responsive? What have you learned?

    Brandon Duderstadt20:52

    Yeah. So, I think one of the most interesting things I've learned was: be open to your thesis about core audience shifting. So, when I originally built and envisioned Atlas and sort of brought it to Andrej and the team to help me realize it, I very much thought that the ICP, ideal customer profile, would be a machine learning engineer. And a lot of our early users were machine learning engineers. We still have a lot of them using the system. But something that struck me as very interesting is Atlas started to get picked up by a lot more kind of, like, I don't want to say less technical teams, but an archetype that's more of a business analyst or someone doing business intelligence, where we see consulting companies that get these datasets from their customers and maybe they have domain expertise, but they can't code.

    Launching Nomic Embed

    21:33
    Brandon Duderstadt21:33

    And so, never before have they been able to interact with this data with this level of velocity and this level of granularity. And so, we've seen a much wider adoption in terms of how technical the actual users of the systems are than I originally anticipated, which has been really, really interesting to see.

    Matt Turck22:03

    All right. I think we've come to the moment of the conversation where we are going to talk about the, drumroll, new offering launched today. So, what is it? Yeah. So, we're launching Nomic Embed. It is the first open-source, reproducible, long-context text embedder that beats OpenAI Ada as well. We are releasing the model weights, the training code to replicate the model, as well as the training data. And I think this is something that's near and dear to my heart, as I spent the last few months kind of painfully going through all the data, curating it, and searching for needles in a haystack in papers and at the edges of GitHub.

    Matt Turck22:51

    So, really excited to get this in the hands of people and see how they use it. All right, very cool. Congratulations. So, maybe as a refresher for everyone listening to this, what is an embedding model? Yeah, so I think the way that I like to think about it is an embedding model is a function of semantic meaning in English, or whatever language, to semantic meaning to computers. And it's becoming increasingly important for AI applications now, especially retrieval-augmented generation, which helps update basically models without having to retrain them.

    Matt Turck23:27

    So you can augment pre-trained models with new information, which obviously is very useful and less expensive than having to retrain yourself. So, to play it back, it's effectively translating text and unstructured data into the kind of numbers that machine learning models understand and digest.

    Brandon Duderstadt23:28

    Yes, correct.

    Matt Turck23:52

    And what has happened to this whole space? It's interesting. Maybe call it a year ago now, I was having this kind of conversation with vector database founders and CEOs, and it was very much presented as a solved problem and something that was like, yeah, you need this thing at a moment to input data into the vector database, but it's no big deal kind of thing. And it seems like over the last few months, it's become a huge area of activity, with a lot of people competing.

    Matt Turck24:03

    So what happened to that space?

    Brandon Duderstadt24:26

    Yeah, I think the fundamental thing here is people realize that there's a new data primitive, and that data primitive is here to stay, and that's the vector, the embedding. And just to bring us back to Atlas for a second, in the same way that we see vector databases taking that new data primitive and trying to play out the implications of that at the database layer of the stack, I think we're gonna see a radical shift like that at all layers of the business software stack.

    Brandon Duderstadt24:52

    And the current conception of what I believe Atlas to be in the limit is the logical extension of what happens when you make embedding vectors a fundamental primitive at the sort of visualization or business analysis, business intelligence level of the stack. So, internally, we love to throw around calling it sort of like a Tableau for unstructured data.

    Matt Turck25:21

    So, maybe you mentioned OpenAI Ada. Maybe paint a quick picture of the space? Like, what are the other embedding models that either have existed for a while or are just coming up? Yeah. So roughly in this space, on the open-source side of things, there's a bunch of really great models that have been released. The model weights have been released, and a few of them are being hosted on embedding endpoints. But the problem with a lot of these models is that they have a context length, so they're limited to 512 tokens, which I think nets out to maybe a few paragraphs, which becomes a problem when you have large documents.

    Matt Turck26:04

    Maybe you have long financial documents and you want to reason over the whole document versus chunks, just because it gets a little confusing and it's a little hard to reason about the right way to chunk up a large document if you need to recall something from the first paragraph and the last paragraph. So that's the open-source side of things. There's one or two long-context models. The ones that are good are either too big, they're in like the 7 billion parameter range, or they don't actually beat OpenAI Ada's model.

    Matt Turck26:24

    And then, for the long context on the closed-source side, Ada is kind of the de facto for many people's applications today.

    Brandon Duderstadt26:24

    Yeah.

    Matt Turck26:49

    But with closed source, the big problem is you have no idea what data they've trained on, and you kind of have no auditability on what's going on with the data, the training, basically anything that's going on under the hood. So this is kind of the gap that we're filling: an open-source version of Ada that outperforms it on a few benchmarks and full auditability. Embeddings are super important also because they're part of this, I guess, process called RAG that a lot of people are now becoming familiar with.

    Matt Turck27:12

    Do you want to maybe re-explain quickly what RAG is and how embedding fits in? Yeah. So to me, RAG kind of—there's three main components to RAG. It's the embedding model, a language model, and a vector database. So the way that I like to think about it is, if you have a collection of documents, say it's Wikipedia, and you want to ask a bunch of questions about it, but you don't want to actually read every document, what you can do is you first embed all the documents of Wikipedia using the text embedder, store them in the vector database, and then you can start to ask questions about your documents.

    The Importance of Data in AI

    28:10
    Matt Turck28:13

    So if you wanted to ask, where was Beethoven born? You would take that question, embed that question, get the nearest or basically most similar documents to that question, and then pass those documents to your language model, where your language model can then reason over and answer your question given sufficient documents. Basically, you're treating your language model as a reasoner versus a lossy database. What are some technical details that you learned while building this new model? Yeah, I mean, it's, I think, a boring answer, but data really makes or breaks your model.

    Matt Turck28:48

    There's a lot of times where we were trying all these fancy tricks and, at the end of the day, we uploaded some data to Atlas and we're like, wow, this data is crap. This doesn't make any sense. So I think that data auditability, like we've said a few times now, just understanding what's going into your model, making sure that it makes sense to you as a human, I think gets you most of the way there. And I think that's one of the bigger points that I've had to relearn over and over again in every project that I do.

    Brandon Duderstadt29:11

    And so, I think also, one of the things that I'm most proud of that we're able to do with this, which is distinct from a lot of other open-source releases, is we're not just releasing the model weights, we're releasing the dataset that Zach spent all of his time curating. And what this is going to do is it's going to allow people that want to build upon and adapt the model to, one, be able to reproduce it, and two, be able to start tweaking it and learning for themselves without having to go through that arduous process.

    Brandon Duderstadt29:42

    And as far as we're aware, none of the other top models on any of these benchmarks have gone out and released their datasets. Most of the time, we just see open weights. Sometimes we see a codebase, but you almost never see the datasets. So we're really proud that we're sort of delivering the full end-to-end package to the open-source community.

    Matt Turck29:59

    Yeah, that's awesome. Why don't people do that? Is that a reproducibility kind of issue? What is it? Is that marketing? Is that coming across as open source but not being very open source?

    Brandon Duderstadt30:25

    There are a couple of answers here. I think the truest answer is that the data is often the secret sauce. When people talk about, like, oh, what is the secret sauce of your AI model? It's almost always the data. Sometimes it's clever scaling and you'll have architectural tricks that move the needle, but almost always it's the data. And so if a company wants to appear open source without maybe really being open source, they'll release the weights but not the data. For some companies that we're now seeing, you're seeing Voyage pop up, for instance, and Cohere is another one that are really investing heavily into embeddings as a service.

    Brandon Duderstadt30:51

    That's their IP. That's their core IP. And so for their business to work, they can't release it. But then you start to get into these very weird conflict-of-interest scenarios where it's like, okay, but if they don't release it, one, there's an auditability question on that model. And then there's another question, which is, if they don't release the data and they don't have necessarily the tooling to actually comb through all of it, it's really hard to make sense of if the benchmark scores that they get are super legit or not.

    Benchmarking LLMs

    31:10
    Brandon Duderstadt31:12

    Maybe they're legit, but you can't actually check, right? And so I think that's an element to it as well, perhaps.

    Matt Turck31:50

    How does the current benchmarking system work to that point, compared to closed source in particular? Yeah. So with respect to the embedding models, there's this benchmark called MTEB. I think it's called the Massive Text Embedding Benchmark. And it comprises a few different tasks: classification, retrieval, which kind of ties back to the retrieval-augmented generation point, clustering—how close can you get the embeddings of related documents together and exclude any false positives?—and a few other tasks.

    Matt Turck32:22

    It's really helped, I think, improve the field a bunch. But with any benchmark, I think once you start measuring it, it quickly becomes a game of, like, how can we make this number go higher, with some maybe questionable tactics on how to make that number go higher. That's on the, I guess, general embeddings, like, how good are your embeddings. Since most of these models don't have a long context, there aren't a ton of benchmarks, but there recently have been two long-context benchmarks that have been released.

    Matt Turck32:55

    One, I think, was released in, I want to say, October by Jina AI, comprises four datasets. And then recently, Hazy Research from Stanford released a long-context benchmark called LOCO. So those were the three that we focused on and compared against some of the other long-context models.

    The Future of Nomic AI

    32:56
    Brandon Duderstadt32:56

    Okay, very cool.

    Matt Turck33:24

    So congratulations again on the launch. So as a result of that, you're now a three-product company, product/models. You've got Atlas, you've got GPT4All, and you've got Nomic Embed. So how does that fit in, especially the latter part? We talked about the other two from a go-to-market perspective a bit earlier, but how does that latter part fit into the overall go-to-market model?

    Brandon Duderstadt33:50

    Yeah. So, one of the key portions of the Atlas pipeline is actually embedding all of your data. And so one of the things that we're excited to start doing is, as we build up this core competency of building these incredible embedding models, we can release them to the public and also incorporate them to improve the Atlas system. And so there's sort of direct product value and sort of direct auditability on the Atlas side of things that I think is really important. And beyond that, going back to if you want to talk about go-to-market and how GPT4All fits into that, this idea of if we can do these things where we're really just giving as many people as possible access to this technology, be it what you need for retrieval-augmented generation with Nomic Embed, or what you need to actually get up and running with the generative models themselves with GPT4All, regardless of your hardware, that's just going to bring more people into the AI ecosystem faster.

    Brandon Duderstadt34:42

    And I think that demand sort of directly translates to need and sort of usage of Atlas. Because again, as people are producing massive unstructured datasets and consuming massive unstructured datasets and tweaking these models and really starting to learn the lessons that we constantly relearn, which is, like, it is literally always the data. If something is wrong, it's almost always the data. The understanding that having tooling to really dive into your data and understand it at an incredibly detailed level is just going to proliferate more widely.

    Matt Turck35:13

    And so, what's the roadmap/vision for the company? Fast-forward a couple of years or three years. Is the idea that you guys are going to be very prolific in releasing new products based on where the market is going? And then what's the end result? Is it like a suite of different products that you're going to integrate into some kind of platform? How do you think about it?

    Brandon Duderstadt35:36

    Yeah, it goes back to the core mission of Nomic. Our objective function is to improve the explainability and accessibility of AI models. On the accessibility side, we want to, as much as possible, continue releasing new open-source tools for the community for free to bring more people, regardless of where they're at, regardless of what resources they have, into the post-AI world. I think this addresses what I would say the number one risk of this technology is, which is the sort of unequal access element to it.

    Brandon Duderstadt36:09

    The version of the world that I fear most is like there's two or three mega-companies that lock everything down in the early days. And so the sort of, like, I don't want to say inequality necessarily, but the inequality induced by that only gets worse and worse and worse. That, I think, is the number one future we have to be watching out for. And then the explainability side of things really is, I think, Atlas. I cannot emphasize deeply enough how important looking at and understanding your model inputs and outputs are for actually building models that are safe and doing the sorts of things that you intend them to.

    Brandon Duderstadt36:22

    To do.

    Matt Turck36:48

    Cool. Awesome. So, to close on a kind of a light note, one of the cool things, seen from my perspective as a New York-based VC, is that you guys are building the company largely out of New York. And I'm curious about your experience doing that. Obviously, there are awesome, awesome things happening in San Francisco, the whole, like, Silicon Valley thing. There's a little bit of a dominant narrative that AI can only happen there, which I find, depending on the day, amusing or irritating.

    Matt Turck37:05

    And I'm curious what your take is and what your experience is as a New York AI startup.

    Brandon Duderstadt37:30

    The New York AI ecosystem is incredible. It's definitely smaller than the Bay, but I think that gives it a level of, like, tight-knittedness that you almost lack in the Bay. And as someone that's sort of been a part of both, I spent a good amount of time living on Steiner Street in Hayes Valley, and I spent the last couple of years living here in New York. I think I prefer New York. And the main reason being, I think the challenge with Hayes Valley is in the sort of cerebral Valley, and now the AI arena, it can get a little bit like, I don't want to say groupthink necessarily, but there is a dominant narrative.

    Brandon Duderstadt38:09

    And it's like, I think when you're surrounded by that, it's quite hard to kind of think a little bit outside of that box. And one of the things that I love about New York is you're not just surrounded by tech. You are surrounded by incredible people in all walks of life, be it art or acrobatics or companies that are doing things that are not necessarily AI but up-and-coming. That aspect is really key. Plus, we have sick AI heavy hitters in New York.

    Brandon Duderstadt38:38

    Hugging Face has an office in Brooklyn. Where's Hugging Face SF? It's not there. And especially as Paris continues to be an up-and-comer in the scene, it's a faster flight from New York to Paris than it is from SF to Paris. So I think if I were to bet on where the global nexus necessarily would be, if you're accounting for the fact that there are things happening outside of the US, I think New York is a very strong candidate.

    Nomic AI is hiring

    39:10
    Matt Turck39:10

    Obviously, music to my ears. I just don't think all of this needs to be an either-or. That's what I find a little much about the SF narrative. SF is fantastic, and you can have fantastic things happening in different places. It doesn't need to—ultimately, we're all trying to build, or in my case, finance, companies that are largely going in the same direction across the world. Anyway, rah-rah New York. I appreciate it. Cool. All right.

    Matt Turck39:23

    So, time to talk about—you’re hiring, I’m sure. What kind of roles? How do people find you? How do they find you online? All those good things.

    Brandon Duderstadt39:39

    Yeah, you can follow us on Twitter, @nomic_ai. We are always hiring for particularly interesting people, is how I would describe it. I think everyone at Nomic is not only excellent at their work craft, but every single person has this interesting kind of plus-one to them. So, I know Zach hates it when I mention this, but for Zach, it's not only that he's an excellent machine learning engineer that can single-handedly drive a project to open-source something as good as one of the top companies in the world, but he also played D1 baseball, right?

    Brandon Duderstadt39:58

    It's that little extra something that I think really interests me.

    Matt Turck39:59

    Yeah, I like it too.

    Brandon Duderstadt40:25

    And also, we hire a lot out of open source. So, if people want to get involved, we hired Zach out of open source. A couple of people that we picked up during the big hiring spree after we raised our A—Adam, Aaron, Jared—these are all people that had interacted with us on open source in Discord before. Not necessarily before it was cool, but we saw them shipping and giving back to the community. And that's something that I think is really baked into what we want to do here at Nomic.

    Brandon Duderstadt40:37

    So, if you want to work with us, hit our Discord, come work with us, and we'll see what we can do.

    Matt Turck41:01

    Do you find that there is any difference between people that come from the prior world in AI and have done, I guess, what on some day in November 2022 became overnight classic AI, or classical AI? Or do you need generative AI-native folks, or is that an irrelevant question?

    Brandon Duderstadt41:08

    It's a false dichotomy. Yeah, it's a false dichotomy. You need all of the above, and it's questionable if that's even a distinction that should be made.

    Matt Turck41:26

    I've heard it made. Okay. But that's interesting and reassuring. Okay, awesome. This was super great. I really appreciate you guys' time today. Again, I appreciate that you chose The MAD Podcast today to talk about the new announcement, and congratulations on Nomic.

    Brandon Duderstadt41:50

    Thank you to everyone for coming back and joining us for the first episode of season two of The MAD Podcast. We will be back to our regular publishing schedule weekly, each Wednesday, with new conversations with leaders in the machine learning, AI, and data landscape. If you like the show, you can find the video recording of this episode, along with many, many more, on the Data Driven NYC channel on YouTube, and you can find all the important links in the show notes.