MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Vector Databases and the Future of AI-Native Applications with Weaviate’s CEO Bob van Luijt

    Bob van Luijt is the Co-Founder & Co-CEO at Weaviate. We cover why RAG is less risky than fine-tuning but remains primitive, how hybrid search combines semantic retrieval with BM25 for exact identifiers, and why doubling embedding dimensions doubles memory needs as deployments scale to billions of objects.

    02/15/2024

    Hosted by Matt Turck · with Bob van Luijt, Co-Founder & Co-CEO, Weaviate

    Vector databasesRAGHybrid searchEmbeddingsAI-native applications
    Listen now
    YouTubeApple PodcastsSpotify
    33 min · 14 chapters
    Contents

    Transcript

    What is RAG?

    1:00
    Matt Turck0:37

    It is Thursday, February 15th, and The MAD Podcast is back. Join us for conversations with leaders across the machine learning, AI, and data landscape with Matt Turck, partner at FirstMark Capital. Today we welcome Bob van Luijt, the co-founder and co-CEO of Weaviate. Weaviate is the open-source, AI-native vector database that helps developers create intuitive and reliable AI-powered applications. As always, if you love the show, hit the follow button to get the latest episodes every week. And off we go. Bob, welcome.

    Matt Turck1:10

    You are the CEO of Weaviate, which is an exciting company in the exciting field of vector databases. The company has raised $67 million in venture capital money. For those who count, you're a Series B company. And I thought a good place to start would be to do a level set around this emerging architecture, around whether that's called the modern AI stack or whatever, that's mostly focused around this concept of RAG, for retrieval-augmented generation. Can you help paint a picture of what that looks like, and where do vector databases fit in?

    Bob van Luijt1:51

    Yeah, sure. So first of all, thank you for having me. Thank you all for being here. So RAG, an abbreviation for retrieval-augmented generation. So the name kind of describes what it is, right? So you augment the generation of the generative model by retrieving something, right? And RAG was the first unique use case, if you will, that emerged around the ecosystem of vector databases. So I'm at this for quite some time already, before it was cool, basically. And back then, most was focused on what we call better search or better recommendations.

    Bob van Luijt2:24

    But RAG was something really new that emerged. And what RAG basically does is that the moment that you have a generative model, but you have your own data, and you want the model to generate something based on your own data, you somehow want to feed that information into the model. And this is where the vector database plays a role. And because, based on vector search or hybrid search, you can retrieve documents from the database, you can inject them into the generative model, and then you can just generate something based on your own data.

    Bob van Luijt3:17

    That is something that works better than fine-tuning, for example, because if you fine-tune, then you're still dealing with potential hallucination. Fair enough. With RAG, that's possible too, but it's less risky. But what's interesting to mention is that the way we—and with we, I mean everybody in this room—are doing RAG right now is actually pretty primitive, right? So we retrieve the data from the vector database, we inject it in the prompt. But a lot of exciting work is happening where the model actually knows how to retrieve based on the vector embeddings from the database itself.

    Bob van Luijt3:34

    So the model and the database start to weave, no pun intended, together.

    Matt Turck3:43

    Okay. And while we are in the definition part of this conversation, maybe define what an embedding is to make this interesting to a broad group?

    Bob van Luijt4:11

    Yeah, sure. So an embedding is, if you will, a geographical representation of your data. So I see some faces now go like... So a very easy way to think about it, if this is new for you, is that the example that I always give is based on a supermarket. So you say, like, if you have a supermarket and you look at the map of a supermarket, you have a 3D representation of stuff that's in the supermarket. And if you have a shopping list and the shopping list says, "I need apples, washing powder, and bananas," you kind of know that if you are looking at apples, the bananas are closer by than the washing powder.

    Bob van Luijt4:54

    And you know that if you walk to the washing powder, that you move away from the fruit section. And vector embeddings are a way, not only in language—but language is the use case we see the most right now—to represent how things are different from each other. And that's just by—I mean, if you want to, we can double-click on how that's done. But it's a way of describing similarity. And the power is in the fact that if you store a data object of, let's say, the Statue of Liberty, that based on the text, you knew that the distance between, for example, Paris and the Statue of Liberty was bigger than New York and the Statue of Liberty.

    Bob van Luijt5:19

    And the cool thing that we could do with that now is that we could basically—

    Matt Turck5:26

    The original Statue of Liberty is in Paris.

    Bob van Luijt5:54

    Yeah. So the point I want to make with that is that now the thing that we could do is that we could retrieve data from the database where we did not have exact matches or keyword matches. And that was, like, the unique thing that we saw based on vector embeddings, right? So these dimensional representations of how we store the data, basically.

    Matt Turck6:19

    So to play it back, you have a large language model, which is your AI engine. And then RAG is the system that enables you to, at query time, check the results against your data. And your data is stored in the vector database. And the way to get the data into the vector database is through the embedding models. Is that fair? Or is that too simplistic?

    Why is embedding models is such a hot topic right now?

    6:20
    Bob van Luijt6:38

    That's fair to say. And the common way that developers do that today is that they store the actual data objects combined with one or more vector embeddings in the database. So it works. The UX is like a traditional database, but it's purpose-built to deal with these vector embeddings. So the uniqueness sits in the index.

    Matt Turck7:08

    And the embedding model space seems to have gone from something that very few people talked about, that seemed like a solved problem—this idea of transforming your data into a format that vector databases can understand. That went from what felt like a solved problem to a very hot space in the LLM stack, including a race to the bottom in terms of price. What do you think that is, and what is happening there?

    Bob van Luijt7:36

    Oh, that's very simple. We're now in this transition from people who are trying it out, building POCs, prototypes, the usual stuff. And now they want to go into production. And the thing to bear in mind is that if you have an embedding of 768 dimensions versus 1,536, you need double the memory to store the latter, right? And if you do a POC with like 100 or 1,000 or a million data objects, that's fine. If you now scale that up to hundreds of millions or billions, it gets expensive quickly.

    What is your assessment of RAG?

    8:06
    Bob van Luijt8:06

    So a lot of work has been happening, one, in retrieval speed, because again, if you want to index 10,000 documents, you can wait. If you need to index 10 billion, you need speed. Speed becomes of the essence. So that's one thing. The second thing is work in the models themselves. So can we lower the size of dimensions that we need to store? And the third thing that sits a little bit in the overlap between the database and the models themselves, and that has to do with compression algorithms, binary representations, the whole shebang, just to make it as easy as possible to run this stuff in production.

    Matt Turck8:51

    So everyone these days talks about RAG as this kind of solution to the problem, like hallucination and all the things. A lot of people talk about it almost like a fait accompli, like it's something that everybody has agreed is working. What's your assessment? Is it really working? What are the issues, and how do you evaluate if it works in the first place?

    Bob van Luijt9:19

    So it's a first step. So what's very important for people to know, if you run a database company, then the big question that you always ask is, like, and people like yourself, what you guys always ask, is, so what's the unique use case? Right? And then, in the beginning, yeah, we can do stuff with search and recommendations. Look how cool it is. But yeah, we can see, what's the unique use case? So the moment—so RAG is quite old. I mean, old in the sense of how things are young and old in the space of AI.

    Bob van Luijt9:51

    So it's a relatively older concept. For those interested, you can actually, if you go into the Hugging Face library, see first iterations of it. It's very interesting. But so when ChatGPT came on the scene, a lot of people were like, we want to do this with our own data. And people were like, hey, it's great to use vector databases for this because the input queries are often not fully working for keyword matching. So this was, like, a beautiful use case for the vector database.

    Generative feedback loops

    9:53
    Bob van Luijt10:20

    And it's not always right because it's like pure vector search. So vector search alone is often not enough, right? So that's stuff like hybrid search, but we might get to that. But not only that, RAG is just the first step in new things we can do with vector databases. So one thing that I'm extremely excited about is something that we call generative feedback loops. So it's based on RAG with the vector database, but what it does is that it stores data back into the vector database.

    Bob van Luijt10:57

    And we have on our website, we have a blog called Generative Feedback Loops, and it's based on Airbnb data. And what you see in the dataset is that you have Airbnb data, but there's missing information. So descriptions are missing or stuff is incorrect. So what you basically do is you query, in this case, Weaviate. It runs through a generative model and it says, "Hey, this is Airbnb data, but it's incorrect. Fix this for me." It fixes the data or it fixes the description, creates a vector embedding, and stores it back in the database.

    Bob van Luijt11:28

    So now the model itself starts to interact with the database. So just give it a prompt, if you will. There could be a problem like, "Hey, there's something wrong with my data. There's an issue with my data. Fix it for me." And that is the thing that I'm super excited about. And I would not be surprised if this year, 2024, is the year of generative feedback loops, because this is really a seismic shift in the database landscape.

    11:46
    Bob van Luijt11:57

    It's now not people directly interacting with the data, but you give an assignment to the model, and the model knows how to interact with the vector database. And they do that by using vectors. So the point I'm trying to make is RAG is beautiful. It's a first step, but it's one-directional. And we're now going to go to looping it, so making it bidirectional, if you will. And with the multimodal models, we even have multiple directions that you can go into.

    Bob van Luijt12:04

    So it's exciting times.

    Matt Turck12:22

    So that's RAG as an emerging architecture for generative AI. Let's get more specifically into what you guys do at Weaviate. So, in particular, hybrid search seems to be something that you have spent a lot of time working on. What is it, and what are the benefits?

    Bob van Luijt12:47

    Yeah, so what's important to know is that Weaviate is open source. And if your database is open source, you get something beautiful, which is called a community. And the community starts to work with your database, and they tell you what doesn't work. And that can be, on one hand, just related to operational issues, but it's especially interesting from use case issues. And one of the things that people started to tell us—actually, two things that people started to tell us—was this.

    Bob van Luijt13:23

    One, they said vector search is amazing, but sometimes it's not enough. And in a bit, I'll share with you what that is. And the second thing that they said was, these vector embeddings, we need to work with these models, not everybody knows how to do that or how to efficiently do that. I mean, there are a lot of developers, like the people in the room here, that are very smart, that know how to do that. But not everybody knows that, right?

    Bob van Luijt13:53

    Or people want to just speed up their development. And so we learned these two things. And so, back to the first point. So if you have a query that says, for example, let's say that you have a dataset with customer support tickets and you say, like, "How was customer support ticket ABC123 handled?" The query itself is ideal for vector search, but matching on ABC123 is horrible. The model is super bad at that. Why? Because it was never trained on your knowledge base.

    Bob van Luijt14:23

    So what you do with hybrid search is that you kick off two queries simultaneously. So one, the vector search query, a score comes out. But with hybrid search, you do a BM25 search that works very well for ABC123, and it merges these results together. So now it basically upvotes, if you will, re-ranks based on the tickets that you have in your system or in your database that are related to ABC123. And then they organize based on that semantic query that you have.

    Bob van Luijt14:49

    Now, what does that have to do with the second part? If you build a vector database, you can take two approaches. Approach number one is you say people need to generate their own embeddings, and they just query the database based on the embeddings. You could do that. I mean, you can do that in Weaviate too if you want to. But what we figured out is that if we incorporate the models in Weaviate as well, so it doesn't matter if it's an OpenAI, a Cohere, an Anthropic, or an open-source model, it doesn't matter.

    What makes Weaviate special?

    15:15
    Bob van Luijt15:21

    If we incorporate them, we can bring the whole hybrid search feature out of the box. Because now you throw just the text at the database or the image at the database, and the database figures out where it needs to create embeddings, and it figures out where it needs to do the hybrid search. So the reason we built that is because the community told us the most optimal way for us to build applications is actually if we can do that out of the box.

    Bob van Luijt15:33

    So that's why we have that.

    Matt Turck15:38

    What are some other features that you would want to highlight that make Weaviate special?

    Bob van Luijt16:10

    So, the first thing is the deployment model. Because we live in 2024 now, we've learned from all these amazing existing infrastructure companies, like the generation before us, what the most optimal ways are for people to use databases. So you can use Weaviate serverless, BYOC, through marketplaces, embedded, open source in Docker, open source in Kubernetes—you name it. However you want to use it, you can use it. That's one. The second thing is that you can store the complete data object.

    Bob van Luijt16:43

    So it's very common to what you're used to from existing databases, but now it's really focusing on this, what we call AI-native stack, first. But the third one, and that's by far the most important, is that we've built the database to help you build AI-native applications. So what we mean with that is, if you're building something and you want to sprinkle some machine learning stuff over your application, it's great. I mean, you can use Weaviate too, but you can do that in other ways too.

    What about security?

    16:53
    Bob van Luijt17:07

    But if you say, no, I'm building something that has AI at the core, we give you all the tools and all the infrastructure to build it out of the box, to get you up and running in minutes. And that is the core value prop, if you will, of Weaviate. How did I do? Was it good?

    Matt Turck17:15

    Yes. How about security? It seems to be something that you guys spend a lot of time on.

    Bob van Luijt17:42

    Well, security is a— the short answer is yes, but the more elaborate answer is that no, we bake security in. Exactly. Yeah, that would be funny, though. So what starts to happen is that people start to move into production, and then you get just everything that you would expect from running in production. Like, hey, can we actually separate tenants? What kind of security features do you have? And so on, backups and so on and so forth. So, more the basic things that one would expect from a core piece of infrastructure.

    Does RAG accelerated the need for real-time data?

    17:45
    Bob van Luijt18:03

    Again, this is really led through a community of users and customers. They're just like, great, we want to move to production, but we really need this one thing. And then, great, we add it to the roadmap, the community upvotes it, and that's how we build it.

    Matt Turck18:28

    You have a lot of integrations. One that caught my eye in particular is Confluent. Is the idea that RAG needs to be increasingly in real time? Real time in this area that, for the last 10 years, people have been saying, "Oh, this year is the year of real time," and then the next year, "No, this year is the year of real time," and it seems to be happening. But do you think RAG accelerates that need for real-time data?

    Bob van Luijt18:57

    So I'm kind of laughing because of what you're saying with Confluent. So Confluent is, of course, amazing in real-time stuff, but the bottleneck is now these models. So these people at Confluent are waiting. It's like, "When are you guys real time?" Right? When is the serving real time? But on a serious note, why this becomes so interesting: at some point, we saw, like, it started for Weaviate in an optic that we saw community contributions to our Spark connector. That's how it started.

    How to define good vector database?

    19:27
    Bob van Luijt19:28

    And what started to happen was that—and this is an example of AI-native—is that a lot of companies said, just all the data that's going to come in, doesn't matter if it's transactions or text objects or whatever you can think of, we're not only now going to send them through our streaming pipelines as JSON objects, but we're going to attach these vector embeddings to them. Because now people don't have to run their machine learning model separately. We just run it.

    Bob van Luijt19:45

    And if you need it, you can just use it with the right model. So that's why it's so interesting to work together with companies like Confluent, because more and more companies start to literally stream the data with pre-calculated vector embeddings.

    Matt Turck19:59

    The vector database space has been very exciting to watch over the last year. Certainly a hot space in this whole list of interesting startups. Just at this event, we had Pinecone, we had Chroma.

    Bob van Luijt20:01

    I don't know.

    Matt Turck20:12

    How should people think about the criteria to evaluate one vector database versus the other? And you can name names, which would be fun, or not.

    Bob van Luijt20:33

    So the first thing is this. For a long time, we were waiting, like, when—okay, no, let me rephrase. So the vector embedding is just a data type. So it's just an array of floating-point numbers. I mean, you can store that in an old Oracle database, right? But the thing where it becomes interesting is the index. So the index in how you store and retrieve the information—you can do it in memory for speed, you can do it on disk if you want to optimize for storage and those kinds of things.

    Bob van Luijt21:14

    So what we start to see now is that basically every database under the sun supports vector embeddings. That's a great thing. That's, for me as a vector database founder, a champagne problem, right? Because that means that people see the value in using vector embeddings. So now if you look at the vector database space, you get to these purpose-built databases, right, that are really good at dealing with these embeddings first. So maybe if you're building something small or you're trying something out, you might want to store them somewhere else.

    Bob van Luijt21:42

    That's fine, right? Because it became a universal data type. And now you can make a distinction between that. You're saying, okay, do we just want to care about these vector embeddings, how we store them, or do we think a little bit bigger? And do we think about the developer that wants to build AI-native? So if you're building something, I would recommend thinking about, do we just want to store a couple of embeddings, or do we need to have help building AI-native from the ground up with model integrations, the embedding integration, that kind of stuff?

    What do you think about general purpose databases entering the field of vector-based databases?

    22:11
    Bob van Luijt22:29

    And that is what we try to solve. And there's another—it's like, you also have libraries like FAISS from Meta, right? That's more a library to store embeddings in. There's certain use cases where that's great, right? But again, we really want to become this core platform that people use to build AI-native applications on. And yes, at the heart is the vector database. But that's not the endpoint. That's the starting point where people build on top of.

    Matt Turck22:40

    I suspect the answer is going to be the same. But what do you make of the more general-purpose databases entering the field of vector databases? MongoDB in particular?

    Bob van Luijt22:53

    That goes back to my previous answer. Good for them, right? It's great that they've seen the light too. A bit late, though, but good for them.

    Matt Turck22:54

    Shots fired.

    Bob van Luijt23:26

    Yeah, but on a serious note, it doesn't really matter because the thing is that, in the end, what's important is how you help developers be successful in what they're building, right? And so MongoDB is amazing for developers building web apps and that kind of stuff. So if these developers want to do something also with vectors, great, good for them. And the same thing is, I don't know, if you think about Snowflake, it's great for your OLAP use cases. We are amazing if you want to build AI-native applications.

    Interesting use cases of Weaviate

    23:47
    Bob van Luijt23:57

    And that's how we want to be seen top of mind to the developers. And yes, the vector embedding and this vector storage sits at the heart of that. But that's just the core of the onion, right? And all the layers around that are the tooling, the integrations, the model integrations, the developer experience, and so on and so forth to help you build AI-native applications. So them entering the field, great, because that shows that there's a market fit.

    Bob van Luijt24:06

    And we specifically focus on those who want to build AI-native.

    Matt Turck24:17

    Talking about customers and go-to-market for a second, what are some fun examples of what you've seen people build with Weaviate? Some interesting use cases.

    Bob van Luijt24:42

    So the use cases are still pretty much—I mean, we see some clustering around e-commerce and those kinds of things, but it's still very much all over the place, I think. So let me give you two examples of things that I'm very proud of, just that come to mind. One are just tools that I use myself. It's kind of nice that if you see a blog from Stack Overflow saying that they use Weaviate, oh, thank you, I use that.

    Bob van Luijt25:18

    That's one thing that I'm proud of. And the second thing that I'm proud of is if you are a core part of the technology of new applications that people are building. I mean, I came in today and this gentleman came up to me and said, like, hey, we're using Weaviate, core part of our stack. So great, that's wonderful, right? We saw that with, I don't know, Instabase, for example, how they rewrote some core infrastructure with Weaviate.

    What’s your sense of the current state of the market?

    25:27
    Bob van Luijt25:46

    So those kinds of things are something that I'm extremely proud of. That's just really cool to see. So it's a combination of tools that I use myself that I'm very proud of, or just that it's such a dream come true that you're helping people to be successful with the businesses and applications that they are building. That's surreal. That's amazing.

    Matt Turck26:00

    What's your sense of the current state of the market? You mentioned earlier in the conversation people starting to scale with a lot more objects. Where are we? Are we still super early and people are just starting to deploy, or are you seeing some evolution?

    Bob van Luijt26:31

    Yes, I mean, people start to deploy, but the cases that people are deploying are kind of lagging behind the really new use cases. So the most deployments, like really the billion-scale deployments, are very much just similarity search, right? And then the next wave now is, like, the hybrid search cases, and then the big use case after that, the RAG use cases. And I hope that later this year we see those generative feedback loop cases. So really those core AI-native cases. But as one would expect, the core production cases lag a little bit behind what people are experimenting with.

    Open source vs commercial product on Weaviate

    26:53
    Bob van Luijt26:54

    So that's where we are right now. But it's still—I mean, may I ask who had never heard about a vector database before they came to this event? Okay, so we still see—okay, this is so funny. It's like, this is fewer hands than I expected.

    Matt Turck26:57

    There's all the people that didn't dare raise their hands.

    Bob van Luijt27:12

    Exactly. But the nice thing is, we still have a lot of work to do when it comes to education, right? We need to help people understand how they can build with these kinds of databases. So, long story short, it's still early days.

    Matt Turck27:22

    So you're very much an open-source company. It was an interesting question for me to hear how people think about how much do you put in the open source versus how much you put in the commercial product.

    Bob van Luijt27:49

    So open source is a jobs-to-be-done problem. The database itself, you're not selling the database. What you're selling are the services around the database. So we create proprietary software that is like our serverless offering. Those are our BYOC things that have to do with monitoring, the whole shebang, right? And around that, with a graphical user interface, Weaviate Apps, that kind of stuff. So the core database is open source, and it's a way for people to tinker around with what's happening.

    Bob van Luijt28:21

    Important to know, open source also builds trust with customers. It's not a black box. You can just be super transparent about what you're doing and how you're doing it. So that's the role of open source. But the job to be done, the value to be captured, does not per se sit in the open-source technology itself. It sits in the layers around that. Somebody today forwarded me a Harvard article. It's called, like, "The Value That Open Source Is Creating" or something.

    Bob van Luijt28:52

    I'm not sure if you've seen that. It was, like, two weeks ago. And the article was super interesting. The article is about—and I'm saying this from the top of my head, so the numbers can be off a little bit—but they said how much the economy is thriving on open source. And they also made a calculation that if open source would not exist, how much companies should invest in rebuilding it. And so I believe that I saw in the abstract of the article that they estimated that the value of open source to companies right now is $8 trillion, right?

    How did it all get started?

    29:23
    Bob van Luijt29:30

    Five times be that number. I might be slightly off; I don't remember it correctly, but the point is, if you're building $8 trillion in value with all these open-source companies, you can capture some of that value in the layers around the surface that you're building, right? So that's basically the premise of open source: that the cake is so big that you just try to figure out how you can capture value from that cake. So it's like a small cake and you eat all of it, or you have a big cake and you eat a little bit of it.

    Matt Turck29:59

    So maybe to close, because then I want to open up to people if they want to ask questions, tell us more about the company. So presumably you're recruiting and all the things, but tell us maybe the quick version of the founding story, what you did before, how you started it, and give us a sense for how many people you are, what you're recruiting for, that type of stuff.

    Bob van Luijt30:22

    Oh, sure. Yeah. So I've been working in software for a long time. That's because I was born in '85. So in 2000, when I was 15, we just got an internet connection, and there were people who were like, "We need a website." And I was like, "I can build you a website." So that's kind of how I started. So I'm Dutch, so then you go to the Chamber of Commerce and you start your little company. And it was time to go study.

    Bob van Luijt30:50

    So I studied music, as one does. And after that, I just continued building my software consultancy business. And in that capacity, I was working for a publisher. And this is 2015, and they were looking for new ways to build new products based on the articles that they were selling. And it was like, "Hey, there's this new thing with machine learning." Back then it was GloVe and FastText, single-word embeddings. Maybe we could do something with it.

    Bob van Luijt31:15

    So I started to play around with it. I started to vectorize a little bit of these things. Back then it was just a CSV file with words and embeddings. And I was like, "Hey, there's something in here. There's something here." And then I went to Google I/O in 2016. I fact-checked this, so this is actually correct, that Sundar Pichai said, "We're going to move from mobile-first to AI-first." And I was like, "I know what they're doing."

    Bob van Luijt31:45

    They're using these vector embeddings to vectorize, I guess, web pages and that kind of stuff. So that was when we started Weaviate. And back then we didn't call it vector. Nobody called it a vector database back then. So we started to look, and more people seemed to have this idea, right, that you can do something with vector embeddings. And then all of a sudden it took off, and OpenAI had a search endpoint that they deprecated. This is pre-ChatGPT. So they deprecated that search endpoint and said, "Well, you can just buy our embeddings."

    Bob van Luijt32:14

    There are three ways that you can index them and search through them. They said one is FAISS from Facebook, the other one is Weaviate, and the third one, I forgot the name of the company that they had on that list. And then I can tell you, if OpenAI adds you to the website, everything goes up. And then you had the whole ChatGPT thing, and then people were like, "We want to do that with our own data." And then it was through the roof.

    Bob van Luijt32:40

    But I can tell you, I mean, it was a long journey. This was not overnight. Oh, and by the way, to answer the rest of your question, so we are 100% remote. We have everybody from San Francisco and Berlin. We're like 60-some people, always interested to hear from interesting people. There are like a couple of Weaviate people there. Nick, Erica, say hi. So you can talk to them and you can talk to me. Did I answer everything?

    Matt Turck32:45

    Yes, that was wonderful. Thank you so much.

    Bob van Luijt32:46

    Thank you.