It is Thursday, February 15th, and The MAD Podcast is back. Join us for conversations with leaders across the machine learning, AI, and data landscape with Matt Turck, partner at FirstMark Capital. Today we welcome Bob van Luijt, the co-founder and co-CEO of Weaviate. Weaviate is the open-source, AI-native vector database that helps developers create intuitive and reliable AI-powered applications. As always, if you love the show, hit the follow button to get the latest episodes every week. And off we go. Bob, welcome.
You are the CEO of Weaviate, which is an exciting company in the exciting field of vector databases. The company has raised $67 million in venture capital money. For those who count, you're a Series B company. And I thought a good place to start would be to do a level set around this emerging architecture, around whether that's called the modern AI stack or whatever, that's mostly focused around this concept of RAG, for retrieval-augmented generation. Can you help paint a picture of what that looks like, and where do vector databases fit in?
Yeah, sure. So first of all, thank you for having me. Thank you all for being here. So RAG, an abbreviation for retrieval-augmented generation. So the name kind of describes what it is, right? So you augment the generation of the generative model by retrieving something, right? And RAG was the first unique use case, if you will, that emerged around the ecosystem of vector databases. So I'm at this for quite some time already, before it was cool, basically. And back then, most was focused on what we call better search or better recommendations.
But RAG was something really new that emerged. And what RAG basically does is that the moment that you have a generative model, but you have your own data, and you want the model to generate something based on your own data, you somehow want to feed that information into the model. And this is where the vector database plays a role. And because, based on vector search or hybrid search, you can retrieve documents from the database, you can inject them into the generative model, and then you can just generate something based on your own data.
That is something that works better than fine-tuning, for example, because if you fine-tune, then you're still dealing with potential hallucination. Fair enough. With RAG, that's possible too, but it's less risky. But what's interesting to mention is that the way we—and with we, I mean everybody in this room—are doing RAG right now is actually pretty primitive, right? So we retrieve the data from the vector database, we inject it in the prompt. But a lot of exciting work is happening where the model actually knows how to retrieve based on the vector embeddings from the database itself.
So the model and the database start to weave, no pun intended, together.
Okay. And while we are in the definition part of this conversation, maybe define what an embedding is to make this interesting to a broad group?
Yeah, sure. So an embedding is, if you will, a geographical representation of your data. So I see some faces now go like... So a very easy way to think about it, if this is new for you, is that the example that I always give is based on a supermarket. So you say, like, if you have a supermarket and you look at the map of a supermarket, you have a 3D representation of stuff that's in the supermarket. And if you have a shopping list and the shopping list says, "I need apples, washing powder, and bananas," you kind of know that if you are looking at apples, the bananas are closer by than the washing powder.
And you know that if you walk to the washing powder, that you move away from the fruit section. And vector embeddings are a way, not only in language—but language is the use case we see the most right now—to represent how things are different from each other. And that's just by—I mean, if you want to, we can double-click on how that's done. But it's a way of describing similarity. And the power is in the fact that if you store a data object of, let's say, the Statue of Liberty, that based on the text, you knew that the distance between, for example, Paris and the Statue of Liberty was bigger than New York and the Statue of Liberty.
And the cool thing that we could do with that now is that we could basically—
The original Statue of Liberty is in Paris.
Yeah. So the point I want to make with that is that now the thing that we could do is that we could retrieve data from the database where we did not have exact matches or keyword matches. And that was, like, the unique thing that we saw based on vector embeddings, right? So these dimensional representations of how we store the data, basically.
So to play it back, you have a large language model, which is your AI engine. And then RAG is the system that enables you to, at query time, check the results against your data. And your data is stored in the vector database. And the way to get the data into the vector database is through the embedding models. Is that fair? Or is that too simplistic?