MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Pinecone: Long Term Memory for AI with CEO Edo Liberty

    Edo Liberty is the Founder & CEO at Pinecone. We cover why AI systems should retrieve external data at inference rather than fit it into models, how vector databases serve as long-term memory across billions of embeddings, and why adding retrieved context can measurably reduce hallucinations.

    05/31/2023

    Hosted by Matt Turck · with Edo Liberty, Founder & CEO, Pinecone

    vector databasessemantic searchLLM retrievalAI memoryhallucination reduction
    Listen now
    YouTubeApple PodcastsSpotify
    27 min · 1 chapters
    Contents

    Transcript

    Full episode

    0:00
    Matt Turck1:17

    Cool. Welcome, Edo. You are the founder and CEO of Pinecone. Pinecone develops a vector database that makes it easy to connect company data with generative AI models. The company was started in 2019. You've raised $138 million in VC money, at least according to good old trusted Crunchbase, including a $100 million Series B that was just announced. Congratulations on that, and welcome. So I would love to start with a bunch of definitions. We're all in a moment where we're trying to understand what's going on in generative AI, and there's all these terms that are sort of flying around that I think most of us are trying to understand.

    Matt Turck1:41

    So I'd love to start with vector embeddings. What are they, and why do you need them?

    Edo Liberty2:09

    All right, so when you look at the deep learning models, they are mathematical objects. They're number-crunching machines, and in fact, the information that propagates from layer to layer or within the network is always a set of numbers. It's just that. There is no other way to propagate anything inside a neural net other than that. In fact, that's exactly the same thing that's happening in your brain, say in your visual cortex or in your auditory cortex or in any other parts of your brain.

    Edo Liberty2:50

    One process communicates with the other as a set of activations of neurons. They either fire or they don't. That's a set of numbers. So that's the way neural nets and AI in general represent data. That set of numbers is called a vector or a vector embedding. The embedding is a mathematical term that refers to placing one object in another space. So taking an image, which is in pixel space, and putting it in another space, which is the high-dimensional number-array space, in a way that somehow keeps some of the representation of it.

    Edo Liberty3:16

    So that's the embedding. The vector is the array of numbers, and the embedding is the translation kind of preservation property.

    Matt Turck3:27

    And vectors help you understand whether there's similarity or not between concepts. Is that the right way to think about it?

    Edo Liberty3:51

    Correct. So in the same way that you would see somebody that you know, and your visual cortex would process the face, and then with your temporal lobes, you would access the kind of face data stored in your brain and figure out, oh, that's my cousin, that's happening in the vector embedding space. That's happening with the similarity search. That's happening with a mechanism that looks a lot more like memory than it looks like processing like a neural net.

    Matt Turck4:33

    So let's say I have video data, audio data, text data. So if I understand correctly, I need to translate this into numbers to fit it into a machine learning model, an AI model. How do I do that? How do I take video and translate that into numbers?

    Edo Liberty4:51

    We can talk about the mechanics of how to do it, which is maybe the less interesting thing because the less technically inclined of you don't care and the more technically inclined of you can figure it out. If you zoom out, though, the interesting part of it is that I think there is a significant shift in the way that people think about building these smart AI systems, and they don't try to fit the data into the model as much as give the model access to the right data when it's doing the inference.

    Edo Liberty5:43

    And so, I gave this example several times, but if you study medicine, you already go into medical school knowing English. And then you study a body of knowledge. You study medicine. You don't study medical English. You read a book and then know medicine, and then you speak about it because you know how to speak, right? You don't become a doctor by just doing the rounds for 10 years and then know enough, know how to sound like a doctor.

    Edo Liberty6:16

    So that's the answer in some sense. You don't fit it. You allow the model to access external memory and use it correctly. And you give it the right data. So if you want it to be a doctor, you give it medical knowledge. If you want it to use videos, you give it video data.

    Matt Turck6:31

    And that is vectors and vector embeddings. And that leads us to vector databases. So we'll talk about Pinecone in a second, but what is a vector database in that context?

    Edo Liberty6:51

    A vector database is the infrastructure that supports that long-term memory. So if you have a very small amount of data, you can fit it on the machine and you can kind of play around with it. That already is becoming almost unnecessary with these larger contexts and so on. But when you have millions or billions of documents or images and so on, you really have to have a very specialized system to be able to do this cost-effectively, with low latency, high persistence, and so on.

    Edo Liberty7:14

    You really have to start thinking about things in a fundamental way and not just about the data.

    Matt Turck7:43

    And for anybody that spends time in the MLOps world, in the last couple of years before vector databases became a thing that everybody's talking about, the concept of storing stuff that would be fed into machine learning models, people heard the term maybe feature store. Is that completely different? Is that part of the same family? How would you think about it? Yes.

    Edo Liberty8:16

    No, I'm just kidding. Yeah, no, it's a completely different thing, right? Feature stores are used mainly either for marshaling data into training or using real-time features for objects that change very frequently, right? So, if I'm a user and I just clicked on something, maybe you want to incorporate that last feature that just happened a second ago in my next classification, in my next action or something. Vector databases are completely different. It's really about this long-term memory, about these billions of embeddings of documents or images and so on, and making them available to large language models, all these other multimodal models and so on, where they operate more like parts of your brain than like a feature store.

    Matt Turck8:47

    Okay, fantastic. Maybe one last concept, and thank you for explaining all of this so clearly. Semantic search is something that, again, people hear a bunch. What does that mean?

    Edo Liberty9:17

    So semantic just means by meaning. And search means search. Traditional search is keyword-based, not because keyword search is that great, but just because we have great mechanisms to run it at scale and it's very efficient. Keyword search is a pretty old technique. In fact, books had inverted indexes in the backs of them. They're even called an index. They've had them for a while. In fact, you had indexes in books before we had print.

    Edo Liberty9:48

    So that's a pretty old—well, before we had print in the West. The Chinese were printing well before the Gutenberg Bibles were even thought of. But anyway, it's like early 1200s at least. So keyword search existed for a while. Great technique. Works great for Google, works great in other places, doesn't work when you want to search something by meaning, right? When you remember a conversation that you had, I don't know if you've ever searched your inbox for an email you know for a fact you read, right?

    Edo Liberty10:18

    And the only way to do it is try to hack the search system to figure out what word was there that I didn't use in another email. Like, you try to backwards-engineer the crappy search. With semantic search, the whole idea is that you know what it means. You should be able to search by the meaning in free text, and the match would not be because, oh, like three words matched and they're next to each other, but rather the meaning is similar, and so I can retrieve that.

    Edo Liberty10:43

    And again, that looks a lot more like how humans remember things, and not like how inverted keyword lookup stores look.

    Matt Turck11:05

    Okay, so all of this is a wonderful, hopefully, intro into Pinecone and what Pinecone does. Tell us about the company, the history, maybe your background, how you came to start this company, especially because you started it, I believe, in 2019, which was well in advance of this whole craziness right now?

    Edo Liberty11:19

    So, I did my PhD in theoretical machine learning and big data algorithms, and my postdoc in applied math. And that was at Yale, right?

    Matt Turck11:22

    Yes.

    Edo Liberty12:04

    And already then I was working on high-dimensional geometry and functional analysis and kind of the foundational basis of machine learning. And I was always drawn into this search problem because it seems especially gnarly. I then did two internships at Google on that topic. I joined Yahoo as a scientist. I became an adjunct professor at Tel Aviv University. And through this entire time, I kept publishing and working in industry on infrastructure for big data and for machine learning. Stayed at Yahoo for about seven years, which moved me back to New York after a short stint in Israel.

    Edo Liberty12:45

    And then in 2016, Swami Sivasubramanian, who now runs all databases and AI services in AWS, back then he wasn't a VP yet; he was only a director. But he called me up and he said, "Hey, we're starting this AI thing in AWS. Do you want to come?" And I said, "Okay, who's we?" And he said, "Me and Alex." It was Alex Smola, who is—I don't know if you know him. Look him up.

    Matt Turck12:46

    Very.

    Edo Liberty12:58

    And yeah, that was a journey I couldn't resist, like come to AWS and basically build SageMaker and all these other great services out of—

    Matt Turck13:00

    It's huge now. There must be thousands of people now, right?

    Edo Liberty13:01

    What?

    Matt Turck13:01

    How big is the group?

    Edo Liberty13:35

    Several thousand now, yeah. But I'm saying this, I'm kind of recounting the whole story because throughout that time, vector search and vector representations and embeddings and that kind of basic concept kept getting more, gradually more and more traction all the time. More and more people kind of understood what it does and what's happening with it. And then at some point, the BERT embeddings, kind of the transformer models, kicked in, and then there was a big spike in understanding. And then that's it.

    Edo Liberty14:03

    In 2019, in some sense, I felt like that was the right time. I felt like it was going to get us about two years to kind of build the right kind of infrastructure correctly. And we kind of timed it correctly. And so this is what happened. We didn't, by the way, foresee any of this ChatGPT thing happening. We knew it would keep growing. But that, I think, completely took everybody by surprise, including us.

    Matt Turck14:28

    Yeah. So I can only imagine how wild a time it must have been at Pinecone over the last few months. And so vector databases are becoming very much a central part of what seems to be emerging as the AI stack. So, when people talk about, like, how do you do this whole thing? Vector database is the term that everybody comes back to. What does a stack look like, I guess, based on what you see customers do?

    Matt Turck14:44

    This vector database, this LangChain, there's different models. How does that all fit? What does that look like now?

    Edo Liberty15:25

    Yeah, so there is a—I mean, it depends what you call—I’ll focus very narrowly on the large language model stack. Of course, there are the models themselves. There are vector databases. There are libraries that connect them in all sorts of different ways, like LlamaIndex and LangChain and others. There are, I think, a lot of connective services, but I think we're going to see an emergence of agents in some form or another. I think that's still very—

    Matt Turck15:29

    Do you want to briefly explain what agents are?

    Edo Liberty16:02

    So agents are these recursive pieces of software. I don't even know how to define them. It's sort of like this life hack on large language models. It's like, how do you use these large language models, enable—like, let them use tools like search or different APIs and so on—and then use their own answers to somehow feed back into an input and try to accomplish very complicated tasks, right? So it's not like a one-time prompt.

    Edo Liberty16:38

    Like, if you want to get in shape, right? I mean, that's a plan. You need to fit it into a schedule. You have to find a gym. You have to whatever. There's a sequence of things you need to do, and you can research every one of them separately. And you might think as a human, it wouldn't be like a language generation thing. It would be like you sit there and plan it. So these agents are now coming up as this new paradigm of building more complex sequences and plans.

    Edo Liberty16:56

    Yeah, I put an asterisk next to that. I think this will shake up some surprising ways in the next few months. We'll see.

    Matt Turck17:22

    Great. Let's talk about some of the use cases that you've seen with Pinecone. Maybe if you can use some customer examples, but as I was prepping for this, search, generation, security, personalization, data management—pick a handful of those and sort of double-click on how people use Pinecone.

    Edo Liberty17:56

    Yeah, so first of all, one of them is obvious semantic search. Like, people are literally just reinventing their own search stack. I don't know if you ever tried to truly optimize an Elasticsearch cluster. I mean, and after three weeks wanted to throw your keyboard out the window, right? So with large language models, a lot of that is sidestepped and improved in a significant way. And people use Pinecone to just drive the backend search and storage based on these vector embeddings rather than keywords.

    Edo Liberty18:40

    But that's very obvious and immediate. Nowadays, we see a huge wave of people using us to create context for ChatGPT-like and Bard-like applications. So call them chat as a whole. But chat is very general. It could be customer service. It could be—you name it. But there is a very wide variety of use cases, from anomaly detection to—one of my favorite applications is somebody built face detection for cows. And we're like, I didn't even know you can face-detect cows.

    Edo Liberty19:01

    I grew up in a city. I didn't even think cows looked different. So, I mean, but then I told this to a farmer and they got deeply offended. So I'm probably offending some people here. I apologize.

    Matt Turck19:25

    Maybe that's related to the cows question. But one really important part of vector databases is because they are the long-term memory, as you were saying, they are a very important tool against that problem of hallucination that people talk about a lot. Is that the right way to think about it?

    Edo Liberty19:55

    100%. In fact, I was looking at experiments today from one of our teams, and we literally measure reduction in hallucination as one of the core metrics that we try to drive. So, 100%. And by the way, I mean, this is—it's almost like, obviously, if I ask you a question for which you don't know the answer, but I compel you to say something anyway, you're going to make something up. But if I give you the right context, you can answer a lot more accurately.

    Edo Liberty20:26

    And that's exactly what's happening. And so the question is, can we retrieve the right context out of which an informative answer can come out? And so, in some sense, it's obvious that that should happen, but it's harder to actually get the whole thing to cascade correctly and to work measurably well.

    Matt Turck20:56

    So very practically, I'm a company. I want to deploy GPT-4 in the enterprise, but for a mission-critical application like chat, where I don't want it to say stupid stuff to my customers, I would use Pinecone. I would put basically my source of truth—this is like the history of customer conversations—into Pinecone. And then how do I make GPT go fetch the information in Pinecone?

    Edo Liberty21:04

    When you want to query, when you have a question, say, or some prompt, you want to add context to it. The way you would do it is you take that prompt, you would embed it with some language model, search in Pinecone for the context, which could be maybe the top, I don't know, 100 relevant documents or 100 relevant parts of sentences or paragraphs, and put that back with the context and say, "This is the prompt, and here is more context out of which you can give an answer."

    Edo Liberty21:48

    And that ends up being—even when you do that relatively crudely, because you can do what I said right now in like 1,000 different ways. You can ask, which language model should I use? And how should I parse the data? And how should I chunk it? And why should I put it? And what's the order? You can play with it in a million different ways. But even if you do it relatively naively, that already gives you a huge bump in accuracy or reduction in hallucination, depending on how you want to measure it.

    Matt Turck22:11

    Let's see. Look, I'm a VC. I'm always interested in that part as well. Let's talk about go-to-market. Like, how are you building the company? How are you finding customers? How are you selling to them?

    Edo Liberty22:44

    So we started Pinecone, or I started Pinecone, as a managed service. Because I'm a really big believer in the buyer-driven journey, right? I think in the modern world, if you want to use some technology, the last thing you want to do is send an email and wait for three days. That sounds very unnatural. You want to use it immediately, right? And if it's great, you want to just start using it and forget about it.

    Edo Liberty23:22

    Be happy and move on. So we really built the whole company around the self-serve journey. Today you can start with Pinecone, run some example in five minutes, build a demo app in two hours, integrate some kind of a micro-POC for your team in a day, and in a week be in production and never talk to us once, if you choose not to.

    Matt Turck23:22

    Right.

    Edo Liberty23:47

    Some companies today—very large companies—tend to also want to talk to us. And so we're now building a sales team that is happy and willing to engage and help and so on. But it's more about assisting in the journey if that's wanted, or kind of standing, like, not being in the way when we're not wanted.

    Matt Turck23:59

    All right, last question from me. I'm curious about the space in general. It's become what feels like a very sort of rich and vibrant slash competitive space. Very, very—

    Edo Liberty24:00

    What has become?

    Matt Turck24:24

    The space, the vector database space. And I'm curious, less in the context as a prompt to badmouth competitors, but more as a way for us to understand the space. Like, all those companies that are vector databases now, are they very different approaches, or are you all sort of going in the same direction and it's more of an execution kind of play to win the market?

    Edo Liberty24:34

    Yeah, I'll try not to badmouth anyone. I'll just say stuff about us, and you can extrapolate.

    Matt Turck24:36

    Oh, feel free. This is recorded, but, like, feel free.

    Edo Liberty25:08

    Yeah. Yeah. Only YouTube will know about this. Yeah. I mean, look, I'll tell you what's driving me and my team, okay? And first of all, it's the technology itself, right? I mean, the internals of these databases are actually incredibly complicated numerical data structures. And really, it's very cutting-edge technology, both on the distributed system, the data management, the persistence. We had to rebuild our own object storage. Like, we basically had to build the whole stack, from the numerical indexes to the object persistence to the distributed data to the query planner, the whole thing.

    Edo Liberty25:49

    So as a systems and algorithms person, Pinecone is a fascinating place to be. And I can tell you that we are far and away better in terms of efficiency of running this service at scale and its stability. And we are just getting started. Like, there's so much to come. Okay, so that's number one. And the second thing is, having spent a lot of time in AWS, I got indoctrinated a little bit to be customer-obsessed rather than competitor-obsessed.

    Edo Liberty26:15

    And we are delighted and happy that we have a lot of very demanding and enthusiastic customers. And so we're very focused on what they want rather than what our competition is doing.

    Matt Turck26:24

    Yeah. And by the way, just as a last word, as I was prepping for this, I read somewhere that you guys got like 10K signups a day.

    Edo Liberty26:26

    Yeah.

    Matt Turck26:50

    That's insane. Anyway, that's so exciting. That feels actually like a wonderful place to leave it. Thank you so much. This was super, super interesting. It's really incredible how Pinecone has just vaulted to the forefront of consciousness in this world of generative AI. So congratulations on all the incredible work that you guys have been doing.

    Edo Liberty27:02

    Thank you. Thanks for listening to The MAD Podcast. If you liked this episode, be sure to leave us a review. com/events/data-driven.