Cool. Welcome, Edo. You are the founder and CEO of Pinecone. Pinecone develops a vector database that makes it easy to connect company data with generative AI models. The company was started in 2019. You've raised $138 million in VC money, at least according to good old trusted Crunchbase, including a $100 million Series B that was just announced. Congratulations on that, and welcome. So I would love to start with a bunch of definitions. We're all in a moment where we're trying to understand what's going on in generative AI, and there's all these terms that are sort of flying around that I think most of us are trying to understand.
So I'd love to start with vector embeddings. What are they, and why do you need them?
All right, so when you look at the deep learning models, they are mathematical objects. They're number-crunching machines, and in fact, the information that propagates from layer to layer or within the network is always a set of numbers. It's just that. There is no other way to propagate anything inside a neural net other than that. In fact, that's exactly the same thing that's happening in your brain, say in your visual cortex or in your auditory cortex or in any other parts of your brain.
One process communicates with the other as a set of activations of neurons. They either fire or they don't. That's a set of numbers. So that's the way neural nets and AI in general represent data. That set of numbers is called a vector or a vector embedding. The embedding is a mathematical term that refers to placing one object in another space. So taking an image, which is in pixel space, and putting it in another space, which is the high-dimensional number-array space, in a way that somehow keeps some of the representation of it.
So that's the embedding. The vector is the array of numbers, and the embedding is the translation kind of preservation property.
And vectors help you understand whether there's similarity or not between concepts. Is that the right way to think about it?
Correct. So in the same way that you would see somebody that you know, and your visual cortex would process the face, and then with your temporal lobes, you would access the kind of face data stored in your brain and figure out, oh, that's my cousin, that's happening in the vector embedding space. That's happening with the similarity search. That's happening with a mechanism that looks a lot more like memory than it looks like processing like a neural net.
So let's say I have video data, audio data, text data. So if I understand correctly, I need to translate this into numbers to fit it into a machine learning model, an AI model. How do I do that? How do I take video and translate that into numbers?
We can talk about the mechanics of how to do it, which is maybe the less interesting thing because the less technically inclined of you don't care and the more technically inclined of you can figure it out. If you zoom out, though, the interesting part of it is that I think there is a significant shift in the way that people think about building these smart AI systems, and they don't try to fit the data into the model as much as give the model access to the right data when it's doing the inference.
And so, I gave this example several times, but if you study medicine, you already go into medical school knowing English. And then you study a body of knowledge. You study medicine. You don't study medical English. You read a book and then know medicine, and then you speak about it because you know how to speak, right? You don't become a doctor by just doing the rounds for 10 years and then know enough, know how to sound like a doctor.
So that's the answer in some sense. You don't fit it. You allow the model to access external memory and use it correctly. And you give it the right data. So if you want it to be a doctor, you give it medical knowledge. If you want it to use videos, you give it video data.
And that is vectors and vector embeddings. And that leads us to vector databases. So we'll talk about Pinecone in a second, but what is a vector database in that context?
A vector database is the infrastructure that supports that long-term memory. So if you have a very small amount of data, you can fit it on the machine and you can kind of play around with it. That already is becoming almost unnecessary with these larger contexts and so on. But when you have millions or billions of documents or images and so on, you really have to have a very specialized system to be able to do this cost-effectively, with low latency, high persistence, and so on.
You really have to start thinking about things in a fundamental way and not just about the data.
And for anybody that spends time in the MLOps world, in the last couple of years before vector databases became a thing that everybody's talking about, the concept of storing stuff that would be fed into machine learning models, people heard the term maybe feature store. Is that completely different? Is that part of the same family? How would you think about it? Yes.
No, I'm just kidding. Yeah, no, it's a completely different thing, right? Feature stores are used mainly either for marshaling data into training or using real-time features for objects that change very frequently, right? So, if I'm a user and I just clicked on something, maybe you want to incorporate that last feature that just happened a second ago in my next classification, in my next action or something. Vector databases are completely different. It's really about this long-term memory, about these billions of embeddings of documents or images and so on, and making them available to large language models, all these other multimodal models and so on, where they operate more like parts of your brain than like a feature store.
Okay, fantastic. Maybe one last concept, and thank you for explaining all of this so clearly. Semantic search is something that, again, people hear a bunch. What does that mean?
So semantic just means by meaning. And search means search. Traditional search is keyword-based, not because keyword search is that great, but just because we have great mechanisms to run it at scale and it's very efficient. Keyword search is a pretty old technique. In fact, books had inverted indexes in the backs of them. They're even called an index. They've had them for a while. In fact, you had indexes in books before we had print.
So that's a pretty old—well, before we had print in the West. The Chinese were printing well before the Gutenberg Bibles were even thought of. But anyway, it's like early 1200s at least. So keyword search existed for a while. Great technique. Works great for Google, works great in other places, doesn't work when you want to search something by meaning, right? When you remember a conversation that you had, I don't know if you've ever searched your inbox for an email you know for a fact you read, right?
And the only way to do it is try to hack the search system to figure out what word was there that I didn't use in another email. Like, you try to backwards-engineer the crappy search. With semantic search, the whole idea is that you know what it means. You should be able to search by the meaning in free text, and the match would not be because, oh, like three words matched and they're next to each other, but rather the meaning is similar, and so I can retrieve that.
And again, that looks a lot more like how humans remember things, and not like how inverted keyword lookup stores look.
Okay, so all of this is a wonderful, hopefully, intro into Pinecone and what Pinecone does. Tell us about the company, the history, maybe your background, how you came to start this company, especially because you started it, I believe, in 2019, which was well in advance of this whole craziness right now?
So, I did my PhD in theoretical machine learning and big data algorithms, and my postdoc in applied math. And that was at Yale, right?
And already then I was working on high-dimensional geometry and functional analysis and kind of the foundational basis of machine learning. And I was always drawn into this search problem because it seems especially gnarly. I then did two internships at Google on that topic. I joined Yahoo as a scientist. I became an adjunct professor at Tel Aviv University. And through this entire time, I kept publishing and working in industry on infrastructure for big data and for machine learning. Stayed at Yahoo for about seven years, which moved me back to New York after a short stint in Israel.
And then in 2016, Swami Sivasubramanian, who now runs all databases and AI services in AWS, back then he wasn't a VP yet; he was only a director. But he called me up and he said, "Hey, we're starting this AI thing in AWS. Do you want to come?" And I said, "Okay, who's we?" And he said, "Me and Alex." It was Alex Smola, who is—I don't know if you know him. Look him up.
And yeah, that was a journey I couldn't resist, like come to AWS and basically build SageMaker and all these other great services out of—
It's huge now. There must be thousands of people now, right?
Several thousand now, yeah. But I'm saying this, I'm kind of recounting the whole story because throughout that time, vector search and vector representations and embeddings and that kind of basic concept kept getting more, gradually more and more traction all the time. More and more people kind of understood what it does and what's happening with it. And then at some point, the BERT embeddings, kind of the transformer models, kicked in, and then there was a big spike in understanding. And then that's it.
In 2019, in some sense, I felt like that was the right time. I felt like it was going to get us about two years to kind of build the right kind of infrastructure correctly. And we kind of timed it correctly. And so this is what happened. We didn't, by the way, foresee any of this ChatGPT thing happening. We knew it would keep growing. But that, I think, completely took everybody by surprise, including us.
Yeah. So I can only imagine how wild a time it must have been at Pinecone over the last few months. And so vector databases are becoming very much a central part of what seems to be emerging as the AI stack. So, when people talk about, like, how do you do this whole thing? Vector database is the term that everybody comes back to. What does a stack look like, I guess, based on what you see customers do?
This vector database, this LangChain, there's different models. How does that all fit? What does that look like now?
Yeah, so there is a—I mean, it depends what you call—I’ll focus very narrowly on the large language model stack. Of course, there are the models themselves. There are vector databases. There are libraries that connect them in all sorts of different ways, like LlamaIndex and LangChain and others. There are, I think, a lot of connective services, but I think we're going to see an emergence of agents in some form or another. I think that's still very—
Do you want to briefly explain what agents are?
So agents are these recursive pieces of software. I don't even know how to define them. It's sort of like this life hack on large language models. It's like, how do you use these large language models, enable—like, let them use tools like search or different APIs and so on—and then use their own answers to somehow feed back into an input and try to accomplish very complicated tasks, right? So it's not like a one-time prompt.
Like, if you want to get in shape, right? I mean, that's a plan. You need to fit it into a schedule. You have to find a gym. You have to whatever. There's a sequence of things you need to do, and you can research every one of them separately. And you might think as a human, it wouldn't be like a language generation thing. It would be like you sit there and plan it. So these agents are now coming up as this new paradigm of building more complex sequences and plans.
Yeah, I put an asterisk next to that. I think this will shake up some surprising ways in the next few months. We'll see.
Great. Let's talk about some of the use cases that you've seen with Pinecone. Maybe if you can use some customer examples, but as I was prepping for this, search, generation, security, personalization, data management—pick a handful of those and sort of double-click on how people use Pinecone.
Yeah, so first of all, one of them is obvious semantic search. Like, people are literally just reinventing their own search stack. I don't know if you ever tried to truly optimize an Elasticsearch cluster. I mean, and after three weeks wanted to throw your keyboard out the window, right? So with large language models, a lot of that is sidestepped and improved in a significant way. And people use Pinecone to just drive the backend search and storage based on these vector embeddings rather than keywords.
But that's very obvious and immediate. Nowadays, we see a huge wave of people using us to create context for ChatGPT-like and Bard-like applications. So call them chat as a whole. But chat is very general. It could be customer service. It could be—you name it. But there is a very wide variety of use cases, from anomaly detection to—one of my favorite applications is somebody built face detection for cows. And we're like, I didn't even know you can face-detect cows.
I grew up in a city. I didn't even think cows looked different. So, I mean, but then I told this to a farmer and they got deeply offended. So I'm probably offending some people here. I apologize.
Maybe that's related to the cows question. But one really important part of vector databases is because they are the long-term memory, as you were saying, they are a very important tool against that problem of hallucination that people talk about a lot. Is that the right way to think about it?
100%. In fact, I was looking at experiments today from one of our teams, and we literally measure reduction in hallucination as one of the core metrics that we try to drive. So, 100%. And by the way, I mean, this is—it's almost like, obviously, if I ask you a question for which you don't know the answer, but I compel you to say something anyway, you're going to make something up. But if I give you the right context, you can answer a lot more accurately.
And that's exactly what's happening. And so the question is, can we retrieve the right context out of which an informative answer can come out? And so, in some sense, it's obvious that that should happen, but it's harder to actually get the whole thing to cascade correctly and to work measurably well.
So very practically, I'm a company. I want to deploy GPT-4 in the enterprise, but for a mission-critical application like chat, where I don't want it to say stupid stuff to my customers, I would use Pinecone. I would put basically my source of truth—this is like the history of customer conversations—into Pinecone. And then how do I make GPT go fetch the information in Pinecone?
When you want to query, when you have a question, say, or some prompt, you want to add context to it. The way you would do it is you take that prompt, you would embed it with some language model, search in Pinecone for the context, which could be maybe the top, I don't know, 100 relevant documents or 100 relevant parts of sentences or paragraphs, and put that back with the context and say, "This is the prompt, and here is more context out of which you can give an answer."
And that ends up being—even when you do that relatively crudely, because you can do what I said right now in like 1,000 different ways. You can ask, which language model should I use? And how should I parse the data? And how should I chunk it? And why should I put it? And what's the order? You can play with it in a million different ways. But even if you do it relatively naively, that already gives you a huge bump in accuracy or reduction in hallucination, depending on how you want to measure it.
Let's see. Look, I'm a VC. I'm always interested in that part as well. Let's talk about go-to-market. Like, how are you building the company? How are you finding customers? How are you selling to them?
So we started Pinecone, or I started Pinecone, as a managed service. Because I'm a really big believer in the buyer-driven journey, right? I think in the modern world, if you want to use some technology, the last thing you want to do is send an email and wait for three days. That sounds very unnatural. You want to use it immediately, right? And if it's great, you want to just start using it and forget about it.
Be happy and move on. So we really built the whole company around the self-serve journey. Today you can start with Pinecone, run some example in five minutes, build a demo app in two hours, integrate some kind of a micro-POC for your team in a day, and in a week be in production and never talk to us once, if you choose not to.
Some companies today—very large companies—tend to also want to talk to us. And so we're now building a sales team that is happy and willing to engage and help and so on. But it's more about assisting in the journey if that's wanted, or kind of standing, like, not being in the way when we're not wanted.
All right, last question from me. I'm curious about the space in general. It's become what feels like a very sort of rich and vibrant slash competitive space. Very, very—
The space, the vector database space. And I'm curious, less in the context as a prompt to badmouth competitors, but more as a way for us to understand the space. Like, all those companies that are vector databases now, are they very different approaches, or are you all sort of going in the same direction and it's more of an execution kind of play to win the market?
Yeah, I'll try not to badmouth anyone. I'll just say stuff about us, and you can extrapolate.
Oh, feel free. This is recorded, but, like, feel free.
Yeah. Yeah. Only YouTube will know about this. Yeah. I mean, look, I'll tell you what's driving me and my team, okay? And first of all, it's the technology itself, right? I mean, the internals of these databases are actually incredibly complicated numerical data structures. And really, it's very cutting-edge technology, both on the distributed system, the data management, the persistence. We had to rebuild our own object storage. Like, we basically had to build the whole stack, from the numerical indexes to the object persistence to the distributed data to the query planner, the whole thing.
So as a systems and algorithms person, Pinecone is a fascinating place to be. And I can tell you that we are far and away better in terms of efficiency of running this service at scale and its stability. And we are just getting started. Like, there's so much to come. Okay, so that's number one. And the second thing is, having spent a lot of time in AWS, I got indoctrinated a little bit to be customer-obsessed rather than competitor-obsessed.
And we are delighted and happy that we have a lot of very demanding and enthusiastic customers. And so we're very focused on what they want rather than what our competition is doing.
Yeah. And by the way, just as a last word, as I was prepping for this, I read somewhere that you guys got like 10K signups a day.
That's insane. Anyway, that's so exciting. That feels actually like a wonderful place to leave it. Thank you so much. This was super, super interesting. It's really incredible how Pinecone has just vaulted to the forefront of consciousness in this world of generative AI. So congratulations on all the incredible work that you guys have been doing.
Thank you. Thanks for listening to The MAD Podcast. If you liked this episode, be sure to leave us a review. com/events/data-driven.