And that is really just what the transformer architecture is. So I would say, and maybe I'm biased because one of my best friends is on the original attention paper, but that was the real breakthrough. It's just like figuring out that you have this attention mechanism that actually allows you to do a much better job at sort of representation learning and, as a result, kind of generating correct text sequences autoregressively.
4.0 version, and agentic RAG and what Contextual AI does. But maybe as an introduction/deep dive, how does that work? So there's a retriever, there's a generator, all those good things. What's the core architecture?
Yeah, the core architecture is very simple. You have a language model, so that's the G, and then you want to give that context. And the way you do that is by augmenting it, the A, using retrieval, the R. So that's RAG. And so how you do the retrieval, that has been changing constantly over time. So in the initial paper, we used a vector database, or FAISS. The words "vector database" didn't exist at the time, so FAISS was the first vector database.
And I think people over time have started figuring out that that has all kinds of limitations. So how a vector database works is you just have embeddings. So you encode pieces of information or chunks of documents, you encode them as a vector, and then you do basic dot-product similarity search. But that has issues where you're just looking for chunks that are similar to the question, but you don't necessarily want to find chunks that are similar to the question. You want to find chunks that are relevant to answering the question.
So you need to do different things with the representations, with the embeddings. So a lot of modern RAG deployments are very different from the original ideas in the paper. So you still have a vector database, but the way you encode things is very different. Then you usually also have a sparse component. So it's not just dense search, it's not just vector search. You're also doing even older traditional keyword-based search, so BM25 or TF-IDF-style algorithms.
Or Elasticsearch, or that kind of stuff.
Yeah. So that's what Elasticsearch is very good at, right? And so then you have this hybrid retrieval system, and you can cast a relatively wide net with that. So it's very cheap to do that, but then it's not going to be great. So on top of that, you need to have a re-ranker that actually filters out most of the stuff that you probably shouldn't have retrieved in the first place. So you get this kind of cascade of retrievals where you narrow down the search, and the model that does the decision-making gets smarter and smarter the further in the cascade you go.
So the final re-ranker step, that can be a pretty beefy model where you can even, in our case, give it instructions, which is awesome, right? So you can really tell it, like, I believe this source much more than this source, and I have a strong preference for recency. And if it's a PDF, then I believe it much more than if it's our internal Slack or something, right? So that type of ranking really factors into the overall retrieval results. And then the final step is that goes to a language model.
And hopefully that language model doesn't hallucinate, but you have no guarantees. So that's kind of how the original naive RAG evolved into sort of advanced RAG, where you have this pretty sophisticated pipeline and it's really a system, right? These are all different models. And so one of the—
To unpack some of this at a very practical level, then what goes back to the model is in the form of prompts. It pushes information into the context window.
Yeah, exactly. So your retrieval results, the context that gets given to the language model as a part of the prompt, it's like, this is the question. And then you go off and ask the question to your retrieval system, you get the results, you put them into the prompt, and then you ask the language model to do its magic.
And for the hybrid search, how do you, or how does the system decide when to do a term search versus a vector search? Is there presumably intelligence there in allocating certain types of query to certain—
In most systems, there is no intelligence there. So in a lot of advanced RAG, that's sort of a hyperparameter that you just tune. So in our case, that is learned. And so it is something that you can learn, but we take a very different approach. So we kind of started from this observation that RAG is really not about models, it's about the system of models. And so the real question is, how do you get all these models to work together in the right way?
So in our case, each of those components, right? So the language model has been trained to be grounded. So it's state-of-the-art at grounded generation. Then we have a re-ranker that has been specifically trained to work together well with that grounded language model. Then we have our retrieval step before that, right? And then all of these components are jointly optimized, so they're trained on the same data distribution, and that means that they have been trained to work well together. And you can really see the difference in benchmarks in terms of performance when you do that.