So let's take—that's super interesting. Let's take those two data platforms. Let's talk about the first one first, and then we'll go into the employee-focused one and the internal-focused one. So, taking the first one, does it have a name, presumably?
We call it ZDP, ZoomInfo Data Platform.
Okay. Yeah, internal name is ZDP.
Okay, ZDP. All right, so ZDP, which is the platform that powers the whole business, the customer interaction and all the things. So presumably it ingests massive amounts—not presumably, that's well documented—but it ingests, like, billions and billions of data points into it. So how does that part work? Where do you get data from, all the sources? And how does that get selected, funneled into the platform, cleaned or cleansed? Walk us through the whole pipeline.
I mean, with respect to, let's say, you're going to get first-party data. If first-party data is, let's say, in Salesforce or HubSpot and platforms like that, then we already have these pipes set up for you. So that's one way for us to get data. With respect to third-party data, we are definitely continuously crawling company websites, merger and acquisition data, financial data, SEC data, all kinds of information that are about people, including news that's coming out of those. That's also continuously happening.
So all the public sources, basically, we will have it. We have a contributor network. For example, if you're using ZoomInfo Lite, then as part of that free access, we get some data from our customers.
What kind of data do I contribute if I'm a ZoomInfo Lite customer? Like, what do I commit to?
Information about your company, your email or phone number, things of that nature. If you're also tied into, let's say, your website traffic, then we can also use that to identify why people are visiting your site.
Yeah, I love those business models. There's not that many actually around, but the quid pro quo is that, just to play it back, I give you my email address, but in return for that, I can use the platform for a lower price.
Yeah. Also, these are not personal; it's all business-related. So our context is these are all needed to make a sale, for you to find customers, to retain your customers, as well as your potential business interactions that you're going to have, just to help for that. So data is completely specified for that.
And then, other than my internal data as a customer, do you presumably have access to more private, proprietary kind of data sources as well as part of the platform, or do you only do publicly available?
Publicly available, but customers definitely have first-party data, right? So some companies, let's say they're using Copilot. As a result of that, they want to actually combine their first-party data with our data. Let's say they have a set of contacts in these companies, and the profiles for each of these contacts that they have may not be complete. So they want to use, let's say, our data and tech to enrich that, to either correct these, to remove duplicates and all that, but on top of it, maybe to get additional attributes that they can use for that purpose.
So for that, we can get that data on their behalf, completely dedicated to them as a single tenant. We combine that so we can extract additional signals for them for selling.
So you crawl all of that, you combine it with some internal data. What happens then? How do you clean and normalize everything?
Depends on the type of the data, definitely. If you're, let's say, getting public data from public sources, first of all, you got to crawl it, right? You got to crawl it at a certain frequency so that you are going to get all the changes that are relevant for our purposes. You have to put the right tech over there so the data that you are extracting from those sites will be correct. We also have a good number of researchers, about 300 of them.
If there are specific cases that we need a human touch for them to actually correct the data, it's very difficult to find, definitely in certain cases, to verify that this is accurate data. Then also people might be involved in the loop as part of that.
Is that a machine learning thing? Do you have a system that flags that the data may not be correct?
Almost for every attribute that we are getting, we have a machine learning algorithm or model to tell us how to interpret it, whether it's correct or not, what kind of anomalies we should be sending over there.
And that's homegrown? That's stuff you developed yourself?
Yeah, we use some crawlers we might be using from third parties, but many of these components are homegrown, yes.
And so you have all of that, and then you funnel it into the platform. Or I guess maybe you do all of this once it's into the platform already. And you said the platform is Google for that?
Yeah, ZDP is built on Google BigQuery right now. So all these pipes are continuously pushing the data to the platform, or sometimes we are pulling the data. Then you have to definitely store it on a platform like that, and then you got to do processing on top of it, right? So you have to sometimes join this data or extract certain attributes for each of these, create a graph because these are related to each other. A person works in a company; this company maybe sells a product to another company.
All these connections are there that we are extracting out of that, and that is done through different technologies, either Spark code or Dataflow that GCP has, all kinds of basic processing. Machine learning, definitely, right? Model building. Some of these are just traditional machine learning. Let's say you're going to have a tree-based model that uses that data to create, I don't know, classification or prediction. Also GenAI, right? Using that data, either create an additional context for GenAI, also just to make sure that you can prevent almost hallucination, saying that, do not go beyond that.
Here is your data and try to extract insights out of that. So all that processing is done over there.
And while we're on that phase, in addition to the tools that you described, do you have sort of data infra, I would say, modern data stack kind of tools for, I don't know, cataloging? And if you can talk about it, what do you all use?
So we're using some third parties on that. And part of the catalog, we also do internally, so it's homegrown, to make sure that we can build it in the way that we can. It's not a single system, right? We have a cataloging solution that we are using from a third party, and there's a version of that internally. And there are different cataloging systems, component cataloging, for example, pipeline cataloging. And these are also connected to each other so that they're almost like a unified source for whatever you're looking for.
It's accessed through that. And with respect to processing on top of it, catalog is telling you exactly where things are, a single source on that one, all the metadata. Then how are we going to actually create these schemas in such a way that they themselves are in a way that you can process them very easily, right? So creating all the readers and writers automatically, writing the processing in such a way that whoever is going to write a processing engine internally, I don't want them to literally go that deep into every schema, every data type, and it should be very unified in a way that we may store the data in, let's say, one or more tables, but they should see it as a unified, let's say, 360 view as a person or as a company, right?
So that it's a logical representation, and it gets mapped to a physical representation underneath. And through the cataloging system, actually, you have access to that one, right, with respect to attributes and how the schema looks, owners, who are the owners, right, what kind of updates that we have on top of it. That logical layer, actually, as well as an API on top of it, simplifies basically application writing and creates a very fast turnaround.
So you got the data from the sources, it's all in the data warehouse. You were talking about the ecosystem around the warehouse and the various tools you use. And then the next step after that is that you push the data to the applications. And again, the applications can be SalesOS, MarketingOS. Those are the applications.
Those applications are using additional platform components. For example, search. If you go to SalesOS, there's lots of filters: companies of that size, that revenue, people of that title, and intent of those categories. So you do that search, which means the data that we just extracted should be pushed to search indexes. So we are using Solr internally. There's also some Elastic used for different purposes. If you're going to join the data, then they go through what we call entity resolution.
So there's a separate pipeline for that one because you cannot just perform a regular join. As you know, that join is pretty involved, and lots of machine learning and rule engines are used for that. So that's one place. For example, let's say you need access to—you find what you're looking for through search, but you need more data about that key that you just found out—then we have a separate system to give you additional information about that entity, whatever that is, let's say company or contact, right?
So all the right places it has to go. Then through the APIs, our products are accessing those. So it's also unified in that sense. So there's no need for creating a separate search for talent versus sales, because it's a big investment. It's very real-time indexing, real-time search, very fast, like milliseconds. It's also expensive to run something of this scale. It's already elastic. So we try to basically push it into the right platform foundation components so products can access it very easily.