MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    AI at ZoomInfo: Superpowering GTM teams | Ali Dasdan, CTO, ZoomInfo

    Ali Dasdan is the CTO at ZoomInfo. We cover how ZoomInfo combines customer first-party data with third-party data in single-tenant environments, why Copilot tells sellers which accounts and signals to prioritize while users report saving 10 hours a week, and why the company uses external models for hundreds of emails but self-hosted Llama for 100 million website-scale tasks.

    09/19/2024

    Hosted by Matt Turck · with Ali Dasdan, CTO, ZoomInfo

    GTM AIdata platformssales CopilotLLM infrastructureZoomInfo
    Listen now
    YouTubeApple PodcastsSpotify
    46 min · 11 chapters
    Contents

    Transcript

    What is ZoomInfo

    2:03
    Matt Turck0:58

    Welcome, Ali.

    Ali Dasdan1:00

    Thank you for having me.

    Matt Turck1:28

    So we are going to do a deep dive into all things data and AI at ZoomInfo. But maybe let's start by setting the stage and talk about ZoomInfo, the company itself, what you do, all those good things. So maybe to start with some quick stats. ZoomInfo is a public company trading on the NASDAQ, $2 billion in annualized revenue as of Q2 2024. I should note that unlike some other tech companies, you are very profitable, with a 28% operating margin, more than 35,000 customers worldwide, 3,500 employees worldwide.

    Matt Turck2:01

    So those are the stats. But first and foremost, tell us what the business does. ZoomInfo is one of those companies that everybody that's sort of in the B2B business knows, but maybe the broad public doesn't. So what does the company do?

    Ali Dasdan2:31

    Yeah, so we help companies to sell better, if I just summarize it that way. So the personas that we work with are salespeople and marketing people, and every company, in principle, is our customer, potential customer. And we want to make sure that the sellers can not only find new customers very easily, but also retain their existing customers. So we provide them the data, insights, and signals to achieve that.

    Matt Turck2:51

    So if I'm a B2B software seller and I'm looking for my next 100 customers, you'll provide insights, perhaps alerts, and then all sorts of information about that company that I can then reach out to.

    Ali Dasdan3:17

    Yeah. So you're going to come to, let's say, our product. We also notify you. You don't have to come, but let's say you're looking at our user interface, and you can do searches for anything you care about. Companies, let's say you're selling chairs: which companies are in the market right now for buying chairs, right? Or if these are existing customers, definitely you can figure out whether they are researching, they are doing similar searches. Once you find out that this company is a potential customer for you, then we provide information about this company.

    Ali Dasdan3:47

    If it's a startup, let's say whether they got funding recently or not, you can get information about the organizational hierarchy. So who are the people who are going to be decision-makers for that? All the information about them, how do you reach them? We have the mechanisms for you to reach them using, again, AI. We can actually create even the email to send, which is personalized for that use case. So fundamentally, basically not only finding the right customers and right people, but reaching out to them and making it very easy for you actually to make the sale.

    Matt Turck4:07

    And you target salespeople, but also marketing people, talent people, operating people. The platform expanded quite a bit.

    Ali Dasdan4:28

    Expanded, yeah. So since we have a very strong data foundation, we have lots of information about companies and business people in these companies. As a result, you can literally use that information for multiple purposes. Also, tying back to, since I'm going to help you sell better, marketing is part of that flow, part of that funnel. So that's one of the aspects of it. And also, the fact that you have a basis like that, let's say you want to find the best people from a business context, and we have a TalentOS product that you can actually use for that.

    Data as service

    4:47
    Ali Dasdan4:47

    And OperationsOS product is to help you clean, for example, your data, to remove the duplicates, to make sure that it is ready for you to go.

    Matt Turck5:02

    So you do all the things as an application, but as I was prepping for this, you also have a data-as-a-service business. I don't know how big each one is, but what's that part, the data-as-a-service business?

    Ali Dasdan5:29

    Yeah, data as a service, we call it DaaS, is one of those businesses. So some companies, even though they are customers of ours, they have licenses, seats that their sales teams are using our product. But sometimes, let's say they want to get a copy of that data and enrich within their own platform. Let's say they have their own first-party data. Our platform enables you to bring first-party data. We have lots of third-party data. We can actually combine this on your behalf very easily.

    Ali Dasdan5:46

    But at the same time, some companies prefer that. And as a result of that, you can just get a copy of this data through multiple platforms. You can go to AWS, Databricks, Snowflake. It's ready to go for you. It's in their marketplaces.

    Matt Turck5:49

    So it's an API, and you'll feed the data.

    Ali Dasdan6:10

    You can get it through API, you can get it as a file, or you can get it ready to go in the marketplaces for these companies, Databricks and Snowflake and GCP and AWS, and it's ready to merge for you. Basically, you can just do that processing over there. So either way, basically, my point is, our point is not to create any friction. The point is to help you make the sale, to make your sellers better salespeople. And as a result, if you need the data, we also have a capability like that.

    Ali Dasdan's story

    6:15
    Matt Turck6:26

    Maybe to wrap this intro, a quick word about you and your background. Maybe walk us through your story. So you were at Atlassian right before ZoomInfo. What was your journey?

    Ali Dasdan6:35

    Yeah, I mean, originally came to the U.S. for a PhD, worked in lots of small startups as well as big companies.

    Matt Turck6:36

    What was your PhD in?

    Ali Dasdan7:08

    My PhD was mostly in graph theory and algorithms. It was for doing timing analysis for embedded real-time systems. So it had applications to chip design software, electronic design automation. That was the reason my first company was Synopsys, which is one well-known company in this space. And after that, Yahoo Web Search, did lots of big data and data science and advertising and search over there. eBay for e-commerce. And I did real-time advertising in a startup. I worked in the UK for two years in retail, but built their data platform and marketing platform.

    Organization of ZoomInfo

    7:31
    Ali Dasdan7:31

    Then Atlassian in the productivity space. I was leading one of the three business units over there, which was Confluence, Trello, Jira Work Management, and some of these products were part of my team. And then after four years over there, I moved to ZoomInfo.

    Matt Turck7:48

    What does the engineering organization look like today in terms of size across—I don't know, we'll get into how you organize—but across engineering, data, AI, what's the total number of people?

    Ali Dasdan8:12

    It's roughly over 900 people, close to 1,000 people. So the way we are organized, maybe we're going to come back to it. Like you're saying, definitely, we have a strong platform team, and data platform is part of that. Data platform covers data engineering, data science, and the data platform by itself. There's a separate enterprise engineering team. We call it the Enterprise Productivity Team because we care about productivity a lot. And that deals with serving the internal customers: revenue team, finance team, HR team, and all that.

    Ali Dasdan8:34

    Definitely operations team, site reliability, security teams, IT teams. And on top of the platform, we have multiple products. These product teams you mentioned: sales product, marketing product, operations, talent. We have Chorus for conversational intelligence and a couple more.

    Matt Turck9:05

    Okay, so to dive into the nitty-gritty a little bit of how that all works. I've been looking forward to the conversation because it's talking data with somebody who runs data at a data company. So it's very meta in some ways. So to start from the start, I guess there is a core data platform which is a foundation for everything. Is that correct? And if so, how does that work?

    Ali Dasdan9:30

    So the way it is structured is, definitely, we have a cloud infrastructure team underneath. We are using both Google Cloud as well as Amazon Web Services. And on top of it, we have actually two data platforms. One is for all the data that customers need and our products need. And there's another one that's internally focused for our internal customers. And they're built on different technologies. There's a pipe between them just to make sure that data is exchanged in real time.

    Ali Dasdan10:03

    Then we have teams that are developing applications on top of these data platforms, including data science applications, creating machine learning models using GenAI, using that data. And the fact that we have the data in one place makes it super easy to do all of that. And also, data is huge in our case. We are the market leader in that with respect to company data, intent data, contact data, first-party data, third-party data. Then we have on top of it what we call product foundations teams.

    Ali Dasdan10:34

    For example, search team, identity team, admin hub team that deal with settings, permissions, and all that, platform UI team, for example. And all the teams are what we call product foundations. And then on top of it, now we have the product teams that are interacting with those. And the idea is basically to use these platform components and use the data that is already there so we can build products and features very fast. So through the APIs, third parties can also do that.

    Ali Dasdan10:46

    But more and more, we are trying to also make it easy for third parties to build new experiences, new user journeys on top of our platform.

    ZoomInfo Data Platform

    10:48
    Matt Turck11:02

    So let's take—that's super interesting. Let's take those two data platforms. Let's talk about the first one first, and then we'll go into the employee-focused one and the internal-focused one. So, taking the first one, does it have a name, presumably?

    Ali Dasdan11:05

    We call it ZDP, ZoomInfo Data Platform.

    Matt Turck11:06

    ZoomInfo Data Platform.

    Ali Dasdan11:07

    Okay. Yeah, internal name is ZDP.

    Matt Turck11:43

    Okay, ZDP. All right, so ZDP, which is the platform that powers the whole business, the customer interaction and all the things. So presumably it ingests massive amounts—not presumably, that's well documented—but it ingests, like, billions and billions of data points into it. So how does that part work? Where do you get data from, all the sources? And how does that get selected, funneled into the platform, cleaned or cleansed? Walk us through the whole pipeline.

    Ali Dasdan12:13

    I mean, with respect to, let's say, you're going to get first-party data. If first-party data is, let's say, in Salesforce or HubSpot and platforms like that, then we already have these pipes set up for you. So that's one way for us to get data. With respect to third-party data, we are definitely continuously crawling company websites, merger and acquisition data, financial data, SEC data, all kinds of information that are about people, including news that's coming out of those. That's also continuously happening.

    Ali Dasdan12:28

    So all the public sources, basically, we will have it. We have a contributor network. For example, if you're using ZoomInfo Lite, then as part of that free access, we get some data from our customers.

    Matt Turck12:34

    What kind of data do I contribute if I'm a ZoomInfo Lite customer? Like, what do I commit to?

    Ali Dasdan12:45

    Information about your company, your email or phone number, things of that nature. If you're also tied into, let's say, your website traffic, then we can also use that to identify why people are visiting your site.

    Matt Turck12:59

    Yeah, I love those business models. There's not that many actually around, but the quid pro quo is that, just to play it back, I give you my email address, but in return for that, I can use the platform for a lower price.

    Ali Dasdan13:18

    Yeah. Also, these are not personal; it's all business-related. So our context is these are all needed to make a sale, for you to find customers, to retain your customers, as well as your potential business interactions that you're going to have, just to help for that. So data is completely specified for that.

    Matt Turck13:34

    And then, other than my internal data as a customer, do you presumably have access to more private, proprietary kind of data sources as well as part of the platform, or do you only do publicly available?

    Ali Dasdan13:55

    Publicly available, but customers definitely have first-party data, right? So some companies, let's say they're using Copilot. As a result of that, they want to actually combine their first-party data with our data. Let's say they have a set of contacts in these companies, and the profiles for each of these contacts that they have may not be complete. So they want to use, let's say, our data and tech to enrich that, to either correct these, to remove duplicates and all that, but on top of it, maybe to get additional attributes that they can use for that purpose.

    Ali Dasdan14:19

    So for that, we can get that data on their behalf, completely dedicated to them as a single tenant. We combine that so we can extract additional signals for them for selling.

    Matt Turck14:30

    So you crawl all of that, you combine it with some internal data. What happens then? How do you clean and normalize everything?

    Ali Dasdan14:53

    Depends on the type of the data, definitely. If you're, let's say, getting public data from public sources, first of all, you got to crawl it, right? You got to crawl it at a certain frequency so that you are going to get all the changes that are relevant for our purposes. You have to put the right tech over there so the data that you are extracting from those sites will be correct. We also have a good number of researchers, about 300 of them.

    Ali Dasdan15:13

    If there are specific cases that we need a human touch for them to actually correct the data, it's very difficult to find, definitely in certain cases, to verify that this is accurate data. Then also people might be involved in the loop as part of that.

    Matt Turck15:18

    Is that a machine learning thing? Do you have a system that flags that the data may not be correct?

    Ali Dasdan15:31

    Almost for every attribute that we are getting, we have a machine learning algorithm or model to tell us how to interpret it, whether it's correct or not, what kind of anomalies we should be sending over there.

    Matt Turck15:34

    And that's homegrown? That's stuff you developed yourself?

    Ali Dasdan15:41

    Yeah, we use some crawlers we might be using from third parties, but many of these components are homegrown, yes.

    Matt Turck15:56

    And so you have all of that, and then you funnel it into the platform. Or I guess maybe you do all of this once it's into the platform already. And you said the platform is Google for that?

    Ali Dasdan16:19

    Yeah, ZDP is built on Google BigQuery right now. So all these pipes are continuously pushing the data to the platform, or sometimes we are pulling the data. Then you have to definitely store it on a platform like that, and then you got to do processing on top of it, right? So you have to sometimes join this data or extract certain attributes for each of these, create a graph because these are related to each other. A person works in a company; this company maybe sells a product to another company.

    Ali Dasdan16:47

    All these connections are there that we are extracting out of that, and that is done through different technologies, either Spark code or Dataflow that GCP has, all kinds of basic processing. Machine learning, definitely, right? Model building. Some of these are just traditional machine learning. Let's say you're going to have a tree-based model that uses that data to create, I don't know, classification or prediction. Also GenAI, right? Using that data, either create an additional context for GenAI, also just to make sure that you can prevent almost hallucination, saying that, do not go beyond that.

    Ali Dasdan17:01

    Here is your data and try to extract insights out of that. So all that processing is done over there.

    Matt Turck17:20

    And while we're on that phase, in addition to the tools that you described, do you have sort of data infra, I would say, modern data stack kind of tools for, I don't know, cataloging? And if you can talk about it, what do you all use?

    Ali Dasdan17:44

    So we're using some third parties on that. And part of the catalog, we also do internally, so it's homegrown, to make sure that we can build it in the way that we can. It's not a single system, right? We have a cataloging solution that we are using from a third party, and there's a version of that internally. And there are different cataloging systems, component cataloging, for example, pipeline cataloging. And these are also connected to each other so that they're almost like a unified source for whatever you're looking for.

    Ali Dasdan18:10

    It's accessed through that. And with respect to processing on top of it, catalog is telling you exactly where things are, a single source on that one, all the metadata. Then how are we going to actually create these schemas in such a way that they themselves are in a way that you can process them very easily, right? So creating all the readers and writers automatically, writing the processing in such a way that whoever is going to write a processing engine internally, I don't want them to literally go that deep into every schema, every data type, and it should be very unified in a way that we may store the data in, let's say, one or more tables, but they should see it as a unified, let's say, 360 view as a person or as a company, right?

    Ali Dasdan19:03

    So that it's a logical representation, and it gets mapped to a physical representation underneath. And through the cataloging system, actually, you have access to that one, right, with respect to attributes and how the schema looks, owners, who are the owners, right, what kind of updates that we have on top of it. That logical layer, actually, as well as an API on top of it, simplifies basically application writing and creates a very fast turnaround.

    Matt Turck19:24

    So you got the data from the sources, it's all in the data warehouse. You were talking about the ecosystem around the warehouse and the various tools you use. And then the next step after that is that you push the data to the applications. And again, the applications can be SalesOS, MarketingOS. Those are the applications.

    Ali Dasdan19:49

    Those applications are using additional platform components. For example, search. If you go to SalesOS, there's lots of filters: companies of that size, that revenue, people of that title, and intent of those categories. So you do that search, which means the data that we just extracted should be pushed to search indexes. So we are using Solr internally. There's also some Elastic used for different purposes. If you're going to join the data, then they go through what we call entity resolution.

    Ali Dasdan20:07

    So there's a separate pipeline for that one because you cannot just perform a regular join. As you know, that join is pretty involved, and lots of machine learning and rule engines are used for that. So that's one place. For example, let's say you need access to—you find what you're looking for through search, but you need more data about that key that you just found out—then we have a separate system to give you additional information about that entity, whatever that is, let's say company or contact, right?

    Ali Dasdan20:51

    So all the right places it has to go. Then through the APIs, our products are accessing those. So it's also unified in that sense. So there's no need for creating a separate search for talent versus sales, because it's a big investment. It's very real-time indexing, real-time search, very fast, like milliseconds. It's also expensive to run something of this scale. It's already elastic. So we try to basically push it into the right platform foundation components so products can access it very easily.

    Lessons from building a data platform

    21:02
    Ali Dasdan21:02

    And even APIs, external APIs, are supported through those systems.

    Matt Turck21:14

    What has been most challenging? What has maybe not worked, and you went in one direction and then you're going in another direction? Like, any lessons learned that may be interesting for people to hear about?

    Ali Dasdan21:45

    I think across—so, creating a data platform with respect to its components, right? All the integrations on the input side, output side, storage, processing, governance layer, with respect to cataloging, security, permissions, definitely pushing the data for business intelligence, Tableau, and systems like that. So most of these are well known, but it just gets learned and relearned every time, right? So I'm lucky that I have created huge platforms multiple times in the past.

    Matt Turck21:47

    So experience helps, as it turns out.

    Ali Dasdan21:48

    It helps.

    Matt Turck21:48

    Yeah, it helps.

    Ali Dasdan22:05

    But choosing those components, making sure that data flows, let's say, in real-time streaming mode, you can be assured of the quality of that data if you're going to create, let's say, analytics out of those BI systems or custom analytics interfaces, right? All of these should be easy. I think the lesson is this, again, that gets learned every time, is you have to pay attention to all these pipes and the boxes and how these things are connected to each other so that you can do it in a way that you can trust that data, you can turn it around.

    Ali Dasdan22:40

    And time to market is very important over here. You don't have to start everything from scratch. Let's say you're going to do analytics. There are lots of analytics databases, different choices. They are proven with respect to these capabilities. I mean, in the past when we were doing data platforms, there was no Kafka, for example. We created a Kafka-like system ourselves, but we are lucky that we created an HBase—sorry, Hive-like system ourselves with 100 petabytes of data, right?

    Ali Dasdan23:10

    So there's no need for that right now. I think many of these are very mature. You just have to know which one to start with, or just start with one. And you can adapt to that. I think the lesson is time is of the essence. Productivity is key. Just use all these very mature components and connect them in a standard way. There are lots of best practices on that.

    Matt Turck23:12

    Do you use open source at all?

    AI application at ZoomInfo

    23:19
    Ali Dasdan23:19

    Yeah, definitely. For example, our internal productivity dashboards are driven through DevLake. It's an open-source Apache project.

    Matt Turck23:49

    A couple of things. So you had a big announcement a few months ago in May about the Copilot, and we can talk about this, but even before that, it seems that you've had AI-powered applications for a while, including through the acquisition of Chorus, I guess last year or maybe a couple of years ago, sales call intelligence platform. Do you want to talk about, before Copilot, what was the spectrum of things you did with AI?

    Ali Dasdan24:20

    Almost every key AI technology or model-building technology, we had to have it. Because when you're dealing with data of this nature, some of this data is, for example, unstructured, right? How do you extract that? Or if you're dealing with calls, you have to do natural language processing, which simply, you have to figure out what's being said, who's saying it at what point, analytics, sentiment extraction, right? All of these you cannot do without AI. So the company was deep into AI before GenAI.

    Ali Dasdan24:56

    We had a team almost from the beginning, but all these models were deployed across the board, and many of them are traditional models, right? Natural language processing models, tree-based models, classification models, and SVM and things like that. Even with Chorus, right, Chorus still runs our own homegrown NLP technology. We are doing it because our quality, we believe every time we compare, is actually pretty good. We are beating well-known vendors. Plus, it's cheap for us, right?

    Ali Dasdan25:27

    We are running it with just a couple of people. It's really well done. Now, when GenAI came, especially in our business where these AI models were working pretty well, it created a revolution almost. And the first thing we released actually last year was applying GenAI to Chorus. So you go to a call, let's say you have a customer call for one or two hours, you are talking about things, and you want to share this with your team. Well, what's the best way to share?

    Ali Dasdan25:55

    They can look at the transcript, they can watch the video, but then they're going to spend two hours. And one of the best applications and first applications of LLMs were summarization, right? So we actually provided call summarization and action item extraction. And the amazing thing was, literally, as soon as it got released, as you know, increasing NPS is difficult. Sometimes every time you just increase just a couple of points, and the absolute increase in NPS was just amazing, almost like 20 points.

    Ali Dasdan26:10

    Because people loved what they were getting with that. So that was the first experience.

    Matt Turck26:25

    And so, to play it back, quite literally, for something that's straight into the kill zone, for lack of a better term, of generative AI, you just did, like, a rip and replace for Chorus.

    Ali Dasdan26:41

    Yeah, for Chorus, it was an addition, not a rip and replace. We are still using natural language processing with respect to converting recorded conversation to transcripts. It was extracting summaries and action items from transcripts. We are using GenAI, so it was an addition.

    Matt Turck26:47

    And is that because the earlier model still works better than generative AI for that specific task?

    Ali Dasdan27:11

    With respect to extracting from recorded calls, we have not actually tried it in that sense. There were some experiments on that. I think it's going to be very expensive to run this unless we have our own model. Now we are running on GPUs that we have created. Probably this could be one of the areas that we can do. But since our NLP tech for creating transcripts is working really well compared to what's out there, it's not the first priority to replace that at this point.

    Ali Dasdan27:29

    But anyhow, GenAI, that was the first use case and wow moment. Then with Copilot, as you know, right?

    Matt Turck27:42

    And you replaced, for that part of the Chorus product, the generative AI—I can question that—that was like a vendor, like Anthropic, or that's open source? Or what LLM do you use?

    Ali Dasdan27:44

    Chorus. For that one, we used Anthropic.

    Matt Turck27:45

    Anthropic.

    Ali Dasdan27:50

    So yes, we are using Anthropic, we are using OpenAI, we're using Google. But for that one, we released with Anthropic.

    ZoomInfo's Copilot

    27:58
    Matt Turck28:01

    Okay. All right, so that's Chorus. Great story. What are some other examples of use of generative AI in particular?

    Ali Dasdan28:14

    Yeah. Then Copilot came. Copilot is GenAI, even beyond GenAI. The whole point is to completely turn around the interaction of sellers with our products.

    Matt Turck28:26

    Yes. And then, for clarity, it's ZoomInfo Copilot, a product which was just released, as we mentioned earlier. Exactly. In a world where Copilot means different things across a bunch of different companies.

    Ali Dasdan28:30

    But the name fits really well because it's really a copilot and helper for sales.

    Matt Turck28:35

    So what does that do in terms of product experience?

    Ali Dasdan29:01

    So the problem that sellers will have, right, and they want to retain their customers, they want to find new customers, they have to be on top of so many signals. And if you look at what best sellers are doing, they reach out to different data sources, they extract insights from different places. They do, let's say, lots of searches on ZoomInfo and combine the data that we are providing with what they were doing. But then they had to do all of these. And if you are not as good of a seller as these best sellers, then you were at a disadvantage, right?

    Ali Dasdan29:32

    So Copilot completely turned this around, right? So, at the fundamental technology level, it's a recommendation engine, right? But what it did is, it will tell you exactly what you should be focusing on: which accounts, what are the signals that you should be paying attention to, what action you should be taking on top of it. We are actually helping you to take those actions. So if you're going to send an email, we will actually help you with that email. If you go to what we call Account AI, all the information across all the data sources, all Chorus calls, right, all that extracted data boiled down to a version that you can actually pay attention to.

    Ali Dasdan30:04

    And the chat box is also connected to it. So you can literally ask questions about that account. Let's say you want to get up to speed, you're going to have a conversation with a potential customer, and you have half an hour. There is no way you can look through all that data to extract what is the essence of that, right? And we help you to do that, right? So it just turned this completely around, and people just love it because of how it enables you as a seller to be super effective.

    Matt Turck30:29

    It's amazing, by the way, to the ongoing debate that we all have about what happens between AI and our jobs. A big thing that you hear a lot is that it will make okay-to-mediocre people much, much better, in that it just levels up the state of play across professions. But in the case of sales, what you just described seems like a perfect example of this, where you can not be that good at your job, but AI will just basically put you at the level of people who are pretty excellent at the job, at least in terms of preparation and identification of targets.

    Ali Dasdan30:57

    It will definitely magnify the skills that you have. I mean, you still have to take extra action with respect to what questions, for example, you're going to ask.

    Matt Turck30:59

    Yeah. And build human relationships.

    Ali Dasdan31:12

    Closes most of the gap. Yeah, definitely. Actually, some sellers were jokingly complaining about it, saying that, hey, I am a good seller. I'm doing all the work, and now you're going to help these lazy sellers to compete with me.

    Matt Turck31:46

    Yeah, it's absolutely fascinating from that perspective. So if you think about what makes the superstar seller of tomorrow, it will be, I guess, a lot of human relationship, ability to network, navigate a conversation based on questions, and all the things. But, yeah, different. It will amplify a certain part of the skill set much more than prep and diligence and all the things. Okay, great. And so that was just released. Still in beta?

    Ali Dasdan31:53

    No, it was released in production. So in May. Before May, actually, it was in beta.

    Matt Turck31:54

    Okay.

    Ali Dasdan32:23

    Twenty-seven or so thousand customers. I mean, people were saying that they save about, like, 10 hours, literally one day a week they were saving by using Copilot. And with respect to the use of GenAI, it was not only, like, Account AI or summarization, email generation, but actually the models underneath for certain signals were done using GenAI. It actually simplified our life because the typical way of extracting the features and creating a model and the training and all that. So just in some of the cases where we had more than a dozen models, that completely were regenerated using GenAI for Copilot.

    Matt Turck32:42

    Let's actually go a little deeper on that part, on what powers Copilot in terms of those various models. What do they do?

    Ali Dasdan33:11

    So one is the data, definitely, right? Without data, and that's actually the power. So you see lots of mentions of AI: I just throw AI and I am getting good results. But AI without the data that we have, not only its comprehensiveness, accuracy, how you're actually making sure that the model, you are providing the right context to the model. Not only do you get better results, but you prevent hallucination for that context. So all that actually is a requisite for the model to run really well.

    Matt Turck33:23

    So you use your data basically as a RAG kind of setup?

    Ali Dasdan33:48

    Yeah, let's say Account AI use case. So how am I going to extract that account information? Right. So this is my account. Let's say I'm the seller. What are the pain points of this customer? How many conversations have happened so far? And in each conversation, what was the summary of the conversation? Was there a need mentioned or a competitor mentioned, or was there a problem that was mentioned that I can pay attention to? I know exactly what kind of discussions happened about their problem needs as well as pain points and what kind of responses we have given, right?

    Ali Dasdan34:22

    All that has to be extracted. And this is going to be extracted with not only the information that we have about this account, but also the first-party data. Let's say you're going to bring, because let's say you were using Chorus and all that, you own that data. Yes, our technology is recording it on your behalf, but you own that data, and we extracted this on your behalf automatically. So it's in one place. So this is just one use case of Copilot.

    Ali Dasdan34:50

    The others are actual signals themselves, right? So it's going to tell you, let's say you work with John Smith in this company and the person has moved to another company, right? This new company, not only is the news about that, news by itself is just one signal, but more valuable is this new company maybe is in the same industry or they are in the market actually trying to buy something that you are selling, right? All that has to be extracted and presented to you.

    Ali Dasdan34:56

    You don't have to even search for that anymore.

    Matt Turck35:09

    And so the models underneath are what? It's some third-party LLM vendors, it's some of your stuff. And it sounds like it's an ensemble of models rather than one model.

    Ali Dasdan35:21

    Is that correct? So for each of these use cases, you have a model we created. We're using Anthropic and OpenAI, right? We have our own. We use Llama, basically anything that is—

    Matt Turck35:23

    Which Llama? One?

    Ali Dasdan35:30

    So the one that we deployed right now is 7B, I think. Okay, the 7B. Yes.

    Matt Turck35:44

    Are you finding big differences in performance? Like, how do you pick one versus the other? Open source versus proprietary? Do you think of it in terms of performance, cost, latency? How do you select?

    Ali Dasdan36:09

    There is always a trade-off, definitely. Right, accuracy and cost trade-off. So let's say it is a small-scale run. Let's say I'm going to extract, in that context, a simple use case. Go with the email generation, right? We also provide through the UI, like, you can actually customize it even further for the person. It's not just some random email, actually an email based on all that information that we have extracted. It's to the context, and then you can customize it even further.

    Ali Dasdan36:39

    So for that use case, it's going to be, let's say, you deal with 10 accounts or 100 accounts, it's 100 emails. Let's say using Anthropic or OpenAI for that one, it's cost-effective, right? Plus, accuracy is, since they are on par with the top models, and then you can just use them for that purpose. But let's say another use case is I'm going to extract the scoop or news item from these websites, and I have 100 million websites. So using 100 million LLM calls, Anthropic, they can do it, but it's going to be expensive.

    Ali Dasdan36:53

    So for that one, we can actually use our own, right? And plus, we are actually renting the GPUs in Oracle Cloud. Okay, so your Llama runs on Oracle GPUs.

    Matt Turck36:55

    Okay, interesting.

    Ali Dasdan37:17

    Exactly. Then we can run it at any scale because we already paid money for it, right? And for those use cases, even if it's comparable for most of the experiments that we have done, but let's say even if it's just a little bit below in terms of accuracy, at that scale, the advantage that you are getting is still a big magnifier to whatever else you can do right now. If you do not have that technology like that, what are you going to do?

    Ali Dasdan37:36

    You have to crawl the data, extract all those insights, create all those models. Maybe you have it. If you don't have it, you have to create them from scratch, do feature extraction. All that time saving is given to you by using that.

    ZoomInfo AI toolstack

    37:43
    Matt Turck38:00

    And bearing in mind that all of this is new and very much the frontier, do you have a whole set of tools around those? I'm thinking evaluation, monitoring, orchestration. Are you, I don't know, LangChain users, or what do you use and what have you built yourself?

    Ali Dasdan38:25

    I think you have to create a layer on top of this at a minimum, right? In the end, whether Anthropic or OpenAI, they're companies, right? They sometimes have stability issues, right? If they're down, what are they going to do about that? Since our responses, Copilot, is all real-time, we have to respond to that. There is security: whatever we are sending and whatever we are getting as a result of that has to be scanned, right? So that layer actually includes that.

    Ali Dasdan38:51

    There are companies today offering even further on that one with respect to accuracy and cost trade-offs. So right now we have not implemented that one, but that's one thing that we're actually considering doing within that layer. There's also, as you know, either GCP or AWS providing a layer by themselves for you to make it easy to choose different models for different use cases, right? So we are thinking, like, how much of that layer we should be creating ourselves and how much we can actually rely on what Google or AWS or Oracle will be providing as a layer for us.

    Ali Dasdan39:25

    And we can just use that. And given, I think, how valuable these products are, how well they work, and the enterprise use cases will force enterprise support structure around that security and privacy and model selection, all that. So as a result, probably we've got to be careful with respect to how much investment we are going to do versus what's going to be given to us a little readily through these cloud vendors.

    Working with small vs. big companies in the AI business

    39:30
    Matt Turck39:56

    And of course, the question for all startups and people like us, VCs who fund them: from your perspective, because you're a provider of technology, but in some ways you're a customer of technologies as well, for generative AI, how do you think about working with startups versus working with large companies? And in large companies at this stage, I would probably include OpenAI. What I'm talking about is, like, the small Series A, Series B startup providing whatever LLM evaluation kind of thing. Is that something that you would consider, or basically, would a company like you wait until GCP has the feature?

    Ali Dasdan40:21

    Yeah, I mean, we are very technical ourselves, so we can evaluate them, I think, pretty well. We work with so many vendors, small or large. The point is how it is helping me. Am I definitely saving time? Productivity is top of mind for us, as well as, if I were to do this ourselves, right, we can do anything. In the end, it's software, but how many people are you going to allocate, how much time is it going to take, and all that?

    Ali Dasdan40:41

    If there is a solution out there that we can use, we will do that. But one key is it has to be very easy to integrate and it should work, right? And we are going to make sure that through evaluation, that's the case.

    Matt Turck41:04

    A few, maybe four or five minutes on the other side of the house, or, like, an AI part of the house that we alluded to, which is using data and AI for internal productivity. So you mentioned there was a separate platform. So maybe walk us through how that works and what that does.

    Ali Dasdan41:31

    So one is, I think we have an internal motto that we always use: how fast do you move is probably the only advantage of any company, including ZoomInfo. So as a result, I will actually give a talk today about productivity on that one and share some of these. And I think a couple of months back, at FirstMark CTO Summit, to plug this, and the CTO Guild, again from FirstMark. Thanks for that. Thank you for the invitation. So I had actually given a talk on using Copilot, GitHub Copilot, for developer experience.

    Ali Dasdan42:05

    So again, since productivity is top of mind, we try to use these tools to speed things up, right? So the first easy use case was GitHub Copilot or OpenAI or Anthropic, literally to generate code, generate testing, help with documentation. And people love it. We do regular quarterly developer satisfaction surveys. People love what they're getting out of that. There was a big push, actually, for us to deploy these as soon as possible. We made sure that it is secure, went through all the basic enterprise checklist items before actually deploying this.

    Ali Dasdan42:43

    Now we are trying to go beyond that. It's not only helping while you are coding and writing tests. We are trying to, for example, speed up code reviews using this technology. And Google had an internal, I think, solution that they had a paper about. I'm not aware of any startups doing that. Probably there are, but we are trying to develop ourselves internally to speed this up because code reviews are a best practice for improving product quality, software quality. But at the same time, it takes time.

    Ali Dasdan43:11

    It takes time from both sides, from the coder who wrote the code as well as who's reviewing the code. So if you can speed this up, even the nice thing about these things is you don't have to completely eliminate humans. Even if you save, let's say, 10 minutes, when you multiply it by the number of code reviews you are doing and by the number of developers that are doing it, the savings will be immense, right? Trying to do the same thing for incident detection or creating, for example, telling you exactly maybe what could be wrong.

    Ali Dasdan43:38

    With that incident. And response time for an incident is not only to discover, but to resolve, mitigate the issue, is super important. Again, GenAI can help a lot over there. So we are getting initial signals around that, some good use, but we have not deployed a solution yet. But it's being developed.

    Using data and AI for internal productivity

    43:39
    Matt Turck43:52

    And for GitHub Copilot, what has been the sort of human element of this? Is that something that you've had to tell developers, no, you need to use this, or do they love it and they do it without you asking them?

    Ali Dasdan44:16

    Actually, they love it. Yeah. I mean, not everybody started using it right away, but there was a push to get this deployed as soon as possible. So that was a nice thing to hear from the developers, that they also want to move as fast as possible. So we actually deployed this, I think, last year. We first did a trial with a number of developers just to make sure it works out. Good signals were there, but just to prove that it actually works out for the use cases, again, it's secure and we are going to protect our intellectual property.

    Ali Dasdan44:55

    All of that, we checked, confirmed, then we deployed to all of our engineering. So everybody's using this daily in their interactions. And our acceptance rate for code suggestions was around 25%, actually over 25%, which is interestingly in line with the recent research that GitHub published. I think they were mentioning 26%, which was good that we were actually aligning with the number that they mentioned.

    Matt Turck45:18

    As we just mentioned, we were recording this right before your keynote at the FirstMark CTO Summit, and I'm super grateful that you took the time to do this in addition to the keynote. So I want to make sure I don't make you late. Thank you so much for this deep dive in all things data and AI at ZoomInfo. It's fascinating. Plenty of interesting lessons and tidbits.

    Ali Dasdan45:19

    Really appreciate it.

    Matt Turck45:19

    Thank you.

    Ali Dasdan45:22

    Yeah, this was fun. Thank you for inviting me. Thanks for having me.

    Matt Turck45:22

    Thanks, Ali.

    Ali Dasdan45:22

    Yeah.

    Matt Turck45:43

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.