Welcome, Olivier. Or should I say bienvenue? I hope that people are ready for the sound of two French people speaking with each other in English.
Yes, the problem usually when that happens, the accent devolves into something that's very French. We'll try not to get there.
People have been warned. You've been in New York for 25 years, right? Or something. You arrived in the late '90s?
Yeah, I moved here in 1999.
Okay. Yeah, same general thing. Very comparable journeys, obviously, other than the fact that you built a $39 billion market-cap company and I moved here with a podcast. But other than that, pretty close. Why did you stay in New York?
Oh, first because it was fun. So I moved in 1999, thought I would stay six months. I had an internship at IBM Research, which was a very interesting place at the time. And then it was the tail end of the dot-com boom. So there was a lot going on at that time in New York. There was not much going on in France from a tech perspective. So I thought it was a great time to stay and join startups.
It was interesting. Of course, then there was the dot-com crash, 9/11, all of that stuff. So that was less fun. But by that time, I was very attached to the city. I liked the dynamism, I liked the cultural aspects of the city, and I decided to stay. And I stayed until—it was until 2010 that I started Datadog in New York.
And what do you make of the current or recent explosion of the French tech ecosystem? Because indeed, I do remember the exact story. I came here, and then there was basically not much to go back for in France in tech at the time.
Yeah, I think it's amazing. I think when I left, there was not much going on. When I started Datadog also in 2010, there was also not much going on. It was difficult to start companies. Few people wanted to work in startups in France. It was difficult to get funding, all of that stuff. And I think now it's changed a lot. There are many companies that are getting started that are very dynamic, doing very well. I think we still have to see them scale to large companies in France, in Europe more broadly.
It hasn't happened yet. So the ecosystem is still young and, I would say, still fragile from that perspective, especially, of course, compared to the Bay Area, but also compared to the New York ecosystem.
Yeah. And exits, right? Scale and exit. Because France hasn't had that cycle of people making money and redistributing it into the ecosystem.
It has been, but typically earlier, like smaller exits. And I think they need to have more of the later ones.
Cool. So maybe for context, you and I did something sort of similar a few years ago in the context of Data Driven NYC, which is the monthly data AI meetup I've been running for 11 years now. In some ways, this podcast, The MAD Podcast, is a spinoff from it. But that was a really good conversation that stood the test of time. As I was prepping for this, I listened to it again, and I would encourage people to go back to it to hear sort of the basics, like the Datadog story, the founding story, the initial vision.
Where the name comes from, all the things. And as real YouTubers say, we will put that in the show notes, so people should check that out. But for today, we're going to try and do something a little different, more focused on product, more focused on all things data and AI at Datadog. Maybe a great place to start would be the usual 101 on Datadog, your elevator pitch that you've probably given about 10 million times at this point.
We're Datadog. We do observability and security. We do it for cloud environments. So we sell to engineers, basically, that run applications and infrastructure in cloud environments. And we do it for companies that are big and small. So our smallest customers don't pay us anything, and they're individuals or students. And our largest customers are the largest companies in the world, paying us tens of millions of dollars a year, and we have pretty much everything in between, but a very extremely technical product sold to a technical audience.
And the core has been observability historically. What is happening in that space? It sort of feels like there was infrastructure monitoring, there was application monitoring. It sort of feels like things are starting to collide a little bit. Is that fair?
Yeah, it used to be different categories, used to be called monitoring. Though a lot of the use cases are fairly similar, and it used to be balkanized. So there used to be monitoring, application performance monitoring, network monitoring, log management, user experience monitoring. All of these were different categories with different user bases and different types of instrumentation, different vendors. I think what we've done that, to be honest, was kind of—I mean, it was actually the starting point of the company, but was a bit of a hard sell initially—was bring together those categories into one platform.
The idea behind Datadog was, well, those teams don't see eye to eye. They don't understand the world the same way. For the story there, in the previous company, my co-founder and I used to run two different teams. I used to run the dev team and he used to run the ops team. And even though we're very good friends, we hired everyone on our teams in these companies because it was a fast-growing startup, we did end up with ops that hated devs and devs that hated ops and fighting all the time and finger-pointing.
So the starting point was, hey, let's get them on the same platform, sharing the same language, same view of the world, and working together. So the first use case we brought up was really infrastructure monitoring for cloud. And then we added all of the other parts of that stack. So application, log management, user experience, network, all the other things that would make sense in there.
And very much for anybody that has been following the Datadog journey a little bit, what's been amazing to watch is that evolution over time from that start to a platform that keeps getting broader and more horizontal as you keep adding product after product every year. In fact, it seems you have this chart in your public documents somewhere where actually the number of things that you add every year—features, new products—seems to be increasing, which is amazing to watch. But going back to the beginning of that, so the vision was always to be a platform because there's that kind of cliché in venture and startup circles that you need to be a tool before you are a platform.
So how did you navigate that?
Well, initially not very well. So the idea was, yes, we were going to be this platform. That was our raison d'être. That was why we started the company. And we had this vision of multiple datasets, multiple teams, multiple sets of use cases, all meeting into this platform. So we started a company, we were super excited, we applied to all sorts of incubators, we got to a Y Combinator interview, and then we didn't get into Y Combinator. I have an email from Paul Graham that says, "A platform is only as good as its first product, and you don't have a first product," blah, blah, blah.
So it was not like—you're right, it actually was a tough sell. And to be fair, when we started releasing our product, like the first beta of our product, we did call it a data platform as opposed to calling it a very specific use case. And everybody loved the idea. The users loved the idea, but people were not coming back to the platform. And also, nobody was paying for it. So that was a bit difficult. At some point, we decided to name it monitoring, which was the name of the category before us, really.
Infrastructure monitoring. And without making too many changes to the product, we added a few small things to make sure it would be called that way. Immediately, it caught traction. Immediately, people understood, okay, this is why I need to bring it back, and this is how I convince my boss I need to pay for it. And we were off to the races from there on.
Because it was a known budget item.
Yes. And that comes back—we can talk about it later—but in terms of understanding where you're going as a platform, but also how you bring that back to where your customer base is today, I think it's a key part. You actually have to take your customers where they are and talk to them in terms that they understand, in terms of their existing categories, their existing spend, existing solutions.
So maybe walk us through some examples of that evolution over the years. You started with metrics, traces, logs? What was the next sort of few products over the years?
So we started with just metrics and infrastructure. We didn't have tracing, we didn't have logs. And for the first 7 years of the company, we didn't do anything else. We had the great chance of having amazing traction with that core product. So we were just too busy keeping up with the customer base, both in terms of scaling the systems, but also just adding all the functionality they needed for that and supporting all these new customers.
So first, metrics is just any number for performance of your infrastructure.
Yeah. So metrics, we had a number of integrations basically that would get the data from whatever systems our customers were using, whether that's their cloud provider, the database they're using, the operating system, the CDN. But also, we would let them—we still do—submit custom metrics for their own applications. So if they want to track, for example, the number of times a certain function is called, the number of items in shopping carts, the dollar amounts of ad impressions, they can do all that.
And they did a lot of it. That's what made the product so successful, that it could cover the gamut from how much CPU you have all the way to how much revenue you're making with the product. That was extremely appealing, especially to all of the digital native companies at the time that were moving into the cloud. So we started with that, I would say, from 2010 to 2016, we mostly did that. But our platform was still fairly open. You could add any sorts of data, manipulate the data in all sorts of interesting ways, build analysis and visualizations and things like that.
And what we saw was that a number of our customers started building other products for themselves on top of our platform. So we didn't do application performance monitoring. There were other products that were doing that, like a whole category of products. But we saw that a number of our customers were building a poor man's APM—that's the shorthand for application performance monitoring—into our platform. And that told us a few things. That told us, well, there's a need, but also our customers believe that we're part of the solution.
So they want to see it there, and they're willing to go through the trouble of hacking together something on top of us. So that gave us great confidence that our initial vision was good and that we should keep developing that, and the next product should be application performance monitoring. And we started building that.
And when was that? What year?
2016, I think, is when we engaged with that. And this one actually took us a while to get right, and we can also come back to that. Part of it was going from product number 1 to product number 2, which is hard. Part of it was also that it's actually a category that has very high table stakes, and it is a bit of a slow grind to get everything going with customers. So it took us several years, actually, to get that product to be truly successful.
Yeah. And then there was New Relic, AppDynamics.
There were a number of other companies out there. And at about the same time, we also found that log management is also a category that belongs in what we were doing. It's another aspect of observability that is very important for users. And same thing, we saw customers trying to hook their log management to us in different ways and build and hack solutions for doing that. We also found that the problems that customers have didn't stop at the barrier between application and infrastructure and logs.
They actually cross everything. So it made complete sense to do everything in one go. For log management, we started working on that too. This one we accelerated with M&A. So we actually bought a company in France, a small company that had an early product for that, and we accelerated that development. I think through the combination of this being accelerated by M&A, this being also an easier product to get to critical mass compared to APM, we actually ended up with both gaining traction at exactly the same time, even though we started with one maybe 2 years after the other.
What came after? There was, like, synthetics?
Yeah, we did the rest of what's part of an observability stack. So we did synthetic testing, which is automated testing for APIs, web applications, things like that. We did real user monitoring, which is measuring what end users are actually doing in the application, first for performance reasons, but then increasingly for business analytics reasons. So what are my users clicking on? Where are they going from there? That kind of stuff. We started expanding into security. So entering, doing security on top of logs, there's a category that's called SIEM for that.
We started doing security on cloud environments, both on the infrastructure side and the application side. And we have a thesis that security, 5 years from now, will be a no-brainer that you have to attach your security to your observability, because that is what gets deployed everywhere in your application, in your infrastructure, and what gets used by all your engineers. And to solve your security issues, you cannot just rely on a team of, say, 10 security engineers.
You need to rely on your 50 ops engineers and your 200 software engineers to make that happen.
That's because a threat would manifest into logs, into activity, or—
Yes, but also most of the security issues are introduced by your development team and your operations teams, and most of the fixes have to be made by your development team and your operations teams. Mostly, dealing with security is not monitoring North Korean hackers. I mean, there's a little bit of that, but mostly it's about patching all of the various vulnerabilities you have. Because maybe you made a mistake when you were coding, or maybe you use a library and that library now has a vulnerability and you need to update your code and make sure everything works, or you misconfigured something, or there's a vulnerability that exploits a certain type of configuration.
So you need to continuously go and close all of those vulnerabilities. And the people who do that are typically not the security team, they're the development and the operations teams.
So it's very interesting, right? The first two or three, so going into APM and other things, sounds like your customers basically dragged you into it. They started building the product that you didn't have and showed you the way. Is that true as well as security, or is that more of a strategic kind of decision? Because I totally get how a lot of this manifests in what you monitor, but equally, it's a different industry, it's a different buyer.
I mean, it's a bigger leap. I think it's true in part. Same thing: we saw customers build some versions of that themselves. And for some parts of what we do, it's a natural evolution. Like in log management, for example, the companies that used to sell logs before us then evolved into security companies as part of that. So it's understood that the two sort of live together.
But for other parts, application security and cloud security, for example, these were very separate. So I think it's a bigger shift. It's also a shift in terms of how to think about the security market in general. So I mentioned earlier that observability used to be very fragmented and now is one big category, really. And that's what we're leading today. Security is still an extremely fragmented category. I mean, you're in VC, so you probably can spend the first three hours every morning keeping up with all the new funded companies in cybersecurity.
Yeah. And you can have an entire investing career just doing security, which some people do.
Yes, there's so many of them. And as a result, you end up having so many different subcategories. And we think that it's too complex. It's impossible for the customers to actually understand how to piece that together and to integrate everything into a consistent whole that completely protects them. And so we think that security is going to go through the same consolidation, basically into platforms with some extra products here and there, but mostly relying on large platforms.
I'm very fascinated by that topic of how you keep building products because, as you said, it's so hard to get a second product. Most companies never get the second product. So one aspect is your customers point the way. So that's super interesting. Sort of almost logistically, how long before you jump into an industry do you start planning that? How does that start? Do you do research efforts, do you have an internal consulting team, or is that you as CEO who decides this?
And then how long do you plan before you start building and then launching?
Well, mostly we hear about it from customers. So we have this strong bias in the company that our window into the world is our customers. So, for example, when we think about competition in general, we don't spend a lot of time reading our competitors' press releases or websites. Instead, we talk to our customers and we hear what they have to say about it, whether it registered with them or not, what seems to be valuable to them or not. Basically, the world around us exists through the prism of our customers.
So from that, we get a sense of what's actually a problem for them, what seems to be real, and then we decide what we might try to go and work on. But the way we build it is—we don't do it like Apple. We don't disappear in a basement for three years and then ship a fully formed product that takes the world by storm. Instead, from the earliest days, we work with design partners, we work with customers, and we try to ship something to them as quickly as possible.
So that's really the way we develop.
And you offer it for free initially for those design partners?
Well, there's a whole process. And actually, a key part of building a new product is understanding what's valuable and what's not. And it's difficult because when you start, your design partners are typically your existing customers because you have strong relationships with them. And typically they love you, they love your product. They're happy to spend time working on new things, but they haven't necessarily thought through the whole value of what it is they're asking you to build. So the way we do it is we start by building with them.
So we ask what the problems are, we get a sense of what else they might use for it, what's their next best alternative, or what they were using before that's not good enough. So it helps ground everything. But then as soon as we have enough product, we basically say, okay, it's going to cost you this much to use this product. And what typically happens at this stage is that when the product is ready, when it's great, about half of the design partners just disappear.
The people who are showing up in the meeting every week or twice a week and were very happy to work with us, they start ghosting the product teams because they realize, actually, I'm not able to pay for it. I don't want to pay for this. It's not valuable enough. And so that's the first gate. The first gate is, do you retain enough of the design partners? Is there enough value in what we're doing there, or is it just good stuff?
After that, we typically start opening up the products more broadly. So we typically are going to have some form of a private or public beta for the products. And then we're going to go through another, initially for free, and then we're going to go through another phase of, okay, so now we're going to announce pricing and we're going to start charging the most active of those customers. And same thing, we see: do we have enough retention there? Do we clear the bar of that product being valuable enough that we can keep going?
And then we do that one more time, which is when we completely open it and when we start charging automatically for usage. We see, okay, so when people start using it, do they keep using it after we start charging for it or not? And then that tells you, okay, the value is there or the value is not there. And I think it's very important. We feel that the worst that can happen in software is that you build stuff, but nobody cares about it.
And it's very easy to build stuff.
Yeah. Have you killed products that did not meet those bars?
We did not kill products, but the way we do it is we start small, and we've waited a long time to scale up teams or products as we were still searching for basically the core value or what we would be able to sell in the end. So some of these are still searching for the exit. Some of these are still—we have a small effort going slowly, exploring. And many of those have scaled as we went through all those gates and we validated the value and we understood better the packaging, and we could then have a very clear roadmap in terms of what we need to build for those customers to keep scaling.
Who does this when you launch a new product? Do you have an internal, I don't know, Sherpa team? I know you're active on the M&A front, so I assume in some cases it's whoever you buy. But do you take people from other products to reassign them? How does that work?
Yes. And that's in part why it's difficult to build multiple products, because what happens when you start thinking about product number two typically is you have product number one that is very successful. But if it is very successful, chances are everybody is super busy just keeping up.
And everybody that is super good and you would want to trust with starting a new product is load-bearing on your core product. So that's difficult. You need to pull people away, and that's painful. That's hard. The second part is that scaling a successful product and starting a new product are very different motions. And I would say people feel very differently about it. In one situation, you mostly work in situations where customers love you and you know exactly incrementally what you need to do next.
And you can work with them. In the other situation, you don't have traction yet, you're trying to understand why, and you have to decipher the cryptic feedback you're getting from customers. Because again, these customers in general are good people. They don't want to hurt your feelings. So when you talk to them, they'll say, "Oh yeah, no, this product is great. It doesn't actually work, but it's great."
And the people who are used to scaling the products that work ignore the negative part, whereas that's the only part that's worth listening to.
And I assume that's true on the sales side as well, right? You can become super great at selling something that people want, and then you have to sell the new thing. So how does that work?
That part, I would say, for us is a bit easier because we get adopted bottom up. So we focus on getting usage first: making the product discoverable, getting really short time to value inside the platform, in the product itself. So then the sales side of things is a bit easier. I think as the products get more mature and as the large enterprises start using them more, there's more of a sales job, basically, to work on consolidation. So basically situations where customers were using 12 things before and they're just going to use us after.
So there's more of a sales job to be done there that's more traditional. But otherwise, we mostly do it bottom up. One last thing to mention, by the way, on putting people into new projects is that you need to see it through the angle of people caring for their career too. You have to make sure it's safe to work on a product that's not successful yet, because otherwise people just want to stay on the star of the show and not go into that thing, and they don't know if it's going to take off or not.
Now, how do you measure the success of a new product? Do you have an internal metric, I don't know, around attach rate or whatever you call it?
I mean, we try to get very good signals. So I mentioned that we establish value by starting to charge for products initially. So we try not to do too much bundling. Of course, you do some of it because it's enterprise software. You can't just piece out everything, every single feature. But we try to get clear signals in terms of: is it getting adopted? Is it getting used more and more? So both on the revenue side, but also on the straight activity in the product.
So how many users do they have? What's the footprint in terms of the datasets that are being sent, the impact on infrastructure, all those things? So we track all of those. We also optimize for very short feedback loops. So in particular, and that dates back from the early days of the company, we always start with very low commitments and very short timeframes. So for new products and new companies, I would argue month to month is great because your customers can churn at any time, which means you'll get the hard reality to hit you in the face and you can't ignore it. Which is a problem when you sell one-year, three-year deals. Even though you might see concerning adoption or usage metrics, you can fool yourself into thinking that you'll be able to fix it.
So when you have a very short cycle time on that, you really have to fix things very quickly.
And is that just an impression, that you keep releasing more and more products? Is that like on those slides that I was mentioning earlier, or is it that you've just built a muscle and you're just cranking because you know how to do it and you keep doing it at industrial scale?
Well, it's more of a factor of demand. So I think we see that there's a lot more we need to do in observability. And there are other categories. I mentioned security. There's a number of new things about AI. Of course, we'll have to talk about it at some point.
Yes, right after this. Coming up.
There are things that are getting closer to developers too that are very interesting.
Yeah, it's a thought that came to mind as you were describing the fact that security is a developer problem. It sort of feels like that's a world of DevSecOps. You've been very focused on the metrics and what comes out of the machines, but it sort of feels like you need to go into the world of code. Is that—maybe you're already doing that—but is that part of the idea?
I mean, part of the idea is you have to tie it back to what the developers are doing. Yes, I think the act of coding itself, historically, hasn't been an area that's conducive to being solved with software products so much. I think it tends to be more of a commodity side of the world, as opposed to everything that relates to production systems. But we think we definitely need to close the loop between what's happening, what developers are doing when they type a line of code, and what the impact is actually going to be on their application in production, on their end users, on their business, on the security posture of their business.
This is the hard part. Again, I think I'm getting a little bit ahead of myself there in terms of where AI is going to take us. But one way we like to think about it is the wave of, or the exponential increase of, developer productivity. If you go back 40 years, 50 years, people were coding in machine language on punch cards. Like, my dad started coding on—
It was like rulers as well, right?
Yes, rulers. So fairly low productivity. Then you coded on keyboard and screen, but still in machine language. Then you had more advanced languages. Then you had more advanced languages and libraries that other people wrote that you could use, but you still needed to pretty much write all the code yourself and buy books from the library to understand what to do. Then you had the internet. Then you had open source with all of the libraries you could use everywhere. Then you had SaaS and PaaS and basically all of those things you can just plug into, these APIs that are going to do a lot of the work for you.
So we've seen over the past 40, 50 years orders of magnitude of productivity increases. And I think AI on top of that is going to give us maybe another order of magnitude or two in terms of what a human can produce functionally. The flip side of that, though, is every time you add more productivity, you add complexity, meaning that you have less and less of an understanding of what it is you're doing. Because you did something super complicated in five seconds.
You can't possibly know what's going to happen on the back of that.
So a lot of the value shifts from just creating that thing to understanding how it's going to behave, how it's changing over time, what its failure modes are, what impact it has on its users, and basically how it can be abused from a security perspective. And that's where we see ourselves playing a role in the long term.
So as I think about it, there's different themes for Datadog and AI, and you touched upon a couple of those. There's AI: friend or foe for Datadog as a business. There's building AI into Datadog products to make them better, faster, smarter. And then there's probably a lesser but interesting theme around how Datadog employees use AI to maximize performance, productivity, all those things. So maybe starting with the first one. So maybe playing it back, it does feel like the increase in complexity is your friend, right?
So the fact that we may have all those machines that will do things for us that we may or may not understand is a good thing. Do you think there is a world, just to play a little bit like an AGI doomer, where AI actually becomes a problem for Datadog in, I don't know, automating or being able to do all the things in a way where humans do not need to be involved?
Well, I think at some point someone needs to control the AI in some form, right? Or maybe not. Maybe we just, in a way—or the AI does. Yeah, we're all pets in the end. But I don't subscribe to that vision of the future. But I do think that at the end of the day, I do see that as one more evolution in the history of innovation, which is we just do more. And because we can do more more easily, we do even more.
And then we need to manage it and understand it. I think that's the overall arc of things. And that's where we have potentially an even bigger role to play in the future as that happens. And when you think of the impact on our business, there's a few ways to think about it. I mean, the most straightforward—and the ones happening right now—is the emergence of AI just pushes more digitization and move to the cloud. Because, I mean, to capitalize on AI, you need data, so it needs to be digital.
And you probably are not going to build your data center for AI yourself. If you are a handful, maybe 10, 20 companies, yes, you are. Because you are at such large scale and you're in the business of providing your services for others. But otherwise, you're not, because you don't know what you need. How would companies even know today what to buy and what it looks like three years from now when they've realized those investments? That's impossible.
There seems to be a little bit of a trend around, actually, let's not move data to the computer, data to the AI, but bring AI to the data. And therefore, actually, on-prem or VPC, to some extent, is a good place to do AI, and sometimes open source, hence a part of the NVIDIA rise, but also Dell.
To be clear, what we call on-prem at this level is the cloud. It's on-prem, but it behaves—it looks like and behaves exactly like the public cloud. So it's the same thing. And there's a dynamic also right now where the compute itself, the GPUs, are way more expensive and don't need that much bandwidth compared to everything else. So it's actually okay if your data is in Australia and your GPU is in the U.S. That's not a big issue. That's not necessarily how things are going to be in the long run, though.
I think it's more of a side effect of where we are right now in terms of the technical solutions, the evolution, the size of the models, the price of the GPUs, and all those things. So I don't necessarily think this is a long-term trend.
And from your perspective, all of these are just systems and machines that spit out machine data that needs to be monitored, right? So, for example, I was reading somewhere that you integrate with whatever NVIDIA part of the NVIDIA platform that enables you to monitor the health of GPUs.
Yeah, I mean, look, again, straightforward, that's more infrastructure. So we need to understand the GPUs, we need to understand there's new components in there. Like, the models themselves, from the outside, are components with latency and error rates and cost and things like that. We need to monitor these new databases, new vector databases, and a whole new stack basically around AI.
But a vector database does not behave any differently than an OLAP database from your perspective?
No. And also, the AI applications, typically the part that's a model on GPU is only a small part of that app. Like, the rest is—it talks to a database and it talks to a web server and it talks to its source files, all that stuff. So that part is, I would say, more of the same. I mean, it's not exactly the same. There's many new technologies. We have many people working on things like GPU profiling that are very different from what you would do with a CPU.
But I would say you can think of it as very similar, just another iteration of what we've been doing the whole time. There's a second thing which is a bit different, which is understanding the models themselves and how they behave. And that one is quite nice, completely open-ended, mostly because the models are changing very fast, but also the applications around those models are still fairly early in their lifecycle. I think there's a few things that we understand well. So we understand what an image model looks like, and we understand how to build a chatbot.
And those form factors are sort of there. But for the rest, we're still looking for the right form factors. And the models themselves, as I said, keep changing underneath that. So I think it's going to take maybe a few years for that category of observing the models themselves to fully flesh out in terms of what the use cases are, the needs, and where they go.
And you just announced something, right, at DASH, which is your annual conference, an LLM observability product that does some of this.
Yes, we named it so it's easy to pronounce: LLM Observability. Try to say it four times in a row. I had to practice hard for the keynote for that.
And from what I read about LLM Observability, you do indeed operational performance, but very much to the point that you were just making, you also evaluate functional quality. So the quality of the results, which is new territory for Datadog, right?
Yes. I mean, look, functionally observing software was easier when everything was deterministic. Now that the software changes over time and humans might actually disagree on whether it behaves properly or not, I think it becomes a much harder task. But again, that means there's more value in understanding that and bridging the gap between the humans and what the application is actually doing. And we're still fairly early there. We have a product that is out. We have real customers that use it with production applications, which is great.
I mean, a year ago it was not the case. Everybody was talking about ChatGPT. But other applications in the wild were very few and far between. Now we start seeing that happen, but I expect that field to change quite a bit in the near future.
But it is out and you have customers in the sort of evolution of a new product that you described. It's at the design partner stage where people are sort of—
But I would say it's different from other products in that most other products go after categories that are fairly mature, where the use cases on the customer side are fairly clear. I would say this one is still in motion. It's very possible that two years from now, the applications look very different. The applications built on top of LLMs or models in general look very different.
So that's one theme of the impact on the business and the opportunities. And it sounds like one more way Datadog is just like this unbelievable business, which is that any new thing that pops up fundamentally is just another thing to monitor.
Yeah. And look, the strength of the business is that it usually doesn't make sense to look at one aspect in isolation. You are not going to manage your AI separately from your databases, separately from your network, separately from your security. Everything makes more sense when it is managed together and when we can assemble the full picture for you. So that's what we focus on. The last aspect of the impact of AI on our business is, obviously, there's a lot we can do with AI ourselves on the back of it, which is we have all this data, all these use cases.
What can we do to automate that for you? And I would say this is an area that was, like, for the longest time in the history of the company, we've been careful about using the words AI. We thought they were bullshit words, mostly. People would say AI, but in practice this was statistics or, even less charitably, just additions, subtractions, multiplications. I mean, obviously now it's different. We definitely can do a lot more with it. But the promise here is, since we have the cleanest data, the data you use today to wake people up at night, so when it's not clean, it gets fixed pretty quickly.
And we are in the middle of the use cases, like the transactional use cases of, hey, I need to fix something and verify that it works. We are at the ideal insertion point to automate a lot of that.
I vividly remember that part of the conversation from a few years ago when we did that chat, where you were precisely using the example of waking somebody up and false positive versus false negative. And yeah, you did seem careful about machine learning at the time. Having said that, so fast forward to today, you were saying this shortly after launching Watchdog, I believe, which is the AI product. I'll let you describe it better than I can. But there was a little bit that you were doing.
And so generative AI has increased not just the scope for what you can do, but also the level of confidence you have in terms of not waking people up at night.
Yes, but I think most importantly, we don't have to lead too hard with it. The problem with AI in general is if you start solely with the promise of automating with AI, you're sort of on the hook to doing it all the time. And the general case for what we do is not something that you can solve with AI. I think right now we'd be lucky to solve 2%, 3%, 5% of the cases with AI. Maybe in a year it's going to be 20%, maybe, but it's going to be gradual.
I think if you try and insert yourself saying, I'm going to automate it for you, the incentive will be to try and do too much, break the confidence with the user, and then that not working out in the end, which has been the story of AI in our business. I mean, in systems management for the past 20 years, basically, that cycle has repeated again and again and again. I think our strength comes from the fact that we can pick and choose. We can say, hey, we're here for observability, we're here for security, we're here to make sure you can solve your issues.
And little by little, we'll do more of it for you, up to the point where maybe you have only 1% of the issues to solve yourself.
Maybe it might make sense to just double-click on what AI means for the kind of data that you're dealing with, right? Because we're not talking about creating avatars or voices or marketing copy. We're talking about detecting anomalies across mostly structured numerical information everywhere, which, to the point of having that Watchdog product back in 2018 or '19 or whenever that was, that's very much, I guess, what is now known, incorrectly so, as machine learning versus generative AI. So what is the mix of different kinds of models you use?
Yeah, so the first generation of products we've built there were built on statistical methods and traditional machine learning, which means the best way to describe it is: not transformers. And that's worked really well. This fit-for-purpose approach works really well. I think with the new generation of models that we've seen come out, what's relevant to us is, first of all, the language models themselves are very relevant because now we also can access not just the numerical data we have, but we can access the documentation our customers have written about their systems, what they're saying on Slack and email about what's going on in their incidents, and piece together the semantic meanings of the various things they even put into the events, the logs, and things like that.
So that's very interesting and opens up more things there. There's more we can do maybe on the reasoning side. I would say even with the latest releases from OpenAI, it's still fairly early in terms of the quality of the models and what can be done there. And then what we're seeing also is how we can apply this amazingly scalable Transformer technology to what we used to do with traditional machine learning, which is time series events, and even before that, statistical modeling.
And speaking of which, another thing that you announced, which I believe is not in production yet, is you've just announced your own foundation model, which is called Toto. What is the idea behind doing a foundation model for time series data?
Well, so the idea is, now that this technology has been proven to work really well, what can we do with our own data? Can we actually improve forecasting for our own time series, for observability time series, and maybe also for other kinds of time series? So Toto was our first attempt. And what's really interesting is that even though it's our first attempt, the first time we incorporate transformers in what we do, it's also not a gigantic model. It's not billions of parameters.
It's much smaller than that. This model from day one was state of the art. It beats all the other models, of course, on observability data. We have special benchmarks for that, but also for other things like weather data, which was very surprising to us. The reason for that is that we've just got amazing data for it. We've got, of course, tons of time series data, but we also have really strong metadata to understand the quality of that data. So going back to what I was saying earlier, we know which time series are waking people up at night.
We know which time series are being looked at and how often, on which dashboards, which gives us really strong signals in terms of quality. So the plan now is to double down on that. So there's more we can do to make that model bigger: more data, larger model size. And there's more we can do also to incorporate more types of data into it. So multimodal data, though it's not in the sense of images and sound, it's in the sense of mixing text data and time series data in the same models, basically.
All right, so we talked about LLM observability, we talked about Toto, we sort of touched on Watchdog, which again is a product from a few years ago that you've kept on improving. What does Watchdog do?
So Watchdog does the anomaly detection. So basically the idea is, when you use us, you're going to generate millions or billions of time series and logs and things like that. And it's just impossible for humans to watch all of them and understand if something's going wrong. So the idea with Watchdog is it just watches them for you and it tells you, okay, so that thing is deteriorating, you should pay attention, or that thing is a sign that other issues are going to happen later down the road.
Going back to the conversation we had a few years ago, the most important thing when you do things like that is to not generate false positives. So these systems tend to generate more false negatives than false positives. Basically, if the system is not sure, it's probably not going to tell you anything, because if it tells you things are going wrong twice and you don't believe it, you're never going to look at it ever again. So I think as the technology improves, as the quality of the models and the forecasting and the detection improves, we can go from having something very interesting to say in 20% of the cases to maybe all the time, which would be amazing.
Bits AI was the first time we incorporated the LLMs into our product. And we did the thing that was pretty much everyone's first attempt at incorporating GenAI, which was, hey, we have all this data and all these things in our product. In addition to our UIs, let's add a chatbot to it so you can ask for anything, mix data from various parts of the product, and interact with it in text. Which we find is—and I think we've been through the same path as most other companies have done—that it works great for some use cases, but you need to handhold the users a lot more.
Like, users don't necessarily know exactly what to ask for and how to ask for it, or where to go next once they've asked a question.
And we've all been there. Like, we all tried the various bots, whether it's the Google one, the Microsoft one, and you sort of try them, you force yourself to try them, and then you run out of ideas, so you don't come back.
Which is all ironic, obviously very ironic, considering it's all supposed to be easier, but in fact it's harder.
Yes. But the next step for that really is—so that's great. I mean, there's still some advanced use cases where it's fantastic. People want to mix data and do things. That works great. However, in most cases, what the models and the AI should let us do is get ahead of the issues and ahead of where the customers are going and help them get there faster. And so that's the next generation of that. We announced that at DASH. I'm going to fill your AI bingo card, but this is the agentic workflows.
But the idea there being is the machine is in its own loop and it's going to figure out what to do next and tell you about it. You don't have to ask. And so there's interesting ways for us to do that.
Is that experimental? Is that working?
Well, it is experimental and working. It's not something we've rolled out widely, but the idea there is, say you get a page because something broke on your application. And by the time you get on Slack, the bot is there already and is telling you, "This is what I looked into. I checked this piece of data, this piece of data. This is a notebook where I put my investigations. You can follow what I did there. I think this is probably that."
My recommendation is to restart the service. Do you want to click here and restart it?
Turn it off, sir or ma'am.
Yes, exactly. This is really exciting. It really changes the interaction. It also watches what people do on Slack. So when we ask a question about something, it can say, "Actually, I found the data for that, and this is what it is." Again, I think there we still need to make sure it works well enough in enough of the cases, because the models themselves are still imperfect enough that sometimes you don't get exactly what you want. And also, we need to make sure we find the right form factor for the interaction with the user.
What's not enough? What's too much? You don't want to end up with Clippy. And the risk right now is that a lot of the assistants we're seeing from the big companies are too close to Clippy. For the young people who listen, Clippy was the Microsoft Word assistant.
Maybe to make sure that we at least touch upon it, like the foundation for all of this, the reason why you have all this data and are able to leverage the data to feed it into the AI is because you have this unified real-time data platform, which sounds like a complete monster in terms of what does it do? Like, trillions of data points per hour. How is it architected? How does one build a platform that's able to have that kind of performance?
So first, there's the way data is organized. And for that, from day one, we had decided it would be a broad platform. And it would load data from many different sources, most sources we didn't know at the time, with different sizes, different velocities, different shapes. So we organized everything so that we could hook together many different data stores under one fairly flexible data model. So that was the, I would say, the good idea in the beginning that helped us do that, even though it stood in the way of us getting into Y Combinator.
After that, I think the secret has been we just keep rebuilding all of those modules all the time. Obviously, we're not running the same data stores right now, which are getting billions of data points every second, than the things we had on week one, which were running on just one database. And that was extremely naive in terms of the infrastructure. But you're right that it's an extremely high-constraint world, especially since we have to keep that going 24/7. And anytime we have a delay of a few seconds, it actually impacts our customers.
So the bar is fairly high in terms of what we can do there. And the repercussions on the AI side is that the bar is fairly high also in terms of what we can incorporate on the AI side, what the cost of it can be to operate at this extremely large scale, and also what the latency of it can be.
But in layman's terms, is that like a gigantic cluster in the cloud that just cranks this whole thing, and you have connectors that feed data into it and APIs to push data out of it? Is that roughly the idea?
Well, I mean, there's a number of queues and data stores, basically, and some of it is completely homegrown after generations of iterations, and some of it is still open source off the shelf that works very well for us. We use a lot of Kafka, for example, that works well, and we still use a lot of Kafka. Database-wise, we mostly built on—I mean, we still use, in some areas, standard databases like Postgres and things like that, but we obviously shard them and scale them a lot.
But a lot of the core data stores, like the time series, the log data, the event data, all of that is on completely custom data stores that we've built over time. We publish, actually, a series of articles on our new event store, which we use for logs and traces, which is called Husky, another dog name. And people can go on our blog and read about it. We share some of the technical details behind it.
And you store all of this, right? I think I read that there were some improvements around storing log data and that kind of stuff, but you just store massive amounts of historical data as well on behalf of your customers.
Yeah, and we did some things there. I mean, look, the biggest challenge with observability is that any application can generate an arbitrarily large amount of logs. And so the data volumes grow much faster than our customers' revenue, which is a problem. And so to solve that, there's a few things to do. One is you create the right feedback loop so people understand what they need and what they don't need in terms of data produced by the application. And when they send too much, you can fix it.
But the other one is also you just need to be more and more efficient in terms of how you can send more data and store more data, with it costing less. One thing we've done over the past few years is we decoupled the storage from the compute. So it allows us to store data, which is much cheaper, differently from the compute, which is a lot more expensive.
And you have about half of the team that works on that platform?
So, yes, the breakdown is roughly half of our engineering team is on the platform and half is assigned to specific products.
And for the AI stuff, did you have to build a mini lab with some AI PhD types?
So we don't have a lab. In general, we're a little bit careful with labs. I think it sets the wrong expectation. And we've been through the same iterations as most companies, which is you centralize it a little bit too much, then you decentralize it a little bit too much, then you recentralize it a little bit. So you sort of go in between those two to make sure you get the right amount of fundamental and common stuff, and at the same time, the right amount of work that is directly applicable to the real problems of real products and real customers.
But we've been building that. I think the main difference in terms of the way we invest in and build the AI products compared to the others is that because that part of the market is moving so fast, it's often a little bit more speculative. You have to start building and experimenting before you know exactly what it is you need, just because, one, the customer doesn't know yet, but two, you also need to learn as you go. So I think we place the dial a little bit differently there than we might in some other parts of the company.
Incredible. Well, look, this has been a wonderful conversation. Maybe to close, and going in a completely different direction, but since you mentioned PG, and by the way, it's wonderful to hear that even YC makes terrible mistakes and terrible passes from time to time, that they don't just have mega hits. But since you mentioned PG a couple of times, and I heard you say either in our conversations or in other things I listened to that you, at some point, were a little bit of a control freak.
What's your take on founder mode?
Look, I think most of what's said in there is fairly obvious. I mean, you don't just tell people, hire people, and then let them go and do whatever they want. You actually have to care about the details. Yes, you do. I think that part's fairly obvious. My worry about founder mode and everything that gets boiled down to a very short piece like that is that I think it's going to be used and abused in all sorts of different ways because there's so much context that goes into every single word in there.
When you say manage someone, what does it mean? It's going to mean very different things for very different people, or detail is going to mean very different things for very different people. So I think, as is, my worry is that it's going to do more harm than good in terms of the impact on the ecosystem. I've already been on the receiving end of a few people quoting founder mode to disagree with things that I was doing. So we'll see where it goes.
But look, the short of it is, yes, of course you have to care about the details. I think that's how you build a company. And I think that there's no way to remain relevant unless you do that.
Okay, amazing. Thank you so much for doing this.
Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.