MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Why Influx Rebuilt Its Database for the IoT and Robotics Explosion

    Evan Kaplan is the CEO at InfluxData. We cover why InfluxDB rebuilt in Rust to handle cardinality and separate compute from object storage, why the company abandoned Flux in favor of SQL, and how Tesla Powerwalls use time-series data to trade energy in four-second intervals.

    05/08/2025

    Hosted by Matt Turck · with Evan Kaplan, CEO, InfluxData

    time series databasesIoTRustSQLopen source
    Listen now
    YouTubeApple PodcastsSpotify
    36 min · 7 chapters
    Contents

    Transcript

    The InfluxDB origin story and why time series matters

    2:22
    Matt Turck1:44

    Thanks for being here. This is the third time that we have InfluxData at the event, and it's actually interesting. We have your CTO and original founder, Paul, here in the audience.

    Matt Turck2:06

    And I was looking this up. The first time we had Influx at the event was actually in 2015, which is kind of crazy. And then we had Paul back in 2019, I guess around the time of InfluxDB 2. So, as it turns out, it takes a long time to build those great companies.

    Evan Kaplan2:08

    It's an overnight success.

    Matt Turck2:40

    Yes, exactly. And part of the reason why it's super interesting to have you guys tonight is that the big news is InfluxDB 3. So we're going to talk about all of that in detail. But first things first: what is Influx, and maybe what is time series? Again, in an effort, as in prior conversations, to make this broadly interesting to folks that may not follow every single detail of this space.

    Evan Kaplan3:08

    So first of all, I should make Paul stand up and take a bow. So if you have any questions, Paul is the original author of InfluxDB. He's a native New Yorker, and I love telling this story because he can't stop me from telling it in my own way. And I met Paul in 2015. I was working at a venture firm, helping out with other CEOs. Paul and I connected on something totally unrelated to technology: we were both CrossFitters. And so, those of you who know the joke, if you've ever met a CrossFitter, they tell you in the first 30 seconds.

    Evan Kaplan3:39

    So we bonded around that. And over time, I met him just after he spoke here. And over time, he asked me if we would work together. And I said yes. And so we've been married now for almost 10 years. And I could do a whole speech on what it means to have a technology partner and a CEO coming together on that. But anyway, Influx started in 2013. And Paul had, as many companies, originally started to build basically a Datadog-kind-of server-monitoring SaaS service, and then had to build a database to service that, and realized in a pivot that the database actually was more important than the SaaS service that they could build, and built Influx. It was built in Go, and it was built specifically for the time series use case.

    Evan Kaplan4:34

    Paul had done a lot of work on Wall Street and had done multiple implementations of time series databases and felt like he could build one with no dependencies. He was deeply committed to the open source community, which he remains today. It's so funny actually saying this with him in the room. Usually I say it without him in the room. But he was deeply committed to the open source community. He ran the machine learning meetup in New York, which is probably how you met him, Matt, originally.

    Evan Kaplan5:08

    And so he built this product, and it was immediately popular. When I met Paul, there were about 3,000 users out there using the open source of InfluxDB. By the time I joined the company, there were maybe 7,000. Three million people use it daily that we can see. And so I shouldn't say people. We say addresses or locals, but it's a rough approximation. And so lots of hobbyists, lots of home users, lots of folks. And the primary benefit, I think the original idea that was so brilliant was: build a package that's really easy to use, build the TICK Stack, which included the collectors, the visualizations, those sorts of things, and make it super easy to build something powerful that eventually became something that enterprises would use.

    Matt Turck5:35

    Great. Fantastic. Maybe the 101 on what time series—

    Evan Kaplan5:51

    So think about it as—the easiest way to think about it is sensor analytics, but anything that's measured in time, anything you want to measure, any telemetry that you're capturing over time that you want to describe: time, time, what happened, what happened, what happened, what happened. So in a very broad case, anything that's really observable that you want to observe over time, it turns out that most datasets are really easily understood in the concept of time, or most usefully understood in the concept of time.

    Evan Kaplan6:10

    And so if you build something optimized for time series, it can be really useful, particularly as it relates to the physical world. Yeah.

    Matt Turck6:20

    And so I guess you sort of just alluded to it, but why do you need a separate database for that kind of data? Why can't a general-purpose database do this?

    Evan Kaplan6:38

    There's always just two kinds of people in the world: those who believe there are two kinds of people in the world and those who don't. So yes, in some cases, for many use cases, you can use a general-purpose database. People have built stuff on Postgres. People have built time series on Mongo. Those are highly useful. But if you're ingesting huge amounts of data, you want the query response time, you want zero time to read, you want low-latency data, you want a truly operational platform, you get to a level where the thing that you use to monitor your swimming pool or your home thermostat isn't the same thing that you'd monitor 50,000 Powerwalls in a locality.

    The cardinality crisis and why Influx rebuilt in Rust

    6:59
    Matt Turck7:16

    3.0, which effectively you all just announced. Congratulations. So from the outside, it looks like a major effort at rewriting a bunch of things. Why did you do that in the first place? What is it that you did?

    Evan Kaplan7:46

    3.0. I don't remember when the first repository opened. Let's call it three, three and a half years ago. So it's been a really big project. And there were certain assumptions after being in the market for, at that point, seven years, and probably at that point having 1,200, 1,300 customers. We just learned a lot about what customers were doing and what issues they were having. And so we set out to solve, let's call it, four or five problems. So the first problem we set out to solve is the issue of cardinality.

    Evan Kaplan8:02

    So when you get a database that's describing a rich set of data around a time series thing, you can run into cardinality problems that can choke off the performance of the database.

    Matt Turck8:11

    Do you want to define what cardinality or high cardinality is for us?

    Evan Kaplan8:39

    So think about a measurement, and think about all the metadata around the measurement, and think about ephemeral measurements like network metrics, source and destination. You get these millions and billions of combinations that are really unique to time series. And so when you start getting those, they can really choke off a database and kill performance. So what we found is at the upper bounds of our InfluxDB 1 and InfluxDB 2 databases, performance would slow down. And customers would have to cut out some of the measurements.

    Evan Kaplan9:07

    And they would have to pull back on the richness of their data, which is not what you really want. We had to solve the cardinality problem. And so that required a different look at the architecture. Two is we had to solve the storage problem. So what was happening is our compute and our storage were linked. And so customers who wanted to keep their data around a long time, it could be really expensive. And so if you were keeping—we had some customers, sort of utilities, things like that, that had to keep data for 10 years.

    Why SQL won (and Flux lost)

    9:26
    Evan Kaplan9:42

    And so all of a sudden, compute and storage becomes a really expensive burden. And so we wanted to have an object-storage-based system, which was non-trivial. To build real-time object-storage-based storage is non-trivial. And so we embarked on that. The third thing was—and Paul came here and talked about it in 2019—we took a hard attempt at building a functional language for time series that was open source. So completely open source, MIT license, called Flux. And it was a really powerful language.

    Evan Kaplan10:11

    It did a lot of different things. It was functional. It was modeled somewhat after the Prometheus model, but much deeper for the kind of tasks and things we were doing. And frankly, there was a learning curve to it, but it wasn't big. But it just didn't see the adoption that we wanted, that we really wanted it to. And so a language without much adoption doesn't serve any commercial purposes, and long-term doesn't even serve developer purposes. And so there weren't a lot of contributors to Flux other than us.

    Evan Kaplan10:38

    And so it was pretty clear. I think if our assumption was—and I'm guilty of this too—if my assumption was back in '18 or '19 that SQL couldn't be the language of the future, like, just couldn't be. Like, how could a language—I think Paul even wrote a blog post—how could a language written 40 years ago be the language of the future? And listening to Mike's talk, we still use Excel from 1985. So there was a little bit of humble pie for us to eat because it was pretty clear that the language—and this is right around the time where you really started to see the clarity of Snowflake.

    Evan Kaplan11:14

    And you started to see the clarity of the emergence in Databricks. Ali Ghodsi is an advisor to the company, has been a friend for a long time. And you could start to see the clarity that maybe we were wrong on SQL. Maybe we should support native SQL because that is the lingua franca for a lot of this stuff. And so we had to solve that SQL problem. So we had cardinality, we had object storage, SQL. But this next point may be the biggest point.

    Evan Kaplan11:39

    We had the assumption that the world was fundamentally changing around what really is a database. What is a database? And I think I'm old enough that I grew up in an industry where really there were two main databases. There were Oracle and IBM DB2. And then there was some Informix and Sybase and these side cases. But if you ran any sort of enterprise, you had Oracle or IBM DB2. I think now our assumption is that we live in a world where there are two main platforms and some variants.

    Evan Kaplan12:16

    And those two main platforms are Databricks and Snowflake. Those increasingly look like databases to us. They're called lakehouses. And so they combine data warehouses and the old notion of data lakes, but they look like databases. And I think, from my historical perspective on the industry, I think lots of capability that sits outside those things will fall inside those things, whether it's data governance, whether it's different ETL, whether it's different conversions. These things are giant sucking sounds. And then there are the Fabrics of the world, the Redshifts, Athenas, and the BigQueries, which also have those same dynamics.

    Evan Kaplan12:47

    And I think that's where increasingly we're going to start. And so we had to ask ourselves the fundamental questions. And like, okay, what do we do? What does a specialized database do in a world where lakehouses are dominant, where they become the sinks for all data and a lot of functionality? And we got very clear about what we needed to be when we grew up, which we're still growing up. And that is, we needed to be an operational time series database.

    Evan Kaplan12:56

    We need to really be oriented operational, which to us meant we need to be great at three things.

    Matt Turck12:59

    What does that mean, operational versus transactional?

    Evan Kaplan13:29

    So most of our customers use us for two broad use cases: real-time monitoring. So you're looking at systems and behavior and alerting on actions that are happening in real time. And so you want your queries to be ideal queries, your last queries to be sub-100 milliseconds, sub-50 milliseconds, and in some cases, more. So not fundamentally an analytical tool. Obviously, you do analytics coming out of the database. But fundamentally, you have to be operational in real time. That had to be because we felt that the lakehouse architectures would never do that well or weren't structured to do that well for the very nature of it.

    Evan Kaplan13:58

    And so we needed to be really great at three simple things. And this is what we say to all of our employees. We need to be amazing at ingest, which means we need to be able to handle a huge volume of sensor data coming in at a fast time. And that stuff has to be available for query very, very fast. Two is we had to be great—and by the way, we're not great at all of this, but we have to be great at all of it.

    Evan Kaplan14:27

    We have to be great at organizing the data in a proper location, in object storage, as an index, in cache, potentially on disk, whatever it happens to be, so that data is available. And that's no small task. If you want that data to be queried fast, you have to be super good at organizing it. And then the last thing, you have to be great at querying the data. And so if we're great at those three things, we feel like we have a very defensible and very relevant platform.

    Evan Kaplan14:56

    If we're not great at them, use Mongo, use Postgres, use something else. And I would say the same for all the specialty databases, whether it's Elastic, whether it's Neo4j. I'm not sure how I feel about the vector databases. But whatever these open-source category leaders are, they have to find a place where they have to live in this lakehouse world. So to answer your final question, the fourth vector we said is, how do we live in a lakehouse world?

    Evan Kaplan15:25

    And so I think we started building this before Delta Sharing and Iceberg were obvious. But we built it around Parquet because we knew that's the file format that the lakehouses would use. And so our notion is we had to live in this world. And our data, whether downsampled or live, had to end up in those lakehouses. So those are the four things that drove it. And I think if Paul were up here, he would talk about the value of Rust.

    Evan Kaplan15:40

    We moved from Go to Rust. We've been able to draw really good developers because we're really pushing at the edge of Rust, both the open-source and our closed-source components.

    Matt Turck15:56

    It's super interesting in so many ways, including having the kind of self-awareness and then will to just burn the boats on a few things and say, well, the world has changed and we need to change dramatically, and then releasing. Kudos to you guys.

    Evan Kaplan16:07

    In retrospect, it sounds like it's like, yeah, it was just a really smart architectural decision. But if you saw Paul and I fighting like cats and dogs, you'd go like, oh, these are really hard decisions.

    Matt Turck16:12

    So the idea is to live on top of S3. Is that the right—

    Evan Kaplan16:16

    S3 or MinIO, anything that would have an object storage interface.

    Why UnfluxData bets on FDAP

    16:34
    Matt Turck16:47

    So you went from disk to basically diskless and sitting on top of S3. Okay. All right. That part, which, by the way, I mean, that's the reason why you did it, but it seems to be such a fundamental trend in the industry of having S3 as a default storage. So that, and then for the sort of Databricks lakehouse, Iceberg kind of world. So first of all, I saw that you guys coined your own acronym. What was it?

    Evan Kaplan16:48

    FDAP.

    Matt Turck16:57

    FDAP, right. So there used to be the LAMP stack, and this FDAP. And FDAP is what? It's Flight, DataFusion, Arrow, and Parquet.

    Evan Kaplan16:58

    You got it. That's right.

    Matt Turck17:05

    So maybe walk us through some of the components at a high level and what role they play.

    Evan Kaplan17:27

    So the idea was to build on largely an Apache stack. We felt very—because of the lakehouse, because of the view of the lakehouse and the view of what a modern data architecture would look like, we decided, okay, how do we—and I say we. I'd say Paul and the technical team decided. I was there to witness it—but decided to build. So building in Rust is a start. But then we built on object storage because we knew object storage would be—it's technically not open source, but anything that's S3-compatible gives you a tremendous number of options.

    Evan Kaplan18:02

    So it might as well be. And then built on Apache Parquet. Those are the standards that the lakehouses are built on, and things like Redshift and BigQuery. And so that was an advantage. Then build on above that the notion of the table format, the open table format being Delta Lake or Iceberg. And now it seems like Iceberg is certainly, with Databricks buying Tabular, that'll be interesting. But Iceberg feels like the base format, although Delta Lake is also relevant.

    Evan Kaplan18:33

    And then above that, the interesting product is DataFusion. We're one of the PMCs for DataFusion. DataFusion is a query engine and query planner, super-fast, vectorized. It's available. It's a super popular open-source product for—it's got a lot of love for something that really doesn't have an end product. You wouldn't just take DataFusion and say, okay, I'm going to run this. It's built into other products. It's built into Apache Comet. It competes with some of the other query engines.

    Evan Kaplan19:03

    But the idea is to commoditize the query engine so that anybody can go out there and build services based on a really sophisticated vector query engine for a columnar database. So commoditize some of the stuff that you'd have to, if you wanted to use ClickHouse, you wanted to use other stuff. Our intent is to commoditize that by an Apache project. And then Flight is the RPC, and Flight SQL is the driver that allows you to run any SQL on top of it. And so that stack, all the way to its integration to the lakehouse at the bottom, at the tables, all the way up, is something that's driven by the open source.

    Evan Kaplan19:34

    And so we feel like that is long-term sustainable. If Influx were to disappear tomorrow, that stack would still be tremendously valuable. And the work that we do on that stack is we've done a lot of the work to convert that stack to Rust. And so we really kind of bet the company on that. So whether FDAP as a name, or FIDAP, catches on is sort of irrelevant. Those projects are really relevant, and they're going to be here a long time.

    Matt Turck20:04

    Very cool. And just to drive it home from a functional standpoint, where would InfluxDB sit in a lakehouse kind of world? Are you moving the data into the lakehouse? Are you analyzing data that sits in the lakehouse? Where are you in the overall kind of data movement chart?

    Evan Kaplan20:31

    I think the way—different people use it differently. And one of the things that developed conviction is, as we went out and talked to our customers, even the most industrial-oriented customers had built something on Hadoop and were now moving it to Snowflake or Databricks and doing it. And so what happens is they run their operational data, they run their dashboards, they run their queries, they do all that stuff in real time. And then our view of how the future—it's just starting now.

    Evan Kaplan21:00

    It's just really starting despite all the hype. Our view is they're exporting a lot of that data into the lakehouses. And the idea of using Iceberg is you effectively have zero ETL. You never really have zero ETL, but you kind of do. You kind of have zero ETL because you can pull on the Iceberg tables directly without moving it. But the intelligence, where the intelligence gets built, is really important. We don't think the intelligence gets built on our platform.

    Evan Kaplan21:32

    I know that's a hard thing to say because everything's AI now. We believe that we're fundamentally mining the data and making it available so that the intelligence models, which are largely going to run out of the lakehouse, where the really structured time series data comes together with the unstructured data and potentially the business data, and you build your models based on that. We think the intelligence gets built at that layer above us. But then the application of that intelligence to real time, we do.

    Evan Kaplan22:04

    Because the application of that intelligence is effectively a query. It's effectively how quickly can you execute a query to change a behavior in a system that's live. Most of our business today is, in fact, real-time monitoring, which is interesting and important. But the real opportunity is control: real-time monitoring control. You build control systems based on this data. Now, in order to build a control system, you have to have that intelligence. You have to have the model and the inference there so you know what to do when the query executes.

    Evan Kaplan22:26

    So in a very real way, we view ourselves as the senses—eyes, ears, nose—of data coming in, pulling that data in, and then the action, the hand that takes the action in the real world. And this is appropriate because robotics is obviously a big place for us.

    Matt Turck22:40

    Because the data about the real world will come to the robot, and you'll feed it into your model, and then you'll send it to—for the action part that you were just describing?

    IoT, Tesla Powerwalls, and real-time control systems

    22:51
    Evan Kaplan22:51

    We'll put the data in the model, build an inference engine where a really smart model will be running, and we'll take action based on what the actions are taking. And those latencies need to be timed.

    Matt Turck23:19

    So this is the—I was going to say the future world, but it's moving all so fast that it might be tomorrow morning. In terms of more pedestrian use cases of machine learning, it sort of feels like if you're in an IoT world, there should be a bunch of use cases where, okay, well, if temperature goes up by a certain whatever, recognize that pattern, machine learning model, and then do something. Do you have a bunch of that already?

    Evan Kaplan23:46

    Yeah, that's what most people do with our stuff. But I would call those pretty simple control loops. A good example of that is the Tesla Powerwalls. I think there are roughly a million of them out there now. They all spew into InfluxDB. And they can trade energy in a four-second interval based on the data that's coming off. So in my house, I have three Tesla Powerwalls. And so I can sell my energy back to the grid based on that.

    Evan Kaplan24:18

    It's all based on InfluxDB, and the app runs. But that's the kind of stuff that people are increasingly doing. Because our fundamental belief is that any human-designed system wants to become increasingly autonomous. Doesn't mean it has to be. Doesn't mean it will be. But it wants to be. Anything that you can do, it wants to be increasingly autonomous with less human intervention. It's a long journey.

    Matt Turck24:30

    And implied in the robot example that you just used, there is a concept that you can work in a sort of embedded manner and you can sit on hardware devices. Maybe talk to that a bit.

    Evan Kaplan24:47

    Yeah, we're really early in that. I think obviously people have been putting us on Raspberry Pis for a long period of time. The new version 3, the open source, can probably be embedded. Obviously, Telegraf, which is our most popular open-source project, is embedded in a lot of different places.

    Matt Turck24:50

    Telegraf being a data capture—

    Evan Kaplan25:06

    It's the collector, yeah. So Telegraf is part of what we call—and the TICK Stack is, I think, 350 or so systems it can collect from. And it puts it in line protocol, allows it to go in the database. Super popular. Microsoft uses it. Amazon uses it.

    Matt Turck25:09

    That's part of the TICK Stack you already described, right?

    Evan Kaplan25:10

    Part of the TICK Stack.

    Matt Turck25:13

    Okay. And T-I-C-K?

    Evan Kaplan25:27

    C and K. So C is Chronograf, which we support, but we don't actively develop. Most people use us with Grafana today, or Superset increasingly, or some of these other tools that are emergent.

    Matt Turck25:29

    As the visualization and analytics tool.

    Evan Kaplan25:57

    Visualization. And Kapacitor—and I'm glad you bring it up—Kapacitor was our, let's call it, task engine built around the database. But version 3 of the database has an embedded Python VM with a set of triggers that allows you to do a variety of stuff at the database, whether it's alerting or taking an action or taking a system action, things like that. It's actually a really cool thing. We're super excited about it. And that's available in the new open source.

    Evan Kaplan26:01

    And that replaced Kapacitor.

    Matt Turck26:13

    And by the way, as an aside to this discussion, a lot of it is—you mentioned monitoring, as in infrastructure monitoring, but there are a bunch of IoT use cases.

    Evan Kaplan26:14

    Yeah.

    Matt Turck26:37

    And I'm very curious about that because the IoT, for people like me as a VC that get sometimes carried away in the hype, the Internet of Things, IoT, was like this big idea in 2000-whatever, '14, '15, '16, '17. And I'm very curious what your view of the reality of it might be today.

    Evan Kaplan27:06

    Well, if you'd asked—since my background wasn't databases, it was networking and security—one of the primary reasons I started Influx in the first place, at the time, 4G networks were just starting to emerge and things were very much connected. And so IoT felt like it was going to be gigantic. And it wasn't gigantic. It was more of a buzzword. But I'd say now, here we are, it's 2025, and probably 60% to 70% of our business is now IoT. And I almost—we have to use the word IoT because it's well understood, but it really is about sensor analytics.

    Evan Kaplan27:36

    And so when we think about our long-term businesses or the time-series space in general, we just think to ourselves, are we going to see more or less sensors in the world? Are they going to be collecting more or less data? And is that data going to be richer? And is that data going to be important? And is that data going to feed the AI models? And so we think a lot about how do we service those workloads coming up.

    Competing with Databricks, Snowflake, and the “lakehouse” world

    27:54
    Evan Kaplan27:54

    And by the way, we say sensors as physical things, but software sensors too. Lots of our business is network telemetry. It's API monitoring, which basically behaves like time series sensors.

    Matt Turck28:04

    Let's get into the fun territory of competitive differentiation, to the extent that you are willing to badmouth your competitors.

    Evan Kaplan28:05

    I'm not willing to badmouth.

    Matt Turck28:30

    All right, well, so we'll remain at a high level. But time series is a very important category in the data world. It's also a competitive category between some products from startups like Timescale and hyperscalers and open source. How do you think about it? How do you position to win?

    Evan Kaplan28:54

    We don't win anything unless people love our open source. So we don't have an evangelical sales force who comes in and finds your front-end people. Our customers are data engineers, people who have to solve problems. So it's not useful for me to have a meeting with the CIO. The CIO does not care about his time series database. I mean, he should maybe, but he doesn't. It is very useful for us to have a meeting with an architect, a CTO, a developer.

    Evan Kaplan29:23

    And so all of our stuff comes from developers liking our open source and us developing around it. And so that differentiates us from a lot of other folks. The people who we see as primary competition is, of course, the substitutes. Like, my problem is not that big. I can use a Postgres database. I can use Mongo's time series stuff. It's perfectly workable. It's fine. I understand that. But when we see direct time series competitors, it tends to come from the hyperscalers because those are easy buttons.

    Evan Kaplan29:57

    I'm already in Amazon. I can use Timestream. I could use Azure Data Explorer. And I can use those things. That's our primary competition. Other players we see. But let's be honest. As an open source vendor, we don't know what's happening. We don't even talk to a customer unless they've been running our stuff for three months, generally, or they've built something. And so we don't get to see the early part. So we have to win the hearts and minds early in the cycle.

    Evan Kaplan30:29

    And the way we have to win the hearts and minds is: is it relatively easy to get started? Is it relatively easy to build? And do I see a way to scale it? The interesting dynamic that's happened for us recently that people might be interested in is, last year, we did a very unique deal with Amazon where they have Amazon Timestream, but their customers kept asking them for InfluxDB. And so they approached us. They could have forked the database. We have a very strict view that open source should be open source.

    Evan Kaplan30:59

    It should be MIT or Apache, and people should be able to use it freely. And companies should be able to build on our stuff and even compete with us. Because if we can't win with our own code, then maybe we don't deserve it. But Amazon could have forked us. They chose not to. They built a relationship with us. They host it, and we build add-ons above that to monetize that. And we have a relationship with them. We do second-level support.

    Evan Kaplan31:26

    And the nice thing about that is now they're not willing to stand everything behind Timestream. They're very comfortable with people choosing InfluxDB. So we have one less dominant competitor. And two, it's a new model for open source. Most people who take open source actually spin up a cloud instance anyway to run it. I mean, a lot of people run it on a laptop, I'm sure, but a lot of people run it. And so I think for a lot of us open source companies, having these relationships really works because now you just go to Amazon, you run it up, and for the price of a server, you're running open source software.

    Evan Kaplan31:46

    InfluxDB. It's a good model. So we'll see how that develops over time. I think it's a change in the industry.

    Open Source lessons, monetization, & what’s next

    31:50
    Matt Turck32:01

    All right. And you started alluding to this, but lessons learned building this open-source company in terms of the community, the go-to-market?

    Evan Kaplan32:27

    I think one of the—actually, to quote something Ollie said a few years ago I thought was really on point—is, it sounds great. You build an open-source project, everybody loves it, then you monetize it, and you ride off into the sunset being super successful. It's actually really freaking hard. So with open source, you have to hit two home runs. You have to build a project that developers love, and they deploy and they use and they love. And then you have the second home run, which you have to figure out a model to monetize it.

    Evan Kaplan32:58

    And you've watched these open-source companies with all these different models. In the beginning, it was, you buy support. Well, after two or three years, I know how to support it as well as you do. I don't need it. Or you change your license where you say, oh, okay, no, we're going to prevent the hyperscalers from ever using this, or people using it for commercial purposes. And so people have played with all sorts of models to make it work. And I would say we didn't get it right the first time.

    Evan Kaplan33:25

    Three million users and only 2,600 paying customers. That ratio is—you'd like it to be a little bit better if we're going to be a really good-going concern. And so you have to figure out what makes sense to empower developers to build cool things, but to hold something back so that you can actually grow a business that can keep funding all of this open-source activity. And so it's tricky. It's two home runs. I thought that was a great quote.

    Evan Kaplan33:28

    Not mine.

    Matt Turck34:02

    I guess, what is next? You mentioned in the earlier part of this conversation that you wanted to remain focused on being an operational time-series database. I guess the broader question is: should specialized people remain specialized? There's a natural tendency to try and want to do more. How do you navigate that tension?

    Evan Kaplan34:26

    It navigates itself a little bit. So if you have a gigantic cash engine, you're generating a lot of business, then you bite off bigger and bigger chunks. And so we have a pretty good cash engine, but it's not quite at the point where I'm willing to bite off higher up the stack or do different stuff. But also the competitive dynamics in the industry—I think a lot of that stuff falls into the lakehouse. And you have to figure out where you can be elite and whether that space will be big.

    Evan Kaplan34:53

    And so we have two convictions, really simple. We believe we can be elite based on the strength of our own open-source efforts and based on the strength of the products that we've built upon. So we believe we can be elite. And we believe this space can be super, super large. We think we're not having less sensors. We think the physical world is getting instrumented at higher and higher rates. We think the requirements for AI and telemetry are going to be super high, through the roof.

    Evan Kaplan35:09

    And so we're just trying to stay the course and build something that we can create long-term sustainable competitive advantage with and still keep to our open-source grounding.

    Matt Turck35:10

    Thank you so much.

    Evan Kaplan35:11

    Really appreciate it.

    Matt Turck35:12

    Thank you.

    Evan Kaplan35:13

    Thanks.

    Matt Turck35:34

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.