So let's talk about lakehouses and all the things. So again, in a sort of effort to make this interesting for a broad group of people, and the hardcore data engineering people can fast-forward through this, but can you compare and contrast data lakes versus data warehouses versus lakehouses?
Sure. So the term data lake really originated in the, call it maybe 2011, '12 timeframe, when Hadoop was gaining momentum. It was described that way because you could basically store anything in it. It was just this distributed file system that was infinitely scalable. You could store infinite amounts of information. And from that, people started to say, hey, we want to do real analytics. We've got to actually optimize the way we store the data, the way that we lay it out in this lake, in this file system, and started to build new file formats that allow you to basically store the data in the same type of way that you would a traditional data warehouse.
So there were early files that are still very popular today called Parquet files or ORC files or Avro files, which were basically columnar representations of the data. And without getting into too much technical detail, storing the data in that way delivers faster read performance. Back to the opening point about analytical systems, you really need fast read performance. And so how you store and lay out the data has an impact on that. And so people kind of recreated the storage of traditional data warehouse platforms, but did it in a lake.
And that's really where the lakehouse evolution took place, which was doing data warehousing activities now in a lake and calling that a lakehouse. And I think that is what has really evolved over a 15-year period. In the early days, again, when I was working on my first company, the gap between doing analytics in a data lake and doing analytics in a data warehouse was pretty significant. Teradata was just dramatically faster than what you could do on Hadoop at the time.
But as a result of the evolution of those file formats and the query engines themselves getting faster and faster and faster, 15 years later, that performance gap is de minimis at this point. And that really has changed the game. And I think now lakehouses are the future. You even see the very traditional data warehouse companies embracing open formats and data lakes, and the rise of Iceberg, which I'm sure we'll talk about at some point, really playing an important role in that as well.
So to play it back, you had the lake, which was super scalable, and you could dump anything in it, but you couldn't really run analytics because it's not super structured. On the other end of the spectrum, you had data warehouses, which are very structured, sort of designed for analytics, but maybe less flexible, arguably, although—
And yeah, almost exclusively all proprietary, right? That was always the biggest complaint when I was at Teradata. It's an amazing database. I will still say that now, seven years later, it's an amazing database. There are things that that system can do that really nobody else on the market can do. But it's proprietary and it's expensive. And I think that the evolution of big data as a term and as a concept, like just the volumes of data growing so much and each of our businesses now becoming so data-driven, necessitates scalable architectures everywhere you go.
And data warehousing is certainly one of those. If all my data's trapped in a proprietary system, that really inhibits my ability to scale and grow and do the types of analyses I need. And so, hence, I think open source plays such an important role in data today.
Yeah. So data lakes on one end of the spectrum, cloud data warehouses, and then lakehouses meant to be the best of both worlds in the middle, combining the advantages of both.
And so obviously the most famous cloud data warehouses, just again, put names in the categories—and maybe I'm too obsessed with this as the guy who does all the industry landscapes, but I like categories and logos in the categories. But in the cloud data warehouses, obviously you have Snowflake, Redshift, Google BigQuery. In the world of data lakes, who do you originally—Hadoop at some point, early Databricks.
Dremio, and then Starburst.
And then Starburst. But then that group evolved to the lakehouse architecture.
Yes, that's right. Or, yeah, I would argue we've been sort of focused on the lakehouse from the beginning, and Databricks has evolved into that.
That's right. Because you always had that vision that storing all your data in one place is a terrible idea. And you were always the federated query engine, and the core value proposition—my words, not yours—where it's like, just leave your data anywhere and we'll find it wherever it is for analytical purposes.
Yeah, I think where our view on that has evolved slightly is we do think data lakes are where you're going to want to store as much data as you can, just because the economics will drive that. But you'll never store everything, to your point. And I used to use, actually, when I was raising venture for Starburst, I used to use your landscape all the time because especially when you show the progression from 2014 or whatever to today, 2025, you see the Cambrian explosion of data sources.
And so being able to federate out to those has an advantage for certain use cases, for sure.
Yeah, it's become the claim to fame of that landscape, is to show the absurdity.
It's not exactly what it was meant to do, but that's what it's become. Okay. And is it fair to say, and I think you just alluded to it, that between those various buckets, the lines have started blurring as well? Because people have federation as well on top of being the repositories and different things. Is that fair?
Lines have definitely blurred. Yeah, I think part of it is the industry's starting to mature. You have a couple giant players in Databricks and Snowflake that are now looking for adjacent markets to continue to grow their TAM and justify their valuations and drive their revenue into the future. And that's leading to more overlaps into boxes that they weren't typically part of.
And so how do you position today, like versus those and I guess Dremio and also the hyperscalers, BigQuery? How do you position?
Yeah, I think it necessitates being even crisper on the differentiation between those things. And so for us, there are a few things that we point to with customers. Number one, we're one of the only hybrid players. So Databricks and Snowflake are cloud-only. So if you happen to have data on-prem, we're pretty much your only bet. And it just so happens that that turns out to be most of the Fortune 500.
Almost the entirety of the financial services sector in particular. And so we do a lot of business in those industries as a result. The second thing that helps differentiate us is the openness of the platform. So we're an open engine querying open formats. And while there's been widespread embrace, I would say especially last year in 2024, around open formats and Iceberg in particular really winning that format, that's new. That's new for this industry. We've been doing this forever, though, and the first queries run on Iceberg were Trino or Presto queries.
So that pairing in the open-source community of Trino and Iceberg has been kind of a reference architecture for years at this point. And that gives us an advantage because we can help manage your Iceberg deployment holistically, everything from streaming ingest, loading that data into Iceberg tables, maintaining it, doing things like compaction, data maintenance, data management, and then, of course, querying. We have that sort of end-to-end lifecycle around Iceberg, and we call that the Icehouse. So that's our play on lakehouse: the Icehouse.
And then the other area where we differentiate, which you touched on already, was federation and being able to query into other data sources. And there are some use cases where being able to join a table that lives in another database system with a table that maybe sits in the lake can be very valuable in getting a fast response time. So we serve anti-money laundering use cases, fraud detection use cases, where very often the patterns to detect bad behavior actually exist in more than one data source.