MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Trino, Iceberg and the Battle for the Lakehouse | Justin Borgman, CEO, Starburst

    Justin Borgman is the CEO at Starburst. We cover why lakehouse performance has closed the gap with traditional warehouses after 15 years, why Iceberg remains independent despite Databricks’ Tabular acquisition, and why complex enterprise data infrastructure resists a true product-led growth motion.

    01/30/2025

    Hosted by Matt Turck · with Justin Borgman, CEO, Starburst

    TrinoApache IcebergLakehouseData infrastructureOpen source
    Listen now
    YouTubeApple PodcastsSpotify
    1h 6m · 24 chapters
    Contents

    Transcript

    What is Starburst?

    1:32
    Matt Turck1:30

    Hey, Justin, welcome.

    Justin Borgman1:30

    Thank you for having me.

    Matt Turck1:37

    You're the CEO of Starburst. Let's just start with the one-minute elevator pitch, just to frame the conversation. What does Starburst do?

    Justin Borgman2:07

    Sure. So Starburst is a data platform for analytics, building data apps, and now increasingly incorporating AI in the applications that you build. We're the creators of an open-source project called Trino, which is a pretty popular project used by a lot of the big internet companies like Netflix, Airbnb, and LinkedIn, and so forth. And it essentially allows you to run fast SQL queries on data in your lake, as well as connecting to other data sources. So you can really run analytics across all the data that you have, be it in a traditional database, a data lake, on-prem, in the cloud, you name it.

    Understanding the data layer

    2:32
    Matt Turck2:46

    So I'd love to start the conversation with a little bit of a sort of broad overview of the space to make this interesting to anyone that may be curious about the world of data infrastructure. And then we can go into all sorts of technical details. But to start with, what world do you operate in? If we think of all of this as databases—so databases, data warehouses—how would you sort of compare and contrast the various databases of the world?

    Justin Borgman3:06

    So I think of the database world as being really divided into two halves: the analytical side and the transactional side. The analytical side are systems that are built to be very read-optimized, so reading data as fast as possible. The transactional world is more oriented towards writing data fast and consistently. So when you think of a transactional system that's maybe powering an application, a canonical example would be like an ATM that needs to record a debit and a credit very quickly and has to be consistent every single time.

    Justin Borgman3:40

    An analytical application would be something like how many customers bought product X last year, and slicing and dicing the demographic profile of your customers and understanding their journey through the website all the way to a transaction at the end. And so those are, broadly speaking, the two worlds.

    Matt Turck3:52

    Yeah. And when you say writes, what matters most is to never lose any data, and the data needs to be 100% correct 100% of the time. Whereas analytical is—what do you optimize for in the read?

    Justin Borgman4:21

    Yeah, so analytical is maybe a little less—I'll call it maybe mission-critical—in the sense of losing a byte is not going to be the end of the world. That's the transactional side, which is very focused on ensuring that consistency. But on the analytical side, you have a different challenge, which is: how do I process massive amounts of information, do very complex joins of information across different tables, to get to an answer as quickly as I can? And that requires a lot of, I'll call it, rocket science in the optimization of how those queries actually get executed.

    Justin Borgman4:31

    And that's what we focus on on the analytical side.

    Matt Turck4:38

    And maybe to just, like, drive it home, some examples of famous company names in the first bucket and in the second bucket.

    Justin Borgman5:01

    Yeah, so on the transactional side, you'll sometimes hear people also call them operational database systems. So I think of operational and transactional as sort of interchangeable words to some degree. You'd think of Oracle or MongoDB as being two very prominent examples on that side of the house. On the analytical side, you'd think of Teradata, Snowflake, Databricks, and Starburst.

    Justin Borgman’s story before Starburst

    5:06
    Matt Turck5:06

    And Teradata acquired your first company.

    Justin Borgman5:06

    Yep.

    Matt Turck5:14

    Right. So you and I, we've done this, or a version of this, a couple of times. The first time, I was looking this up, was in 2013.

    Justin Borgman5:14

    Yep.

    Matt Turck5:20

    When you were running Hadapt. So tell us the story: your first company, what led you to where you are today?

    Justin Borgman5:48

    Yeah, my first company was actually a spinout from Yale University, and it was really commercializing the research of my co-founder, a guy named Daniel Abadi, and his PhD student, Kamil Bajda-Pawlikowski. The research was called HadoopDB, which was really like the first imaginations, I would say, of a lakehouse architecture, which is what's become pretty popular today. But back then, we were sort of the first to really do that. And this was around 2010. Hadoop was just gaining momentum as the data lake of the day.

    Justin Borgman6:04

    And we were some of the first to think about using this for analytical data warehousing purposes. And so we built that business over four years, raised venture capital, and eventually were acquired by Teradata, as you said. And it's funny how history repeats itself in a lot of ways, because the concepts of a data lake, the concepts of doing analytical processing on a data lake, were really things that were being played with and pioneered 15 years ago, but today have now become really the status quo, or what everyone is doing today.

    Justin Borgman6:40

    So it's sort of interesting to see how the space has evolved over that timeframe. And I'd say the biggest thing that has changed over 15 years is that lakehouse technologies have just improved dramatically in terms of performance and functionality.

    Matt Turck6:41

    Mm-hmm.

    Justin Borgman6:52

    And I think that just goes to show that building databases is hard, and it takes a long time to get quite there. So 15 years later, we finally delivered on the vision from 2010.

    Matt Turck6:55

    And so when did you leave Teradata to start Starburst?

    Justin Borgman6:57

    So that was 2017.

    Matt Turck6:57

    Yep.

    Justin Borgman6:57

    Yep.

    Matt Turck7:03

    And maybe tell us about the early days and how that worked.

    Justin Borgman7:28

    Yeah, so when I was at Teradata, my job—I—so Hadapt had been acquired. I became a vice president and general manager of a small business unit at Teradata that was really focused on next-generation technologies and thinking about the future of data warehousing. And this was at a particular point in Teradata's history where they were starting to feel the pain of both open source and cloud. Snowflake was still a young company. It wasn't materially impactful to Teradata yet, but you could see them on the rise.

    Justin Borgman8:01

    And companies like Cloudera were certainly taking workloads away and causing some pain there. And so my job was sort of to solve the innovator's dilemma, if you will, and figure out the future for Teradata. And so in that vein, I actually discovered a new open-source project, which is today known as Trino. Back then it was called Presto. So Presto later became Trino, and it was really how Facebook was running all of their data warehousing analytics. It's very interesting because Teradata was always trying to sell their data warehouse to Facebook and never really getting anywhere.

    Matt Turck8:09

    Yeah.

    Justin Borgman8:33

    And in fact, I had a lot of meetings with Facebook as a Teradata person at the time. And it was like, oh wow, okay, they actually do a lot of what Teradata can do, but they built it themselves with this open-source project, and they've run it at ridiculous scale: hundreds of petabytes of data, thousands of queries. And that got me really turned on to this idea that maybe this could be the future. And I'd already seen the power of open source, just watching Hadoop's rise and the viral adoption of that technology, that I thought, open source is a really powerful distribution mechanism.

    Justin Borgman9:05

    So this could be something. And that began, I'll call it maybe an awkward marriage of sorts, where my team from Teradata actually started working with the team from Facebook to contribute towards this open-source technology and advance it. We built the query optimizer for Presto, or Trino now. We added access control.

    Matt Turck9:06

    So this was while at Teradata?

    Justin Borgman9:06

    While at Teradata.

    Matt Turck9:08

    As, like, a service offering of some sort?

    Justin Borgman9:36

    Yes, yes. So the idea, at least my business plan, if you will, for Teradata was to turn this into an actual product line that would have the virality of an open-source project and that we could build an enterprise offering around it. Unfortunately, during my time there, Teradata had a number of changes at the executive level—the CEO, multiple CEO changes, actually, during my brief period there—and nobody could really get on board with what we were doing. I think they were all concerned that it might be competitive or cannibalistic towards the existing business.

    Matt Turck9:43

    Yeah.

    Justin Borgman10:07

    Which I can understand, those concerns, but that also kind of goes to the innovator's dilemma. Sometimes you have to eat your own lunch if you want to still have lunch. And so, long story short, we parted ways in 2017, and that was, I guess, serendipitous for me and my team because it allowed us to then create a business around this independently. And that's really how Starburst was born.

    Matt Turck10:13

    And then you co-founded the company with some of those top Presto contributors from Facebook.

    Justin Borgman10:15

    That's exactly right. Yep.

    Matt Turck10:21

    So you convinced them to just leave the mothership and strike out on their own.

    The evolution of Presto into Trino

    10:41
    Justin Borgman10:41

    Okay, exactly. And I think the project itself has just grown tremendously since that timeframe. So they've been able to realize, I think, a couple benefits there: A, the rise of a commercial business that's quite successful today, but also seeing the open-source project just become more vibrant and users continue to grow.

    Matt Turck10:50

    Yeah. And the whole Presto-to-Trino kind of evolution was partly because, at some point, Facebook said, "No, this is our open-source project."

    Justin Borgman11:16

    Yes, that was a tricky moment in our history just because it's so hard to create awareness of any brand name. And essentially, because Presto had been created at Facebook, they owned the trademark, just sort of the way intellectual property laws work. And so when my co-founders left Facebook and joined Starburst, they continued development of the community and of the open-source technology, but were still calling it Presto. And that led to a sort of legal challenge where they had to rename it.

    Justin Borgman11:24

    And that's why PrestoSQL became known as Trino today.

    Matt Turck11:24

    Yeah.

    Justin Borgman11:33

    And there was a rebranding, but it was tricky because you're sort of starting from zero in terms of name recognition at that point. Even today, there are a lot of people who still—

    Matt Turck11:37

    And the project. So Trino, at some point, had zero stars on GitHub or whatever.

    Justin Borgman11:38

    Exactly.

    Matt Turck11:39

    Rebuild from scratch.

    Justin Borgman11:46

    Rebuild from scratch. So that was, yes, definitely. If there's a book later about Starburst, that will at least have a chapter for sure.

    Matt Turck12:03

    Sounds like fun times. So how did you build it from there? Was it just like good old—I mean, clearly, I'm sure the community knew what Trino was, but then it was like good old developer relations and meetups and content to build the open-source project?

    Justin Borgman12:32

    That's exactly right. I mean, the hardcore Presto users immediately became hardcore Trino users and followed my co-founders in the transition. But outside of that group of, like, the people in the know, the broader market had no idea what Trino was. And so we had to recreate that awareness, which is very challenging. And yes, a lot of old-fashioned developer relations and meetups. Of course, we also had a pandemic right around this time.

    Matt Turck12:33

    Oh yeah.

    Justin Borgman12:55

    So the in-person element was completely taken out, which was a major bummer because I think in-person events are so great. That was a big challenge for our marketing team, our developer relations team, our community team. But fortunately, today I think the awareness of Trino has grown, and I think more and more people are finding out about it.

    Matt Turck12:59

    And is Presto still active, or is it completely different at this stage?

    Justin Borgman13:19

    Completely different at this point. That's right. They started as identical copies, but the code bases have diverged quite a bit. And today Presto is really just used by Facebook. So it's sort of like their own private branch, in a way, used by a small number of people. And Trino has become the mainstream community branch. Apple, LinkedIn, all the big guys are contributing there.

    Lakehouse vs. data lake vs. data warehouse

    13:20
    Matt Turck13:41

    So let's talk about lakehouses and all the things. So again, in a sort of effort to make this interesting for a broad group of people, and the hardcore data engineering people can fast-forward through this, but can you compare and contrast data lakes versus data warehouses versus lakehouses?

    Justin Borgman13:42

    Yeah.

    Matt Turck13:44

    To make it super clear.

    Justin Borgman14:09

    Sure. So the term data lake really originated in the, call it maybe 2011, '12 timeframe, when Hadoop was gaining momentum. It was described that way because you could basically store anything in it. It was just this distributed file system that was infinitely scalable. You could store infinite amounts of information. And from that, people started to say, hey, we want to do real analytics. We've got to actually optimize the way we store the data, the way that we lay it out in this lake, in this file system, and started to build new file formats that allow you to basically store the data in the same type of way that you would a traditional data warehouse.

    Justin Borgman14:52

    So there were early files that are still very popular today called Parquet files or ORC files or Avro files, which were basically columnar representations of the data. And without getting into too much technical detail, storing the data in that way delivers faster read performance. Back to the opening point about analytical systems, you really need fast read performance. And so how you store and lay out the data has an impact on that. And so people kind of recreated the storage of traditional data warehouse platforms, but did it in a lake.

    Justin Borgman15:23

    And that's really where the lakehouse evolution took place, which was doing data warehousing activities now in a lake and calling that a lakehouse. And I think that is what has really evolved over a 15-year period. In the early days, again, when I was working on my first company, the gap between doing analytics in a data lake and doing analytics in a data warehouse was pretty significant. Teradata was just dramatically faster than what you could do on Hadoop at the time.

    Justin Borgman15:45

    But as a result of the evolution of those file formats and the query engines themselves getting faster and faster and faster, 15 years later, that performance gap is de minimis at this point. And that really has changed the game. And I think now lakehouses are the future. You even see the very traditional data warehouse companies embracing open formats and data lakes, and the rise of Iceberg, which I'm sure we'll talk about at some point, really playing an important role in that as well.

    Matt Turck16:17

    So to play it back, you had the lake, which was super scalable, and you could dump anything in it, but you couldn't really run analytics because it's not super structured. On the other end of the spectrum, you had data warehouses, which are very structured, sort of designed for analytics, but maybe less flexible, arguably, although—

    Justin Borgman16:39

    And yeah, almost exclusively all proprietary, right? That was always the biggest complaint when I was at Teradata. It's an amazing database. I will still say that now, seven years later, it's an amazing database. There are things that that system can do that really nobody else on the market can do. But it's proprietary and it's expensive. And I think that the evolution of big data as a term and as a concept, like just the volumes of data growing so much and each of our businesses now becoming so data-driven, necessitates scalable architectures everywhere you go.

    Justin Borgman17:10

    And data warehousing is certainly one of those. If all my data's trapped in a proprietary system, that really inhibits my ability to scale and grow and do the types of analyses I need. And so, hence, I think open source plays such an important role in data today.

    Matt Turck17:21

    Yeah. So data lakes on one end of the spectrum, cloud data warehouses, and then lakehouses meant to be the best of both worlds in the middle, combining the advantages of both.

    Justin Borgman17:22

    That's right.

    Matt Turck17:50

    And so obviously the most famous cloud data warehouses, just again, put names in the categories—and maybe I'm too obsessed with this as the guy who does all the industry landscapes, but I like categories and logos in the categories. But in the cloud data warehouses, obviously you have Snowflake, Redshift, Google BigQuery. In the world of data lakes, who do you originally—Hadoop at some point, early Databricks.

    Justin Borgman17:50

    Yep.

    Matt Turck17:52

    Dremio.

    Justin Borgman17:54

    Dremio, and then Starburst.

    Matt Turck17:59

    And then Starburst. But then that group evolved to the lakehouse architecture.

    Justin Borgman18:08

    Yes, that's right. Or, yeah, I would argue we've been sort of focused on the lakehouse from the beginning, and Databricks has evolved into that.

    Matt Turck18:29

    That's right. Because you always had that vision that storing all your data in one place is a terrible idea. And you were always the federated query engine, and the core value proposition—my words, not yours—where it's like, just leave your data anywhere and we'll find it wherever it is for analytical purposes.

    Justin Borgman18:44

    Yeah, I think where our view on that has evolved slightly is we do think data lakes are where you're going to want to store as much data as you can, just because the economics will drive that. But you'll never store everything, to your point. And I used to use, actually, when I was raising venture for Starburst, I used to use your landscape all the time because especially when you show the progression from 2014 or whatever to today, 2025, you see the Cambrian explosion of data sources.

    Justin Borgman19:05

    And so being able to federate out to those has an advantage for certain use cases, for sure.

    Matt Turck19:09

    Yeah, it's become the claim to fame of that landscape, is to show the absurdity.

    Justin Borgman19:09

    That's right.

    Matt Turck19:30

    It's not exactly what it was meant to do, but that's what it's become. Okay. And is it fair to say, and I think you just alluded to it, that between those various buckets, the lines have started blurring as well? Because people have federation as well on top of being the repositories and different things. Is that fair?

    Justin Borgman19:54

    Lines have definitely blurred. Yeah, I think part of it is the industry's starting to mature. You have a couple giant players in Databricks and Snowflake that are now looking for adjacent markets to continue to grow their TAM and justify their valuations and drive their revenue into the future. And that's leading to more overlaps into boxes that they weren't typically part of.

    Matt Turck20:05

    And so how do you position today, like versus those and I guess Dremio and also the hyperscalers, BigQuery? How do you position?

    Justin Borgman20:29

    Yeah, I think it necessitates being even crisper on the differentiation between those things. And so for us, there are a few things that we point to with customers. Number one, we're one of the only hybrid players. So Databricks and Snowflake are cloud-only. So if you happen to have data on-prem, we're pretty much your only bet. And it just so happens that that turns out to be most of the Fortune 500.

    Matt Turck20:29

    Yeah.

    Justin Borgman20:55

    Almost the entirety of the financial services sector in particular. And so we do a lot of business in those industries as a result. The second thing that helps differentiate us is the openness of the platform. So we're an open engine querying open formats. And while there's been widespread embrace, I would say especially last year in 2024, around open formats and Iceberg in particular really winning that format, that's new. That's new for this industry. We've been doing this forever, though, and the first queries run on Iceberg were Trino or Presto queries.

    Matt Turck21:03

    Yeah.

    Justin Borgman21:33

    So that pairing in the open-source community of Trino and Iceberg has been kind of a reference architecture for years at this point. And that gives us an advantage because we can help manage your Iceberg deployment holistically, everything from streaming ingest, loading that data into Iceberg tables, maintaining it, doing things like compaction, data maintenance, data management, and then, of course, querying. We have that sort of end-to-end lifecycle around Iceberg, and we call that the Icehouse. So that's our play on lakehouse: the Icehouse.

    Justin Borgman21:56

    And then the other area where we differentiate, which you touched on already, was federation and being able to query into other data sources. And there are some use cases where being able to join a table that lives in another database system with a table that maybe sits in the lake can be very valuable in getting a fast response time. So we serve anti-money laundering use cases, fraud detection use cases, where very often the patterns to detect bad behavior actually exist in more than one data source.

    Why Starburst backed the lakehouse from the start

    22:06
    Matt Turck22:22

    So you were very early to that vision of the lakehouse, now Icehouse. Was it, I don't know, controversial at any point, the Iceberg plus Trino as a reference architecture, or was it hard to convince people that it was going to be the future?

    Justin Borgman22:41

    I would say yes. Until last year, there was a lot of debate over which format was going to win. Databricks had created one called Delta. There was another open format called Hudi, and then there was Iceberg. And if you're just evaluating those for the first time, you think, well, there are three formats. How do I know which one's going to win? I think the reason we had conviction that it was going to be Iceberg was simply that it was the one that had already been adopted by a lot of the super-scaled-up internet companies.

    Justin Borgman23:11

    And I think there's something to be said for saying that you can watch those companies and kind of see where technology is probably going to go because they're the ones running at the most ridiculous scale. These technologies get really tested to the limit that way and can be a good indication of sort of where things are going. So we saw them all adopting Iceberg along with Trino and felt like that was likely going to be a pattern that is adopted by the industry as a whole.

    Starburst Enterprise

    23:20
    Matt Turck23:41

    Okay, well, we'll go back to Iceberg in a second, but since we started talking about the reality of the Starburst product, you have three offerings, I guess, two core ones and one that's kind of newer. So you have the, as you were saying, on-prem version, which I think is Starburst Enterprise.

    Justin Borgman23:42

    That's right.

    Matt Turck23:45

    Then you have the cloud product, which is Starburst Galaxy.

    Justin Borgman23:45

    Yep.

    Matt Turck23:53

    So maybe walk us through those. And then, just to mention it upfront, you have a specific offering with Dell.

    Justin Borgman23:53

    That's right.

    Matt Turck24:09

    Which is the third product. Maybe walk us through those different products, and let's start with Enterprise, which is the on-prem version. Sure. How do you think of a classic question for open-source businesses?

    Justin Borgman24:38

    How do you think of that Enterprise offering versus the underlying Trino functionality, open source? Yeah, so I would call Enterprise a classic open-core model, which is to say that the core is open, meaning there's a Trino engine inside. We're the leading contributors to Trino. You get that as part of the offering. And then around it, we've built all the enterprise functionality that you would need. So fine-grained access controls, row-level, column-level data masking, query auditing. We've also built in a number of extra performance features.

    Justin Borgman24:51

    We have something called Warp Speed. We have a lot of fun with the names of our sort of sub-products. That is smart caching, smart indexing that delivers a 10x performance boost over—

    Matt Turck24:52

    Faster SQL?

    Justin Borgman25:17

    Faster SQL. Exactly. Some techniques there that give you faster SQL. That's exactly right. We have extra connectors, management capabilities, and so forth. And so that's what Starburst Enterprise is, because it is self-managed. What we mean by that is the customer is managing it, and that allows them the flexibility to deploy it anywhere. It could be deployed in an air-gapped facility. It could be deployed in a vehicle if it needed to be.

    Justin Borgman25:22

    Like, you could run it literally anywhere.

    Matt Turck25:24

    Yeah.

    Justin Borgman25:35

    And so again, that works for customers who have on-prem environments, hybrid environments, very complex environments, and gives them a lot of flexibility to integrate with the particular technologies that they have.

    Matt Turck26:03

    Yeah. And classic question again for that kind of business model is the underlying community, because you're the major Trino contributors, but there are contributors to Trino in lots of different places. And presumably, as customers build stuff, they contribute some of it to the community. So there's this natural tension in any open-source business between the open source that keeps getting better and the commercial products, and perhaps a narrowing gap. So what is your experience? Every company is going to be different, but what is your experience with that tension, I guess?

    Justin Borgman26:33

    Yeah, I think you're absolutely right to describe it as sort of nuanced, in the sense that you're always trying to make the right trade-offs between contributing to the community, which generally drives adoption, so that is a benefit, versus keeping something proprietary, which drives conversion. And that's the way I think about it. It's like adoption or conversion.

    Matt Turck26:34

    And you need both.

    Justin Borgman26:54

    You need to grow the pie, and then you need to make sure that you're getting a good piece of it. And so I think it's more art than science, but certainly things around performance and security are very logical places to kind of draw some lines where you know that enterprise customers value that, and that's something that can be monetized.

    Matt Turck27:07

    Do you have any kind of, in your license, anything that forces companies above a certain size to pay or not?

    Cloud vs. on-prem

    27:31
    Justin Borgman27:32

    We don't. And that's something that, at least philosophically, my co-founders have been pretty, I guess, consistent on, is a desire to continue to use the Apache License, which gives widespread flexibility and freedom to users of the technology. And so we haven't created any more, I guess, restrictive licenses like some open-source companies have.

    Matt Turck27:56

    And so you were saying the on-prem versions of Starburst Enterprise are particularly popular with large companies and financial services in particular. It's another super interesting topic of cloud versus not cloud, and who will resist ever going to the cloud. But your experience is that those companies are not going to the cloud?

    Justin Borgman28:20

    Yeah, I mean, they're pretty emphatic about it. I had dinner with the CEO of one of the largest banks in the country. I won't name his name just to protect his innocence, I guess. But he told me point-blank, like, "We're never gonna move everything to the cloud. We're gonna have four clouds," in his words, and the fourth one being his own on-prem, essentially. And so I do think that's the future. I think also there's an interesting case to be made that there may be some repatriation with the emergence of AI.

    Justin Borgman28:53

    And our partner Dell is very much betting on that as well. They're selling a ton of servers right now, AI servers, which basically have a lot of GPUs in them and a lot of NVIDIA products in there, and are seeing customers that are trying to get economies of scale by deploying that infrastructure on-prem and running it themselves. So yeah, I think this dichotomy is going to exist for as far as I can see.

    Starburst Galaxy

    29:10
    Matt Turck29:18

    Yeah, it's been fascinating, actually, to see the Dell stock price. Everybody's talked about NVIDIA for all the right reasons, but it's amazing how Dell, as the quintessential on-prem player, has seen their fortunes accelerate as well last year. Okay, great. So that's Starburst Enterprise and Starburst Galaxy, the managed version. How does that work?

    Justin Borgman29:45

    Yeah, so that's a classic SaaS product that's hosted and managed by us. It is connecting to your storage. So it's your own S3 buckets, your own RDS, your own MySQL, but the compute and the control plane is managed by us. And so we're able to offer a very seamless, easy-to-use, turnkey, push-button type of approach while still giving you all the performance and functionality that you need and that you're looking for. And that product has actually evolved very, very quickly for us.

    Justin Borgman30:14

    We've been able to get a lot of new, interesting features and functionality there. And one of the use cases where we've seen a lot of adoption and interest is actually customers who are building their own data applications and using Galaxy as the embedded engine, where there's a data analytics portion of the SaaS app that they provide to their customers.

    Matt Turck30:15

    Mm-hmm.

    Justin Borgman30:39

    And the analytics are essentially powered by our engine. And we were digging into why they're choosing us for this, and I think it goes down to, if you're going to be part of someone else's margin, essentially their COGS, right? You need good TCO. I think what we're seeing is data application developers are choosing Iceberg for storage because that gives them a lot of flexibility. They want something that's based on open source so that they always have that optionality as well.

    Justin Borgman30:56

    And then we can deliver them, pound for pound, the best cost performance, we believe, in the industry. And so those things all really matter when you're ultimately fitting into their margin.

    Matt Turck31:09

    Yeah, very interesting. And that's in contrast to this whole series of companies that specialize in MLOps, embedded analytics, and BI, so the Sisense and GoodData.

    Justin Borgman31:13

    Yes. Well, I'll clarify: we're not doing the visualization piece, right?

    Matt Turck31:17

    So some of those companies, but you help them do the analytics and then embed them.

    Justin Borgman31:22

    Exactly. Yeah, we're doing the query execution behind the scenes. We're the engine for that. Exactly.

    Dell Data Lakehouse

    31:23
    Matt Turck31:28

    Okay. All right. So that's Galaxy, and then we alluded to the Dell thing. That's more recent, right?

    Justin Borgman31:56

    Yes, that was almost 12 months ago that we released what's called the Dell Lakehouse. So it's their product powered by Starburst. And that's essentially Starburst inside with Dell's object storage, Dell's hardware. And it's become a real centerpiece of their AI strategy in selling this into large customers. Some of them themselves are CSPs that are building out their own infrastructure for analytics and AI. And I think at the end of the day, AI is only as good as the data that you're training the models on, only as good as the data that you're accessing through RAG workflows.

    Justin Borgman32:12

    And so having a lakehouse at the core of that architecture is important.

    Starburst’s data architecture explained

    32:13
    Matt Turck32:42

    Okay, great. So those are the three offerings. And still on the topic of product, maybe walk us through the overall architecture. You have a chart somewhere on the website that maybe we'll put in the video, where you have the classic sources on the left, and then Starburst in the middle and Magic on the right side. So walk us through that. So in terms of sources, how does the data go into, or how does Starburst access the data? What kind of data?

    Justin Borgman32:59

    Sure, yeah. So at the heart of Starburst is this notion of connectors, and we think of everything as a connector. So even if you're just accessing S3 and you're going to be querying Iceberg tables, that's technically our S3 connector, or data lake connector, to access that. And so every connector is basically just connecting to the underlying catalog of the system that you're connecting to, and then as soon as you've connected, which is like a one-time setup thing, now you can run queries. And you can do that at the command line, like just start to write SQL queries, joining tables across different systems.

    Justin Borgman33:46

    Or you can use a BI tool like Tableau or ThoughtSpot or others. And then this goes to the data application side: we're seeing customers build more programmatic ways to interact with the data that we have access to. One of the features that we've built that's really nice, especially for internal purposes, is something called data products, which is basically allowing you to stitch together a view of your data across these different data sources. And that's where you really start to get some interesting optionality because you can decide to materialize that view or not materialize that view.

    Justin Borgman34:19

    And there are trade-offs to both. If you're querying the data source directly, you're going to get very fresh access to the data directly where it lives, but you might trade some performance. And if you're really optimizing for performance to power some kind of dashboard, well, that might be a case where you want to materialize that view in the lake. And so we have a lot of flexibility in allowing you to do either one.

    Matt Turck34:22

    The data can be batch or real-time?

    Justin Borgman34:23

    Yes.

    Matt Turck34:29

    So meaning that you just have, what, a Kafka endpoint ingestor?

    Justin Borgman34:30

    Yes.

    Matt Turck34:34

    Is that a word? Ingestion of some sort?

    Justin Borgman34:58

    Yeah. And actually, that's something we've really optimized recently. It's what we call streaming ingest, and that's connecting to a Kafka stream and landing it in Iceberg tables automatically for you. And we're able to do that at incredible throughput. And that's an architecture I would say that we're seeing a lot of out there, is Kafka as opposed to more traditional batch-oriented ETL. And so you can land your data, have it there, and it might be 15 seconds old from when it was created.

    Justin Borgman35:10

    And that enables, again, more real-time analytics as a result.

    Matt Turck35:14

    The data itself can be structured versus unstructured. Does it matter?

    Justin Borgman35:40

    Yeah, typically it's structured or semi-structured for our SQL analytics offerings. But we're also doing some interesting things around vector search. We can connect to different vector databases. There's some things that we're working on there. But yes, today I would say primarily SQL analytics on structured or semi-structured data. Semi-structured would be something like JSON files or XML files. Yes.

    Matt Turck35:59

    All right. So that's sort of, for the most part, the left side of the Magic chart. So the middle: there's a SQL query engine, which is what we talked about, augmented by Warp Speed. That's what it's called. Okay. And then there's a governance component as well.

    Justin Borgman36:26

    Exactly, yes. And that's very important for our large enterprise customers. The access control piece in particular, having good auditing controls, being able to see who accessed what, who saw what, and being able to do both what's called RBAC, role-based access control, but also ABAC, attribute-based access control. And that's important because that allows you to tag certain datasets as being PII data. Maybe this one has certain data sovereignty constraints. Only people in Switzerland can see this data, for example.

    Justin Borgman36:33

    So it gives you a lot of flexibility.

    Matt Turck36:35

    Are there other bits in that middle that I forget?

    Justin Borgman36:37

    I—

    Matt Turck36:39

    I'm quizzing you on one chart on your website.

    Justin Borgman36:40

    Yeah.

    Matt Turck36:43

    Because as a CEO, of course, you know all the charts on your website.

    Justin Borgman37:03

    I've probably seen them all at some point. Yeah, I mean, I think the other is just obviously if you're using Galaxy, then we have a lot of automation of the cluster management itself. You can scale it up, scale it down, you can have it auto-scale on its own. And all of those things are just going to save you money if you don't have a cluster running all the time.

    Matt Turck37:22

    So I'm totally cheating because I'm bringing up the chart in front of me, but I'm not showing it to you. So you have, joke aside, you have ingestion, we talked about; you have table maintenance. So governance, we talked about; accelerated SQL analytics, we talked about. The two things we didn't talk about are table maintenance and automatic capacity management. Do you want to go into those?

    Justin Borgman37:46

    Okay, yeah. So the table maintenance piece, there's a lot of activities that you want to do to optimize for performance reasons how the tables are structured and laid out. If you've been around for a while, this is kind of like defragging from a long time ago, right? You want to take a lot of small files and bring them together into larger files, and that's called compaction. And that's a particular thing that you want to do with Iceberg tables to get better performance.

    Justin Borgman38:19

    And you could do that yourself. It can be tedious, time-consuming, error-prone. There's a lot of work involved, or you can use Galaxy and we do it all for you, basically. So that's sort of the table maintenance side of things. And then on the capacity management, that's really those auto-scaling capabilities where you can spin clusters up and down and have them automatically scale based on incoming compute. It's a more serverless type of experience behind the scenes.

    Matt Turck38:26

    The data apps you just described, is that a materialized view that you were describing?

    Justin Borgman38:28

    Well, that's the data products piece.

    The rise of data apps

    38:30
    Matt Turck38:34

    That's just data products. All right, great. So there's data products and there's data apps. So what are the data apps?

    Justin Borgman38:53

    So data apps is really what I would say is an emerging use case where customers are building their own applications, leveraging our platform, and doing that because, again, we become part of their COGS, part of their margin. So they want a platform that's going to be very cost-effective to them. And so we're powering the analytics behind some of the large SaaS vendors out there.

    Starburst AML

    38:54
    Matt Turck39:11

    As I'm reviewing my notes, I jotted down things like AML, anti-money laundering, customer 360. So those are examples. So those are apps that are delivered by SaaS companies, just to play it back, but powered by Starburst?

    Justin Borgman39:26

    Well, they can either be apps built internally. So, like an AML use case, an anti-money laundering use case, might be something that a bank has built themselves to meet their own regulatory requirements. AI, and they're a—

    Matt Turck39:32

    That's Amr's company, right?

    Justin Borgman39:33

    Yes.

    Matt Turck39:33

    Yes.

    Justin Borgman39:43

    And they're a cybersecurity-oriented company, and we're powering the analytics behind Vectra. So that's an example.

    Matt Turck40:07

    Yeah. Okay. But so, for the internal use case, like, I'm a bank and I want an AML, anti-money laundering. So walk me through the difference between a data app, like a Starburst data app, versus just a use case. So you formalize, you have, like, a building block, like a Lego block, that helps them build the AML app?

    Justin Borgman40:34

    So not necessarily for AML specifically. I would say this is a horizontal platform, but because of our REST API, they can start building their application directly against that. And so now access to all of the data that they need to either show or power dashboards or analytics within the app itself, or deliver reports, or whatever analysis is taking place, we can be the engine behind that. And we've tried to make it as easy as possible for those developers to build against that.

    “We actually built the Galaxy twice”

    40:41
    Matt Turck41:03

    Okay, very cool. So that's a lot of things you guys have built because you also have—we were talking about connectors—I wrote down, you have like 50-plus data sources. That's a lot of building. What turned out to be the hardest to build in the history of Starburst?

    Justin Borgman41:25

    The hardest to build. I would say building Galaxy, the SaaS platform. There's a story there where we actually built it twice. Customers never really saw the first version because at the last minute we said, "You know what? We can't deliver the seamless, easy-to-use, consistent experience that we're going for."

    Matt Turck41:27

    Oh, wow. So you had to go back to the drawing board.

    Justin Borgman41:52

    So we went back to the drawing board and went with a totally different approach. And actually, what's interesting is that first architecture was fairly similar to Databricks' architecture, where we did not control the compute plane. We were just this sort of control plane, and we were running everything in the customer's environment, and we saw potential benefits to that. But the biggest drawback is you don't have as much control over the experience. And we wanted to take a more—I'll call it—a more Apple approach to really delivering customer joy.

    Justin Borgman42:11

    And so we re-architected, and that basically took another year, so it was a big, expensive replatforming. But ultimately, we're really happy with the way Galaxy turned out.

    Matt Turck42:16

    Was it an obvious decision, or was it, like, a super nerve-wracking decision?

    Justin Borgman42:34

    It was a nerve-wracking decision. It was obvious that we had prototyped this different approach that became Galaxy as it is today, and it was obvious that that was going to be better. But still, it was a nerve-wracking approach because you're sort of throwing away, like, two years of development and starting over again.

    Matt Turck42:35

    Mm-hmm.

    Justin Borgman42:51

    And, of course, we're venture-backed, and we spent a lot of money to build that. And that was actually one of the reasons we raised venture in the first place. We have a little unusual history that we were bootstrapped the first two years, and we were running a nice little profitable business. It was great. But we thought, "You know what? We can't build a SaaS solution, we can't build a cloud platform off our little bootstrapped small business."

    Justin Borgman43:05

    It's just too expensive. And so we raised venture, built this cloud platform, and then threw it out and built it again. So that was definitely stressful.

    Matt Turck43:10

    Good times. I'm sure that must have made for good board meetings and good fun board meetings.

    Managing multiple products at scale

    43:13
    Justin Borgman43:13

    For sure. We're lucky we have very patient investors. Yes.

    Matt Turck43:35

    And the flip side to having a lot of different products, in addition to the effort and time to build them, is that it's a lot of surface to manage, a lot of product management that's involved. How does that work practically, having, I guess, now three different products that need to be sort of synced in terms of functionality?

    Justin Borgman43:42

    Yeah, definitely. There's complexity to it. I think the way that we try to simplify it is find as many common components between them. If you think of a car manufacturer, usually the big ones, the VWs of the world, they'll build on one platform, and that platform will then get used by five different car models, or maybe 10, where they've built the chassis once and they'll use the same engine in five different cars.

    Justin Borgman44:31

    So we've basically taken that kind of approach, where there are common elements that are used across all three of those. But even still, there are differences among those different products. And that requires a lot of intentionality, puts some added pressure on our product managers to really make sure that they have their ICP really nailed down, their ideal customer profile that that particular product caters to. And we touched on some of those things. If you're a big multinational bank with an on-prem footprint, Enterprise is probably your best bet.

    Justin Borgman44:44

    If you're more digital native, then you probably want to use Galaxy, and so each one has its own unique components.

    Matt Turck44:51

    And how does that work from a team perspective? Do you have different teams working on the different versions?

    Justin Borgman44:58

    We do. Yeah. Different engineering teams, different PMs. And again, there are common elements that are shared, but yep.

    Matt Turck45:05

    And then they'll all come to you, and you have to decide on the priority and who gets more money to go faster?

    Justin Borgman45:09

    Yes, we're actually doing that right now for our next fiscal year.

    “We founded the company on the idea of optionality”

    45:14
    Matt Turck45:45

    Yeah, very cool. And it's really one of the very interesting parts of this conversation, precisely that kind of on-prem versus cloud, because the dominant narrative in open source has been that people who did that kind of hybrid approach ended up not really liking it and then recommending to anyone that would listen that people should do cloud only. But you've embraced the complexity, and you do both.

    Justin Borgman46:01

    Yes, for better or worse. That's right. That's exactly right. I mean, we started the company with this idea of optionality. It was probably one of the words that I used, actually, at the data-driven event that you hosted, geez, I don't know, five or six years ago, right?

    Matt Turck46:03

    Yeah, 2019.

    Justin Borgman46:04

    Yes.

    Matt Turck46:07

    That's the second one we did, or the one we did for Starburst. Yes.

    Justin Borgman46:08

    Yes, yes.

    Matt Turck46:11

    You remember words you used in 2019?

    Justin Borgman46:40

    I do, because it was a word that I used a lot, and I still do. Optionality is certainly a word that you may hear in a finance context, but not necessarily used that much in technology architectures. And we sort of founded the company on this idea, which is to really give customers the power of optionality, the flexibility to build a future-proof architecture where you can change out components, you can work on-prem, you can connect to different data sources. And I think part of that was also our bootstrapped orientation, where we were very customer-obsessed because we depended on that revenue to, like, pay the next month's payroll.

    Justin Borgman47:02

    And so optionality became sort of a core ethos. Now, the consequence of that is complexity for us as a vendor. So you're right to touch on that, but it's one of the reasons customers choose us.

    Matt Turck47:17

    Okay, so we talked about the different offerings. We talked about sort of the architecture under the hood. Maybe to close on architecture: we did talk about Iceberg. You support the other two as well?

    Justin Borgman47:19

    We do. Optionality.

    Iceberg

    47:20
    Matt Turck47:23

    Optionality. There you go. So Hudi and Delta.

    Justin Borgman47:24

    Yep.

    Matt Turck47:28

    But you're seeing most of your customers use Iceberg?

    Justin Borgman47:38

    Yes, I feel like the summer of 2024, the world said, okay, it is Iceberg. And that was like a VHS-Betamax decision made.

    Matt Turck47:41

    What triggered it? Was that the acquisition of Tabular by Databricks?

    Justin Borgman48:00

    I think it was two things. I think that was a big, big piece of it. I think the other thing that happened just maybe two weeks before that was that Snowflake also decided to support Iceberg. And of course, maybe some of your viewers are aware there was a bidding war for Tabular, which is what led to such an incredible outcome for them.

    How open-source acquisitions work

    48:01
    Matt Turck48:14

    Yes. And to put this—so actually, I don't know what the numbers were, but I heard that Snowflake was bidding like $300 million and then $600 million. And then what was the final price for? It was like $2 billion, right?

    Justin Borgman48:14

    Yeah.

    Matt Turck48:21

    Databricks. So Databricks won over Snowflake with a $2 billion deal.

    Justin Borgman48:21

    Yep.

    Matt Turck48:36

    And now, which is fascinating, which is probably partly why Ali wants to—one of the reasons why Ali wants to keep the company private for a bit longer—is because he can do that kind of thing, which is a very bold move and exciting move.

    Justin Borgman48:37

    Yes.

    Matt Turck48:40

    But they can do that without the scrutiny of public markets.

    Justin Borgman48:42

    100%. Yes, I totally agree.

    Matt Turck49:04

    And what do you make of the acquisition from a strategic standpoint, while we're on the topic? Because Tabular, for anybody that may or may not follow the space very closely, was a young Series B company that was the commercial company on top of Iceberg. But ultimately, Iceberg is an open-source project.

    Justin Borgman49:06

    Well, Iceberg, yes, is.

    Matt Turck49:08

    What did I say? Sorry. Iceberg.

    Justin Borgman49:08

    Yes.

    Matt Turck49:41

    Iceberg is an open-source project, and it's sort of unclear what you buy if your ultimate strategic goal is to take control over an open-source project. It's sort of unclear what you actually buy by acquiring the company that's a commercial company on top of the open-source project. I mean, there are certainly a bunch of Iceberg contributors, like AWS and Starburst as well. Starburst and Microsoft, right? Snowflake.

    Justin Borgman49:41

    Yeah.

    Matt Turck49:51

    So what do you get then? You spend $2 billion, you buy a company, but to what extent, in your opinion, does that help you take control over that project?

    Justin Borgman50:20

    Yeah, I think this is a really interesting question. I don't think it does allow them to take over the project, to be perfectly honest. I think that the market is actually resolutely determined to ensure that it continues to be independent, which is important, actually, for Iceberg. And that's what made Iceberg popular in the first place over Delta, which was Databricks' own format. The market wants an independent standard. That's what they want, independent of any vendor. And so fortunately, there's enough groundswell of people like ourselves, like Snowflake, like some of the others you mentioned, where it is.

    Justin Borgman50:45

    And it's also, by the way, an Apache Software Foundation-governed project. So you have the Apache Software Foundation also ensuring independent governance, which is important and makes it truly, I think, independent. I mean, Ryan works for Databricks now, and he will for some period of time, but—

    Matt Turck50:47

    And Ryan being Ryan Blue, the CEO of Tabular.

    Justin Borgman51:13

    Exactly. But Iceberg is bigger than Databricks, honestly. And it will continue to be, I think, the ubiquitous format. So if they had hopes of killing it, I don't think that worked. If they had hopes of controlling it, I don't really think that will work either. I think probably the biggest thing that they get out of the acquisition is the ability to tell the market that they can do Iceberg too and not look like they had made a mistake with Delta.

    Justin Borgman51:38

    They get sort of a marketing win out of it. But as a practical, enduring matter, I don't think that they get extra influence beyond that. I also don't think—I think Ryan cares too much about the future of Iceberg. Even if Ali was like, "Do this thing," I don't know that Ryan would do it.

    Why Snowflake embraced Iceberg

    51:39
    Matt Turck51:54

    So yeah. And why do you think Snowflake embraced Iceberg before that? Because that's fundamentally against their interest, right? Snowflake is all about putting all your data in our proprietary database and leaving it here forever. So was it just pressure?

    Justin Borgman52:03

    Customer pressure. I think that's exactly right. I think that the market is basically saying we want to store our data in open formats. That's better for us. And by the way, this was sort of one of my theses, if you will, for starting Starburst. First, I believe gradually over time, enterprise software technology, at least infrastructure technology, wherever there are two things that are mostly the same and one is open and one is not, the open is going to win over time just because economics will eventually drive long-term decision-making.

    Justin Borgman52:36

    And that's what I saw with Cloudera. Again, Teradata was the better system at the time, but Cloudera was the open-source system. And that got a lot of adoption.

    Matt Turck52:46

    That's right. Because, so Teradata, as in Hadoop and Teradata, because Hadoop was basically a competitor to Impala, right?

    Justin Borgman52:47

    That's exactly right.

    Matt Turck52:50

    So we're going down memory lane, but that's right.

    Justin Borgman52:51

    Yes, that's right.

    Matt Turck52:57

    And Impala was open source and Hadoop slash Teradata was proprietary, right? So, lesson learned there.

    Justin Borgman53:06

    Exactly. I am the product of many lessons learned, and that's making me feel old, but that's exactly right.

    Data mesh

    53:15
    Matt Turck53:35

    Yeah, which is a good thing in this space of enterprise software. Having learned the lessons and tried different things is very much a strength. Okay, another thing that seems to be a part of the overall Starburst story and positioning is the data mesh. I would ask, what's up with the data mesh? That seems to be a very powerful idea. And we had Zhamak on either Data Driven or an earlier version of this podcast.

    Matt Turck53:54

    But that was all the rage sort of like a year or two ago. And I guess, what's the current state of it? Is that as vibrant and exciting as ever? What's your take?

    Justin Borgman54:07

    No, I think the phrase or the concept has lost a bit of momentum. I think that there are companies who are deploying data meshes, for sure, but some of the attention has certainly waned. I think the lasting legacy of that, though, is this concept of data products, creating these sort of curated datasets from data that can live in multiple places and thinking about them from a product perspective, with a product mindset, which is to say that there's a clear owner of the... And so I think, like any major sort of wave or phenomenon, there's a lasting impact of that.

    Justin Borgman55:03

    Yeah, I mean, Zhamak has her own company now. She's been very focused on—she probably hasn't been quite as publicly vocal on the data mesh concept. And I think what we've seen with customers is there's a people and process element that's actually harder than the technology part in implementing a data mesh. So we are very capable of helping customers implement data meshes from an architecture perspective, and we do do that. But I think it does require some retraining of sort of how you do things and how you organize your people internally.

    Matt Turck55:24

    Yeah, yeah, and I think we're saying the same thing, but I think the term itself is perhaps a little past the initial excitement, but the fundamental concept remains something powerful.

    Justin Borgman55:26

    Yes, yes, exactly.

    AI at Starburst

    55:31
    Matt Turck55:31

    Of course, the topic that we need to talk about is AI.

    Justin Borgman55:31

    Yes.

    Matt Turck55:35

    I'm hearing good things about AI. Apparently, it's going to be big.

    Justin Borgman55:35

    It's going to be big.

    Matt Turck55:54

    Yes. To what extent is that part of your positioning? We've talked a lot about SQL search and structured or semi-structured data. Equally, you were saying that for Dell, you were part of the AI story. So where do you fit in the sort of AI stack?

    Justin Borgman56:16

    Yeah. So, two places, I would say. First and foremost, on the training of models for those who are actually building their own models. Your models are only as good as the data that you train them on. Access to more data or better data is going to influence that. And so we become this sort of access layer to all the data in your organization. There is no piece of data that we cannot get to, essentially, through our federated architecture and being able to work across on-prem and cloud.

    Justin Borgman56:44

    And so that's one place. One element is you're going to access data, you're going to do some transformation of data, prepare data, train a model with that data. The other is the more practical aspects of putting AI into production, which we see RAG workflows as an essential part of. People are building these agents that need to access contextual information, pass that along to the LLM to get the appropriate response as part of the agent or part of the application that they're building.

    Justin Borgman57:15

    And we think we can play a central role in that, both in terms of the access to structured data, but also as performing vector search. And so we expect to be doing more, and you'll hear a lot more about us, I think, this year around RAG workflows and how Starburst plays a unique role in that.

    Key takeaways from go-to-market strategies

    57:16
    Matt Turck57:33

    Few words on go-to-market: lessons learned selling enterprise software. What has worked? What has not worked? How do you currently sell? Are you mostly outbound sales-driven, presumably given the type of customers? Yeah. Anything you can talk to?

    Justin Borgman57:56

    Sure. Yeah, we're mostly direct. And then, of course, we have an important partnership with Dell that we touched on as a channel for us. Things we've learned: I think the biggest thing that we learned was that we are—and this is going to sound silly now in retrospect, but it was an important lesson—we're hardcore enterprise software. We are not a PLG motion. And I say that, that feels silly in retrospect because it's so clear to us today.

    Justin Borgman58:05

    But there was a period where we thought we wanted to be PLG.

    Matt Turck58:05

    Yeah.

    Justin Borgman58:06

    You know, and—

    Matt Turck58:10

    Well, I think just about any company you name in 2021 thought they were PLG, right?

    Justin Borgman58:31

    Yeah, for sure, for sure. And so, we hired sales leaders who had PLG experience, and we thought that would make us PLG. But the reality is, PLG starts with the P, which is the product itself. And we're pretty hardcore, complex data infrastructure stuff. And our solution architects are some of the superheroes of our go-to-market. They're the guys and gals who know how to deploy this and work within the specific, unique data ecosystem that a particular customer has.

    Justin Borgman58:56

    And each one is different and unique. And that's part of what makes it hard to really create a true PLG motion around this, is that every customer is a little bit different in terms of what they're trying to do and how they want to deploy it.

    Matt Turck58:59

    So, do you charge services? Is services a component?

    Justin Borgman59:23

    We do, yes. And we do that both ourselves. We have our own professional services organization, and then we've trained up some amazing partners. We work with one that's pretty prevalent here in New York City called Kubrick, which does amazing stuff. And we work with a number of SIs outside of that too. But yeah, services are definitely helpful, especially for large, complex enterprises that are deploying this.

    Matt Turck59:40

    So, actually, a really interesting topic: services. For any founder listening to this, how did you go about building that network? So, Kubrick or others, when did you feel it was the right time to start working with SIs?

    Justin Borgman1:00:05

    That's a good question. I would say you should probably learn how to do it yourself first. So, I'd say not at the very, very beginning, but then my advice would be: try to start training up one or two, and probably boutique firms. Don't go after Accenture on day one, but start with a smaller boutique firm, because what's most important, especially in the early parts of developing the services piece, is the quality of the delivery. And you want to control the quality of that as much as possible.

    Justin Borgman1:00:32

    Obviously, you control it when you do it yourself, but that's difficult to scale. And it's not the margins that your investors probably care about. They don't want you to build a services company, so that's probably not going to be the dominant way you want to deliver ultimately. But by choosing just one or two and really focusing on them, there's a mutual benefit there. A, you're helping to bring them business, and you're creating that positive reinforcement cycle that learning how to deploy your software is going to lead to more revenue for them.

    Justin Borgman1:00:55

    But also, you're creating a reliable partner now that you can count on, that when you recommend that firm to deliver services at customer X, they're not going to make you look bad. And that's really important. So I would say, rather than going far and wide and signing up hundreds and thousands of SIs, which we did too along the way, I would say go deep with one or two, probably on the smaller end, just because you'll be able to get more attention, and try to make them as good at delivering as you are.

    Lessons from the Dell partnership

    1:01:18
    Justin Borgman1:01:18

    And then you can build on that relationship, and that can be very symbiotic.

    Matt Turck1:01:41

    And then the Dell partnership, which is more of an ISV kind of—the other sort of side of the partnership world. To the extent you can talk about it, how did that come about? Because a lot of people want to do that. A lot of startups want to do that, and that actually happens very rarely. So what was the history, and how long did it take, and any lessons learned there?

    Justin Borgman1:01:57

    They actually came to us, and I think that's a really important element. Hard to recreate, of course, for any entrepreneurs who are listening to this. But the reason that's important is whenever a partner is coming to you, you know that there's some motivation behind it, maybe that you don't even know initially, on why that's so important. And that's going to just drive momentum and focus and interest, because the biggest challenge working with somebody so much larger than you is, how do they get focused on it, right?

    Justin Borgman1:02:18

    I mean, they have tens of thousands of employees. How do we know this is going to be important to them? Well, because a VP within Dell was like, "This is important. We need to close this gap. Michael himself thinks we need to close this gap." Okay.

    Matt Turck1:02:20

    VP of product.

    Justin Borgman1:02:21

    A VP of product. That's right.

    Matt Turck1:02:23

    Yes. Not a VP of partnerships.

    Justin Borgman1:02:47

    That's actually a good point. Yeah. I think of the partnerships organization, at least in this context, as the facilitators of the relationship. But the real impetus, the real motivation, has to, I think, come from product. It could come from sales. That's also not bad. In some cases, it's a revenue leader who's saying, "Hey, we need a partnership with this company because we know it helps us get into these customers or close these deals." I think that's a great motivation.

    Justin Borgman1:03:11

    Innovation as well. But I think you do need that stakeholder who's willing to stand up and sort of sponsor the initiative to get the right resources around it. So in our case, it was on the product side. They had come to us, they looked at all the players in the market, they did some bake-offs, and decided, "Hey, we are the best technology for what they were trying to do." And that led to this OEM agreement. We started as a reseller, so we did about a year as a reseller agreement, meaning that they could resell our software.

    Justin Borgman1:03:30

    And then that evolved into a true OEM, where we're now embedded in their product offering. It is a Dell product for all their sellers. It's a first-party product. They're getting comped in full. It is treated as a Dell product in every way.

    Matt Turck1:03:38

    So how long was the whole cycle from first call to where you are now? Like a couple of years?

    Justin Borgman1:03:50

    Yeah, I would say probably, yeah, we're probably now starting year three, and I would say it was two years to get it really out the door as a SKU that they could sell. So it is also long. Yeah.

    Matt Turck1:04:01

    Yeah. So people should view those as very long-term projects. And who drove that on your end? Like, how did you make it successful on the Starburst end?

    Justin Borgman1:04:23

    So I think that idea of focus needs to be present on both sides to make it work. And so in my case, I have an SVP of alliances. His name is Tony, who sits on my executive team. He works directly for me, reports to me, and shepherded that deal on our end. And that meant that it was always visible to me. It was something we talked about in every one of our one-on-ones. How's it going?

    Justin Borgman1:04:33

    Where are we? What are the negotiation points? What are we working through? And so I think it really needs to be that kind of priority on both sides to make it work.

    Predictions for 2025

    1:04:40
    Matt Turck1:04:53

    Okay, maybe to close, what's the next year going to look like for you? Or I guess we're at the beginning of 2025, so what's next year going to be for you and then for the industry? So that's the prediction, 2025 prediction part of the question.

    Justin Borgman1:05:11

    The thing I'm most excited about is really seeing AI start to get implemented in real production use cases. I think the past year or year and a half has been a lot of experimentation, a lot of prototyping, and I think we're now just starting to see AI actually matriculate into real production use cases. And that's going to get us over the hype cycle because I think even in my own organization, when we started to get into AI, some people were like, "Is this just a fad?"

    Justin Borgman1:05:49

    Is this like crypto? Yeah, yeah, exactly. No, we're going to quickly get into, I think, a real plateau of productivity, or whatever they call it in the hype curve here, where we're going to actually see material results, which is only going to reinforce, I think, the demand for it. And I think RAG in particular will be a really important element of making AI production-ready.

    Matt Turck1:05:55

    Okay. So that's both what you have on the docket at Starburst and your prediction for the industry?

    Justin Borgman1:05:56

    Yes, exactly.

    Matt Turck1:06:02

    Very cool. All right. Well, Justin, thank you so much. That's been wonderful. Really interesting chat. Appreciate it.

    Justin Borgman1:06:03

    Thank you for having me. Always a pleasure.

    Matt Turck1:06:24

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.