MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Can AI Infrastructure Work Like Magic? | Erik Bernhardsson, CEO, Modal

    Erik Bernhardsson is the CEO at Modal. We cover why serverless fits bursty AI workloads better than traditional web services, why Modal replaced Kubernetes and Docker to start containers in seconds, and why inference spending will eventually exceed training spending.

    10/31/2024

    Hosted by Matt Turck · with Erik Bernhardsson, CEO, Modal

    AI infrastructureserverlessGPU computemodel inferencedeveloper tools
    Listen now
    YouTubeApple PodcastsSpotify
    56 min · 13 chapters
    Contents

    Transcript

    What is Modal?

    1:35
    Matt Turck0:58

    Hey, Erik, welcome.

    Erik Bernhardsson1:01

    Thank you. It's exciting to be here.

    Matt Turck1:30

    Yeah, I'm excited to do this. As I was prepping for this, I was on your website looking at the list of angel investors that you've had. And I think we've had just about everyone, at least on the angel side, on your cap table. So we had Barr from Monte Carlo on the pod. We had Barry from Hex just recently, Jordan from MotherDuck, Tristan from dbt. I'm probably forgetting one or two, but it's funny, like the whole group. So it was long overdue that—yeah, it's like the data influencer.

    Matt Turck1:51

    That's the thing: data influencer. Okay, so you are the CEO of Modal Labs, building a fantastic infrastructure company right here in New York. Let's start with a quick elevator pitch, like the 30- or 45-second version of what Modal Labs does.

    Erik Bernhardsson2:13

    Yeah, so Modal makes it easy to build, scale, and deploy applications in the data, AI, machine learning realm. So basically, you can think of us like: you write a little bit of Python code, and we take that code, we stick it in a container, we execute that in the cloud in a way where you don't have to think about infrastructure. You don't have to think about GPUs, you don't have to think about the cloud vendors. We just magically make that work and scale it to very large—thousands of GPUs, potentially.

    Current state of AI compute space

    2:18
    Matt Turck2:57

    I thought a fun way to start the conversation, to make it interesting to a broad group of people, would be to talk about the AI compute space in broad strokes. So people have heard of GPUs, obviously. People have heard of GPU clouds. Maybe people have heard of inference clouds and platforms, and they've heard of NVIDIA, Modal Labs, Groq, Lambda, CoreWeave, all those companies. If you had to categorize the space, like who does what, the different approaches, what would you say? How would you explain it to somebody that wants to learn about the space?

    Erik Bernhardsson3:13

    Yeah. And by the way, I think it's funny because it's like the typical VC, like, what's the market map? And my role is like, there is no—we defy it. We do all of it.

    Matt Turck3:29

    Yes. And I'll say, as a claim to fame, in your blog post that introduced Modal Labs, whenever that was, three years ago, you hyperlinked my market map as a way of showing the nonsense in the space.

    Erik Bernhardsson3:36

    No, I don't think it was the nonsense. I think I was just pointing out how much stuff there is. I love your market map, by the way. There used to be like—

    Matt Turck3:43

    I'm going to take this and I'm going to clip it and have it in the short. Erik loves my market map. I love the market map.

    Erik Bernhardsson3:56

    People, for the record, we will get back to that later. I built this thing called Luigi many years ago, and it used to be kind of—you could see it, but every year it's getting smaller and smaller. I don't know if it's still, but every time it comes out, I zoom in and I try to find it.

    Matt Turck3:57

    Luigi is still on there.

    Erik Bernhardsson4:21

    It's still in there. Nice. I love it even more now. Anyway, going back to what does the market map look like? There's obviously the clouds, right? Like AWS, GCP, Azure, et cetera. Oracle. I should shout out to Oracle. I actually like them a lot. We use them. Big fan. So they, I think, have been there for a long time and have been doing a phenomenal job building primitives for running stuff in the cloud. More recently, there's been a bunch of sort of newfangled, I call them like the alt clouds.

    Erik Bernhardsson4:49

    CoreWeave, Lambda, what are the other ones? There's like FluidStack, there's Vultr, there's a lot of them, right? As it turns out with GPUs, what people really care about is like, just want a box with the GPU in many cases. And so that's something that's like a little bit more of a commodity, and I think a little bit easier to build in, arguably. They would probably disagree with this. So I think we've seen a sort of renaissance, or not a renaissance, but like a bunch of new sort of cloud entrants, like sort of competing with the hyperscalers, focusing more on just like only GPU availability.

    Matt Turck5:13

    And so those, I'm a developer, if I want access to a GPU, instead of buying my own GPU, I go to those clouds and there will not be much else, but there will be what I need.

    Erik Bernhardsson5:20

    It's like they provision a box with a GPU and you can SSH into it, right? So it's sort of like you get a GPU on your desk.

    Matt Turck5:28

    So those people are in the sort of data center business, meaning that they will have the GPU farms and the electricity.

    Erik Bernhardsson5:54

    Yeah, it's sort of almost like real estate, right? They own sort of real estate and they rent it out, arguably sort of WeWork, you could say, if you want to be a little cynical, but not a ton of value higher up in the stack, which is, I think, where we play, is slightly higher up in the stack. I think if you go up in the stack, start looking at more like applications or infrastructure as a service, there's like a few different providers. First of all, there's a whole bucket of all the LLM inference providers.

    Erik Bernhardsson6:21

    So there's OpenAI, Anthropic, then there's a few ones focusing on open-source models. So Fireworks, Together, I think Anyscale to some extent, Replicate, et cetera. So those I would put in the LLM bucket. And then there's like a slightly adjacent bucket, I would say, of like more like AI as an API. And actually, that's probably more like where I would put Replicate. I also put fal in there. I think we occupy a slightly different bucket, which is more like you're not using a model through an API, you write code.

    Erik Bernhardsson6:52

    We're all about running, taking users' arbitrary code and running it in the cloud, right? Like, we're a serverless compute platform. So we take your code, we execute it in the cloud. You can run whatever code you want. In many cases, that involves GPU, in many cases not. In that space, I think there's less of an obvious—I would say Base10 maybe to some extent is somewhat similar to us. They focus only on model inference. We do model inference.

    Erik Bernhardsson7:12

    Model inference is probably the majority of our use case, but we also have many other use cases. We have people using us for batch, we have people using us for large-scale video encodings or building web services or all kinds of other stuff. And then I don't know, what other boxes should I talk about that we didn't cover?

    Matt Turck7:27

    I think that's pretty much it. At the very beginning, there would be just like the pure hardware, which people can still use direct, obviously, without any of those cloud solutions. But that's a whole different level of effort.

    Erik Bernhardsson7:34

    Yeah. And you mentioned Groq. It's not like an area we're playing in, but I do find it super fascinating.

    Matt Turck7:35

    Right.

    Erik Bernhardsson7:58

    Traditionally, all these LLM inference providers have obviously been building on H100s or NVIDIA GPUs, but I do think it's kind of fascinating how it's kind of going in the same direction as crypto mining towards more specialized hardware. Cerebras is another one. I think it's sort of interesting to see for LLMs. I think the reason why it's happening for LLMs is it's like a kind of well-defined API. It's like text in, text out, and it works really well. And it's sort of a few big models dominate everything.

    Erik Bernhardsson8:20

    The spaces where we tend to compete and where we have most of our customers is more like very custom models, like people training their own proprietary video models, image models, like audio, things like that, or having very custom workflows. And there's less of sort of like power law there. There's thousands of models. Everyone has their own slightly different models. And that's where we focus, is like, if you have your own custom model, you're probably really good at training models and building applications, but we handle the infrastructure.

    Erik Bernhardsson8:35

    So all the auto-scaling and all the making sure utilization is maximized and all these things.

    Matt Turck9:00

    Yeah. And by the way, at a very high level, all of this revolves around the fundamental concept that people are going to want to build their own models, which if you listen to some people, would say, well, it's all going to be provided for you as an API and you're never going to have to worry about this. But you obviously believe in a world where that is not true and people are going to want to build their own AI.

    Erik Bernhardsson9:16

    I mean, I don't know if I necessarily believe in it, but I certainly bet on it. I think there's some sort of existential risk that there's such a good multimodal model that people just use it for everything, right? It's like a model to rule it all. I don't know if that's true. I tend to think it's probably less likely to happen than other people just because having seen so many custom, all the idiosyncrasies around building custom models, let's say you want to build a voice-cloning thing, right?

    Erik's path to starting Modal

    9:54
    Erik Bernhardsson9:54

    Probably, like, a multimodal model is not going to care about all the specific things around that, right? I don't know. I'm a big believer that, in the long run, people will train a lot of different custom models, even though the multimodal models may handle even more stuff. I mean, the pie is growing. I guess it's like a general statement of software engineering in the last 30 years, but the pie always gets bigger.

    Matt Turck10:04

    You mentioned Luigi a minute ago. What was your journey into Modal? What did you do before, and what was the path to starting the company?

    Erik Bernhardsson10:26

    Yeah, so I've done a bunch of different things, but in particular, I was at Spotify for seven years, and I built a thing called Luigi, which was a workflow scheduler. So this is sort of a problem I tried to solve at Spotify: we had a lot of different batch jobs, and you ended up with a very complex graph. You need to run hundreds of different jobs in a certain sequence in sort of very parametric ways. And so I built Luigi to sort of help us figure that out, like how to execute all of it in the correct order. Open-sourced it.

    Erik Bernhardsson10:49

    A bunch of people started using it. I think this was like 2011. Then I stopped really caring about it in 2015. Airflow kind of took over, and then later now there's Dagster and Prefect and Flyte and a bunch of other ones. I mean, I still think it's a good idea, and kind of related to how I ended up starting Modal. Part of how Modal came to be is I actually started looking again at workflow scheduling in late 2020, when I realized I wanted to start something or build something, and started thinking a lot about what makes a good workflow scheduler.

    Erik Bernhardsson11:17

    I even built one. I have some janky code on my laptop. I realized at some point a workflow scheduler is only as good as the underlying compute layer, because at the end of the day, you can't just describe the work you need to do. You also need to describe the code. You also need to write the code. And if that integration point is bad, if you have kind of a system where you have to define the workflow over here and the application code over here, you're never going to have a great experience.

    Erik Bernhardsson11:53

    So I realized, actually, I'm just going to stop working on the workflow scheduling. I'm going to focus all on the compute provider instead. And what makes a good compute provider? I realized, what I want is I just want something that lets me define application code in Python, because I was very focused on data, AI, machine learning applications, and let this infrastructure handle all the rest, like all the scaling, all the provisioning, all the resource management, all that stuff. And I just got obsessed and carried away.

    Erik Bernhardsson12:02

    And three, four years later now, I have a company around it. But that's sort of the genesis of Modal. So it kind of started with Luigi in a way.

    Matt Turck12:13

    Yeah, because originally the plan was to create tools for data engineers and data scientists and the whole AI and GPU... You evolve with the market, right?

    Erik Bernhardsson12:14

    Is that what I told you a long time ago?

    Matt Turck12:16

    No, I don't know, but that sounds good.

    Erik Bernhardsson12:20

    I feel like we had a conversation many years ago, and I probably said something like that.

    Matt Turck12:27

    You tell me. But yeah, I do seem to remember that at the beginning. And when did you start the company? In 2020?

    Erik Bernhardsson12:29

    Yeah, early 2021, really.

    Matt Turck12:38

    That was the time when what was top of mind, I think, for a lot of potential customers was more like data infra, and it was pre—

    Erik Bernhardsson12:38

    Pre-GenAI.

    Matt Turck12:39

    Yeah, pre-GenAI.

    Erik Bernhardsson12:55

    So it's a little bit different. And yeah, I think that's always been the ethos of the company. At the end of the day, I just want to build better tools for data, AI, machine learning teams, and what we actually build, I'm actually somewhat agnostic. At the end of the day, I want to be a platform that helps people build those types of applications.

    Matt Turck13:02

    In your mind, a lot of this is interchangeable: data, machine learning, to some extent.

    Erik Bernhardsson13:21

    I mean, a little bit. I don't know, AI has its ups and downs. I think we had, whatever, four AI winters or whatever. Now AI is like a great term. I feel like 2020 AI was like a bad term. I used to say only, like, data. I'm just going to focus on data. There's been so many other terms coming and going. To me, it's all the same. It's like data, right? You're working with data, and sometimes you're training models, sometimes you're doing, I don't know, data pipelines and batch compute.

    Erik Bernhardsson13:31

    But at the end of the day, it's data.

    Matt Turck13:55

    And we were talking about this market map. I do the MAD Landscape every year, and that's a question I get every year because I steadfastly refuse to break it apart into data infrastructure versus machine learning versus AI kind of thing. But that's exactly the point. For me, all of this is completely tied together. But as a result, that's why it ends up being so populated.

    Erik Bernhardsson13:55

    Totally.

    Core elements of the Modal platform

    13:57
    Matt Turck14:05

    I'd love to get a little bit of a tour of the different pillars of the product and what it does currently.

    Erik Bernhardsson14:16

    So I think the core of it is really just compute and running containers. That's where we started. We always had a vision of building a platform, but the core of it is: we want to make it easy to run code in the cloud. And it turns out that inference is such a killer use case for that, that we effectively spent the last three, four years now entirely focused on just how do we take user code and execute it in the cloud in a way where we handle the scaling and all the infrastructure.

    Erik Bernhardsson14:48

    We sort of let everything else wait until we really nail that experience. How that breaks down is under the hood, and we realized in order to deliver that experience, we had to kind of go deep in the layer of the infrastructure and throw out a lot of existing stuff. In particular, one of the things I always cared about is developer experience. I wanted to build a tool I always wanted to have through my own experience building these types of things.

    Erik Bernhardsson15:15

    I didn't mention it, but at Spotify I built a music recommendation system. So a big part of my life used to be iterating and shipping and deploying things, or just running various types of batch jobs and debugging things. And at the core of that, one thing I realized, for instance, was if you don't have a good, super-fast feedback loop, you can never make it. So much of developer experience comes from having a super-fast feedback loop where you can take code and execute it in the cloud in a way where it almost feels like it's local.

    Erik Bernhardsson15:39

    And if you solve that problem, you solve a lot of different problems. You make engineers more productive because they can iterate and run things quickly. You basically get rid of what I think was the gap between local development and cloud development, which I think with AI and machine learning has always been a big challenge, because you end up having different environments, different toolchains, and then you end up having to take research code and repackage it and rerun it in order to get it out there in the cloud.

    Erik Bernhardsson16:10

    So if you get rid of that distinction between local and cloud and just turn it into one environment, you always run things in the cloud, you solve a lot of different problems. But in order to do that, you have to make it fast. You have to make it feel like you're developing locally when you're running things in the cloud. So what we realized, in order to do that, is you have to solve a number of core technical challenges around container cold starting.

    Erik Bernhardsson16:30

    So what I mean is you have to take code and you have to ship it to the cloud and you have to start that container. And ideally you want to do that in a couple of seconds at most, because that's sort of a human reaction speed. It kind of feels snappy. And then we started looking at, okay, well, what actually happens typically in a system when you do that is you deploy something to Kubernetes, and then you have to pull down a Docker container image, and that then has to start up.

    Erik Bernhardsson17:05

    And all of those steps are super inefficient. So we built our own. We threw Kubernetes out the window, we threw Docker out the window, we built our own file system in order to optimize for how container images are distributed. We built our own scheduler in order to maintain this pool of workers and make it possible to start tasks very quickly. We built our own container image builder because we had our own container image format.

    Matt Turck17:08

    That's a lot of effort. How long did that take?

    Erik Bernhardsson17:23

    I mean, it was too long. And this was during the ZIRP era. So I feel like VCs were like, yeah, cool. People were asking, why do you really have to go so deep? But fundamentally, I think we felt like we're going to solve this problem well. So yeah, we spent basically a year and a half, two years, two and a half years, almost like in a cave, just building a product that no one had used, which is maybe a bad idea.

    Erik Bernhardsson17:54

    And then when we actually kind of emerged out of that cave, it turned out we didn't really have an obvious use case. So we struggled for six months or a year or so to try to find a use case. And then luckily all this Stable Diffusion came out in summer of 2022 or something like that, and a lot of people started coming to us and were like, actually, this kind of makes sense. You have this serverless GPU thing. You can sort of realize this makes a lot of sense for us to focus on.

    Erik Bernhardsson18:04

    And then we started seeing more meaningful traction with users.

    Matt Turck18:11

    The GPUs that I have access to through Modal, where are they? Where do you get the GPUs from?

    Erik Bernhardsson18:28

    They're all over. And this is not a secret, by the way. We use AWS, we use GCP, we use Oracle. We have a number of alternative cloud providers now that we're also using. And we run a lot of different regions. We have a bunch of GPUs in the US, we have a bunch in Europe.

    Matt Turck18:34

    And is part of the trick to be intelligent about which one you get where, when?

    Erik Bernhardsson19:01

    Yeah, totally. And for a few different reasons. One is just capacity management. From time to time, capacity varies, right? There's this sort of cyclicality. If you look at Europe versus the US, it's actually kind of interesting. You see that when it's early morning, the US is not high utilization, but Europe has high utilization, and vice versa. So you can play all these games kind of just aggregating capacity across different regions. Pricing is the other obvious dimension, right?

    Erik Bernhardsson19:30

    We actually have a system that continuously looks at cloud pricing and solves an optimization problem to sort of figure out how do we allocate capacity across 100 different regions in order to deliver the capacity we need to our customers. Because we use a combination of reservations, but also a lot of spot and on-demand capacity, which means we rely on the cloud's ability to scale up and down, because it's a sort of supply-demand matching problem, right? We have volatile demand coming from our customers, which means ideally we also have the ability to sort of match that with on-demand capacity on the other side.

    Erik Bernhardsson20:04

    Which is also why I think we found a good product-market fit with inference, because inference is inherently sort of unpredictable, right? You don't know necessarily when you're deploying something what the usage is going to be. There's going to be a lot of daily variations, going to be a lot of weird spikes when something goes viral, et cetera. And a big part of the pitch of Modal is you can use us, we'll just scale automatically. You don't have to worry about making a cloud commitment, especially not a three-year or one-year reservation.

    Erik Bernhardsson20:19

    You can just deploy things on Modal, and then if you one day need 1,000 GPUs, we can typically get you 1,000 GPUs pretty quickly, like, talking minutes.

    Matt Turck20:29

    In terms of the model itself, I just bring my own model effectively. I go on Hugging Face and grab some open-source model and deploy it on Modal.

    Erik Bernhardsson20:46

    Yeah, you can definitely take a Hugging Face model and deploy it. That's what a lot of people use Modal for. I think where Modal really shines is also people training their own models, like having custom models. One example I always bring up as an amazing use case—I love the product—is Suno, which is AI-generated music. And they have a big cluster; they train their own models outside of Modal, and then they use Modal for the inference side, which means all this sort of generation of AI-generated music happens on Modal at very large scale, right?

    Erik Bernhardsson21:02

    Millions and millions of pieces of music generated on Modal.

    Matt Turck21:05

    Sort of full circle with the Spotify experience.

    Erik Bernhardsson21:28

    Yeah, it's actually funny because sometimes when I talk to them, I end up talking a lot about licensing and stuff like that, music licensing, because I happen to know a lot about that. But yeah, so what I love about that is they can do what they're good at, which is to build an amazing user experience and a user application and train models. And for all this sort of infrastructure mess, they can effectively outsource that to Modal. And that happens to be something we love and we can do really well.

    Erik Bernhardsson21:36

    So I think that's sort of the comparative advantage of specialization in that sense.

    Matt Turck21:39

    The training side would not make sense.

    Erik Bernhardsson22:06

    We are interested in training. We're looking at it. I think training at very large scale, people are very price-sensitive; it's somewhat transactional. I think you have very different sort of requirements, typically, like InfiniBand and very high interconnect. It puts a lot of different constraints on the infrastructure that we don't support today. It's something we're interested in, somewhat looking at down the road, but it's not a super high priority for the business right now. I think another use case that we do support is actually a lot of preprocessing for training.

    Erik Bernhardsson22:36

    So a lot of users use this for, let's say you have millions and millions of audio files or video files and you want to extract some training data from that. Typically, that requires a sort of batch job that extracts features from it. And that's something that Modal does really well too. It's like very high, bursty batch parallelism, just like fan out, parallelize, run this over millions and millions of audio files or something like that.

    Matt Turck22:42

    And then you're done, right? So it's sort of the ability to be elastic and just grow very quickly.

    Erik Bernhardsson22:50

    Totally, exactly. It's sort of pay-as-you-go because everything in Modal is usage-based. So you just pay for the exact time your GPUs or CPUs run.

    Matt Turck22:54

    Fine-tuning, is that emerging as an important—

    Erik Bernhardsson23:17

    I think fine-tuning is like—a year ago, I would have said fine-tuning is absolutely critical. I think now the models have gotten so good that this is sort of a question of: is fine-tuning even needed? I think fine-tuning has a lot of different use cases. I think in particular, if you have very domain-specific models, let's say you're focused on financial data or document extraction or something like that, I think fine-tuning absolutely makes a ton of sense. And so we have a lot of users using Modal for that.

    Erik Bernhardsson23:28

    I think beyond that, I don't know, fine-tuning is still, to me, a little bit of an unproven thing, but clearly there are certain places where it does make sense.

    Matt Turck23:51

    It's a fascinating comment, actually, in terms of the current state of the AI market and where it's going, right? It sort of feels like we're just coming out of a big phase of training and fine-tuning to some extent, which I guess is a part of training, and going into the world of, okay, we built those things, now let's use them, which is the inference.

    Erik Bernhardsson24:16

    Exactly. Yeah. And also, a couple of years ago, a lot of companies tried to train their own LLMs. To me, it makes almost no sense, right? Like MosaicML and stuff like that. So, yeah, I agree with you, right? It's obvious that in the long run, inference spend will dominate just for the obvious reason that that's where you make the money, right? When you're training a big model, you're spending a lot of money training that model. How do you recoup that?

    Erik Bernhardsson24:35

    It's typically through inference, right? If you look at OpenAI, any of these businesses. So to me, it's very clear that maybe in the past we've had like 80/20 spend, GPUs training versus inference. I think in the future it might be the other way around. Much more dollars will be spent on inference than training.

    Matt Turck24:48

    You mentioned batch processing, job queues, and all that stuff, which in some way feels more like the sort of more classic data engineering part. Is that a big use case as well?

    Erik Bernhardsson25:02

    Yeah, there's a bunch of people using us for data parallelism and pipelines and stuff. There's weird sort of unexpected use cases. I shouldn't say weird—it sounds like I'm dismissing that—but actually fascinating use cases. We've seen in biotech, for instance, companies having—I don't know bio super well, so this might be a gross, generalized, bad sort of characterization—but scanning millions and millions of compounds for certain chemical properties, or medical imaging where customers are running computer vision on millions of medical images.

    Erik Bernhardsson25:46

    So there is a lot of sort of batch processing also in those use cases I find fascinating. Not just sort of data pipelines, which is what I think people think about with batch processing, but also all kinds of other interesting applications around biotech or video transcoding or feature extraction and things like that.

    Matt Turck25:51

    And to finish the tour, sandboxed code execution. What does that mean?

    Erik Bernhardsson26:19

    Yeah, we started seeing a lot of customers who build LLMs needing, especially people doing LLMs for code, needing safe code execution. As it happens, we had great primitives for that because that's effectively what we built, right? We built a system for containing user code and executing it in a safe way. So we started thinking about, can we expose that as a service? So we also have primitives to take sort of untrusted code and execute that in a safe way. And I would put that in the sort of category of unproven.

    Erik Bernhardsson26:46

    It's still sort of early days. We're trying to figure out exactly what that market opportunity looks like. But there's so many startups building LLMs for code, whether it's code generation or migrations or things like that, or debugging or things like that. So to me, that remains an unproven but potentially very large upside market opportunity for us that I'm pretty excited about.

    Matt Turck26:55

    Why did you pick serverless as the right level of abstraction for this space in terms of what developers need?

    Erik Bernhardsson27:20

    I think it's funny because I feel like serverless was all that hype. People were really excited about it when Lambda came, and then it didn't really catch on. And 10 years ago, when Lambda came—I actually don't remember what year Lambda came, but I think it's about 10 years ago—I probably would have expected in the future 90% of web development will obviously just be—people don't want to think about containers and processes and things like that. And that didn't really happen.

    Erik Bernhardsson27:45

    And I've thought a lot about why. And one of the reasons I've kind of realized maybe the reason it didn't happen is actually the developer ergonomics around building backend services is actually kind of fine. If you're, like, a backend developer just deploying with Kubernetes or running things locally, it's actually not so bad. It turns out moving things to Lambda is actually kind of worse in many ways. If you look at data teams, though, like data, AI, machine learning teams, I think we never figured out developer ergonomics for those companies, for those teams.

    Erik Bernhardsson28:14

    And 10 years ago, I was the only data engineer, the only data scientist at Spotify. It was kind of fine because it was such a small fringe part of software engineering that probably didn't really matter. But now there's a lot of people building data, AI, machine learning applications. And I think developer ergonomics and developer experience matter so much more. And as it turns out, actually, serverless kind of makes sense, I think, for a lot of the applications that those people are building.

    Erik Bernhardsson28:47

    So serverless always felt to me like, in hindsight, it looks like the right idea but applied to the wrong problems. And I always wondered. Now, looking at what we built at Modal, I'm kind of convinced that serverless actually makes a lot more sense for data, AI, machine learning applications for a lot of different reasons. You have very bursty things, you have a lot of different containers, you don't necessarily care about hyper-low latency. I think there's a lot of different reasons why it actually makes a lot of sense for those types of use cases.

    Is serverless the right level of abstraction for AI compute?

    28:52
    Matt Turck29:08

    As any great entrepreneur, you've been very focused, and you seem to have said no to self-hosting scenarios. First of all, is that correct? And second, just walk us through the thinking.

    Erik Bernhardsson29:12

    Yeah. And it's something we're still debating to a large extent.

    Matt Turck29:12

    Right.

    Erik Bernhardsson29:30

    And by the way, I think of self-hosting not just as a binary thing. It's kind of a spectrum, right? There's on-prem on one side, there's fully multi-tenant on the other side, and there's a bunch of stuff in the middle. There's VPC peering, PrivateLink, cloud-prem, BYOC, bring your own cloud. So it's actually more of a spectrum, but we've been cloud maximalists, and we've been all in on, like, let's just build a multi-tenant service where we just run everyone's code in our cloud environment.

    Erik Bernhardsson30:05

    And to be clear, a lot of potential customers are just not comfortable with that model. I tend to think about that as like the cloud itself. I'm old enough to have started my career pre-cloud. And I remember the first time I heard about AWS in 2007, 2008, or something, when they launched, my first reaction was like, how could anyone ever run their code on someone else's computer? That's nuts, right? And then just a few years later, I was doing it myself, and I was like, this is awesome.

    Erik Bernhardsson30:28

    And then I think there's sort of a similar thinking. I met with the Snowflake team when they were still building the product super early, and I was like, why would anyone put their own data in Snowflake's data warehouse running in the cloud? But now they have very large enterprise customers that are very comfortable with that. And so I think the general wind is blowing in the direction of multi-tenancy and security moving away from the network layer into the app layer and a bunch of different other things.

    Erik Bernhardsson30:58

    And I think one trap that a lot of startups fall into is that you go out and talk to a lot of customers and they sort of insist on self-hosting. In my opinion, sometimes that's actually true, hard concerns. But in many cases, it's actually more of a soft concern, and it really just means you haven't built a product that they're desperate enough to use. And so maybe one day we will open up and offer other deployment models, but I'm actually more interested in the other option.

    Erik Bernhardsson31:30

    Like, how come we haven't built a product that they're desperate enough to use us? Right. Can we push more in that angle? Can we make it so good that they're willing to waive the typical concerns around deployment model and embrace this fully multi-tenant model where we run everything in the cloud? Because the benefits are tremendous. I really think resource pooling is one of the free lunches that exists in cloud economics. You can get much faster capacity, much more liquid capacity, et cetera.

    Matt Turck31:36

    Just double-click on that. What does that mean, resource pooling?

    Erik Bernhardsson31:53

    If you have a lot of different people and they all have very volatile demand, if you just put those pools together, you can actually scale it much more efficiently and drive much higher utilization than if you separate it and say each one of you has to run your own pool of compute. Right.

    Matt Turck32:10

    The spectrum between self-hosted or on-prem to cloud maximalism is something we think about a lot as well. There does seem to be a somewhat surprising pull towards on-prem these days, or self-hosted. Why do you think that is?

    Erik Bernhardsson32:36

    I think it's just easier to grow your business that way, right? First of all, for a lot of products, it's not necessarily a core impact on revenue, right? If you're building a SaaS business and self-hosting versus multi-tenancy, if you fundamentally charge roughly the same thing, or maybe even more for self-hosting, I don't know, then it doesn't really make a big difference. And if the quality of the service is kind of similar, then why even bother insisting, right?

    Erik Bernhardsson33:00

    Let's just go and meet the customers where their demand is. For us, I think the product is much better with the multi-tenant service, and the revenue potential is much larger because we have full control over the underlying COGS, like the underlying cost, and we can do all these optimizations. So for us, I think there's more of a case to be made for pushing for this thing. For many other types of business, I don't think it matters that much. And so I think they're better off just meeting the customers where they are.

    Erik Bernhardsson33:17

    But I do think there's a lot of customers, also a lot of vendors out there, that maybe should push more for a cloud-maximalist approach, because I do think that's where things are going. Who knows? Maybe I'm wrong. Maybe in a year or two, if we do this podcast again, I've been giving up and just banging my head against all these late-stage enterprise customers, and I've just given in and said, actually, you know what?

    Balancing costs: GPU vendor fees vs. customer pricing

    33:35
    Erik Bernhardsson33:35

    There's a big business in on-prem. And maybe that's fine, I don't know. But right now, that's just not where I want to focus my energy.

    Matt Turck33:53

    I know that you already decreased your price once to reflect a decrease in the price of GPUs, but how do you think about that layer of value between what the GPU vendors or the cloud providers that provide the GPUs charge you and what you charge customers?

    Erik Bernhardsson34:21

    Yeah, I think fundamentally, through our software, we have the ability to charge more than the underlying GPU costs. Now, how much more is the question, right? Customers would probably freak out if we charged 10x more, because then they're just going to go, actually, as much as I like the developer experience of Modal, I'm not going to pay 10x. I'm just going to go deploy it directly on whatever, Lambda Labs or something like that. And so the question is, how much more can we charge?

    Erik Bernhardsson34:40

    And it's probably like 2 to 3x. I think one of the things I always think about is, are we a WeWork in the sense that do we make long-term GPU reservations and then we have short-term sort of income, and then we have this massive sort of exposure to the underlying costs?

    Matt Turck34:45

    And then you can have a GPU-adjusted EBITDA.

    Erik Bernhardsson35:02

    Yeah, I think probably the answer to that is we need to be a little bit more conscious about not spending a lot of capital on very long-term reservations, especially in this climate right now where no one can predict the GPU prices. GPU prices themselves are kind of strange, right? For a while, H100 prices were almost going up, and then they were kind of still for a while, and then now they're actually going down quite a lot, from what I'm seeing in the market, right?

    Erik Bernhardsson35:29

    And so we have to be cognizant of that and lower the price. Does that make us a commodity? I don't know. I look at AWS. You can argue, is AWS a commodity? Are they not? But at the end of the day, EC2 is like, what, like 50%, 60%, 70% margins? It's actually not as much of a commodity as people think. So I think the value of software and the value of creating this pool of liquid capacity lets you actually have a decent amount of pricing power.

    Erik Bernhardsson35:53

    And in the long run, I also think there's many other add-on services we can charge for as we go up the stack that let us retain decent margins. I also think a lot about GPU prices. What happens if they crash? What happens if they go up? Are we long or short GPU prices? Fundamentally, I actually think that since we're sort of—the value we create is like the software on top of the GPUs—that in a way, if GPU prices were to crash, the relative value of that software actually goes up.

    Erik Bernhardsson36:26

    So I'm not necessarily sure that a crash in GPU prices would be bad for us. I think it actually could be great for us because if GPU prices go down, there's going to be more applications using GPUs, and then people are going to pay more for software, just like they do for CPUs, by the way. Margins on Lambda are much higher than margins on GPU serverless products. So I spend a lot of time thinking about this, but I do think there's a lot of pricing power in building infrastructure as a service in this space.

    Matt Turck36:41

    What do you think that is? What do you think you nailed to be able to get that amount of love and interest?

    Erik Bernhardsson37:06

    I think we did a few things right, which is, like, one was just me building kind of the product I always wanted to have and just being kind of opinionated about what that looks like. And then not being afraid to actually go deep into sort of layers of infrastructure and change things that sort of got in the way of delivering that experience. To me, it's like we live and die by the quality of the ergonomics of the tool. And so, to me, that's like the core competitive advantage.

    Erik Bernhardsson37:33

    That's the sort of core ethos of the company. We cannot absolutely lose that. That's why people pick Modal. Everything else is sort of secondary. I also have a long background in building consumer products my whole career, so I think that's maybe helped kind of thinking about what is the onboarding experience like? What does it feel like? Sort of unboxing, when you get started, download a client, install it—what does that feel like? I think that's kind of often like a missed opportunity to create the sort of set the expectation that this is a magic tool.

    Designing products for humans

    37:56
    Erik Bernhardsson37:56

    So, spending a lot of time thinking about that, I think it also sort of inherently led to more of a bottoms-up go-to-market approach. We're so focused on developer experience, but yeah, just being obsessed and maniacal, thinking about what the SDK looks like.

    Matt Turck38:15

    Yeah. And you just published, a few weeks ago, a blog post actually a little bit on that topic, right, of sort of onboarding experience, or at least a part of the blog post was about the onboarding experience and how you make sort of the complex easy. Do you want to summarize maybe some of the ideas?

    Erik Bernhardsson38:37

    Yeah. So what I wrote about is generally just like, how do you write code when you're writing code for humans? When the actual end user of that code is a human? I'm talking specifically about packages, frameworks, even programming languages. How do you think about code when the end consumer of that code is a human? So how do you think about their mental model, their onboarding experience, making it easy for those people to take your code as building blocks and compose larger things?

    Erik Bernhardsson39:13

    And a lot of those things are just kind of things I've been thinking a lot about for the last 10, 15 years, building things like Luigi and other things, and Modal, of course, now. And to me, it's actually not like rocket science. There's a bunch of very core principles of just, how do you craft this experience that makes a great experience? And for Modal, it's so important. We focused everything we can on that. But I actually kind of feel like it's somewhat repeatable for many other people too.

    Matt Turck39:42

    Yeah, I mean, it was a really interesting post, and we'll put it in the show notes and all the things. But what were some of the core ideas? I think there was a part about having less, not explaining the whole thing from A to Z, but showing examples of the final product and what that looks like, and what happens if you remove one part or the other.

    Erik Bernhardsson40:00

    Yeah, exactly. I think there's so many products where, like, so many dev tools, or you go check out some code on GitHub and it's like, maybe this is interesting, maybe not. And you're like, all it says in the README is how to configure it and how to install it, but not actually showing an example of how to run it. And I'm the least patient person in the world, and there's like 50 billion gazillion dev tools out there and so many GitHub packages and whatever.

    Erik Bernhardsson40:30

    So when I encounter something like that, I just don't have the energy to sort of visualize, how does this fit into my codebase? And so I think it's such a lost opportunity for all these dev tools. Show me some examples. Show me how to build something. I don't know, maybe my mind is different than other people. I just don't understand when I look at core concepts, stuff like that. When I look at a dev tool and it starts out introducing a bunch of weird things, I just don't have the energy to kind of visualize how those would fit together to solve the problem that I have.

    Erik Bernhardsson40:56

    I'd rather just look at an example that sort of reminds me of the problem I'm trying to solve. A lot of this actually goes back to my time at Spotify and kind of what Spotify optimized for. It's kind of a weird, very different product, obviously, but it has a lot of analogs. And one of the things I think Spotify did incredibly well was the sort of first experience of, like, you download Spotify and you search for music and you double-click on it and it just plays music.

    Erik Bernhardsson41:31

    And you do that within a couple of seconds, right? It's so obvious, it's so visceral almost. It's funny, it actually goes beyond that. You're like, is this music streaming in the cloud? But it feels like it's local, right? And that's, in a way, exactly the sort of experience I wanted to have with Modal, right? Is this code running in the cloud? No, it's actually running—or it is running in the cloud, right? And nailing that sort of first experience and setting that expectation for this is how the entire product is going to be, to me, that's a big part of how do you sort of make a magic developer experience.

    Erik Bernhardsson41:59

    As well as many other consumer products. It goes well beyond developer tools, but I build dev tools, so that's why I write blog posts about that. So to me, that's a lost opportunity in many dev tools, and that's why we spend so much—I mean, to the point of having a bunch of frivolous stuff. We have silly confetti stuff in the web browser and stuff like that, which maybe went too far, but I'd rather go too far, like making the onboarding experience magical.

    Erik Bernhardsson42:12

    Because it really sets that expectation.

    Matt Turck42:19

    Do you cut stuff as well when you look at the product? It's like, oh, this is too complicated, and after the fact, you just remove and simplify?

    42:43
    Erik Bernhardsson42:43

    Yeah, we probably should do that even more. But yeah, we definitely cut out a bunch of features that are rarely used in order to simplify the SDK. If something is not really used by more than some small minority of customers, I'd rather just get rid of it and merge it into some larger concept and make it a special case or something else in order to reduce the number of concepts you need to learn.

    Matt Turck43:06

    Those developers that show you so much love, a lot of them seem to be working at great companies like Ramp and Codegen, Substack, and a bunch of others. How did you find them, or how did they find you? What's your early go-to-market motion, I guess, to use VC terms?

    Erik Bernhardsson43:10

    Yeah, honestly, it's like Twitter shitposting.

    Matt Turck43:15

    I knew this had an actual use case.

    Erik Bernhardsson43:37

    Yeah, exactly. Reinforcing it. I don't know, I'm trivializing a little bit. I think I have a background in consumer products, and I had always been tweeting and blogging and stuff like that. So I think every company has to find their own channel mix and their own go-to-market, and everything is idiosyncratic, and there's no thing that really translates. For us, it just turned out that targeting early data influencers and talking about the product and giving away hats online on Twitter, that's always worked for us.

    Erik Bernhardsson44:10

    And I think now, more recently, we hired our first salesperson six months ago, and I'm trying to figure that out. And to me, that's a whole different world. Maybe that's a different conversation we should have in a year or two, what that looks like. I don't know, but I think that's the next step. So almost everything is inbound so far.

    Matt Turck44:25

    That moment is always very interesting in the life of a company, right? So the salesperson, what do they do? Do they take some of the inbound and just have that first conversation? And I know you're figuring it out, but what do they do currently?

    Erik Bernhardsson44:48

    Yeah, there's a fair bit of that. I also think fundamentally, when you look at later-stage companies, I do think that a more traditional sales model will make sense for those companies. If you're selling to, an extreme case would be Bank of America. You can't expect them to sign up using GitHub and just start writing code and deploying things and swipe a credit card, right? There's going to be some procurement process, there's going to be some proof of concept, there's going to be a bunch of other stuff.

    Erik Bernhardsson45:12

    So we're trying to figure out exactly what that looks like and how it fits into our model. We've been very happy with the inbound and just kind of upselling, and that's been enough to get to the point where we are. But I also think layering on some sort of top-down will help us get to the next layer. To me, actually, the best companies that I look at for inspiration would be Datadog or Mongo or something like that, or AWS, actually, frankly, started with a product that customers loved and nailed the bottoms-up, self-service process.

    Managing early engineering team

    45:32
    Erik Bernhardsson45:33

    But then also, once that got to a certain point, started layering on more of a top-down, traditional sales process.

    Matt Turck45:44

    You've had some interesting thoughts on how to manage the team in terms of giving them a lot of leeway. What is your philosophy on managing that early engineering team?

    Erik Bernhardsson46:08

    I think that in the early days, actually mid-stage too, you can get very far with almost no management. And I look at a lot of the lessons from early Spotify days. Spotify probably went too far. No one told me who my manager was for the first year at Spotify, which is probably bad. But I think it's sort of an extreme example of the freedom under responsibility that existed in the early days of Spotify. At the end of the day, people were super commercial in building the stuff that they thought was the best for the company.

    Erik Bernhardsson46:41

    And kind of self-organized around that. There's a little bit of Scrum and stuff like that, but actually I think that was kind of bad. So, to me, I'm not naive about it, but I think aspirationally, you can get very close to self-organization if you just hire fairly smart people, very entrepreneurial people, competitive people, and just give them the context. You just tell them, this is what we're trying to build, and then you go figure it out. And then obviously, as you grow, you're going to have to kind of put things in different swim lanes and have a little bit more project management.

    Erik Bernhardsson47:17

    But you can actually get much further with a lot less of that than people think, is my experience. And you're actually better off doing that because you sort of retain people's autonomy. You retain people's ability to sort of come up with solutions themselves, especially when you're hiring smart people. They want to do that, and you should do that. You should harness the creativity of the sort of the largest surface area, which is the individual contributors.

    Matt Turck47:35

    And how does that work on a daily basis? Because you sort of don't manage them, but ultimately manage them at the same time by virtue of being the CEO. Do you catch up, or do they come to you when they have a problem, or do you just look at the final product? How does that work?

    Erik Bernhardsson47:59

    Yeah, I mean, I think you can look at the final product. You can talk about what they're working on. I don't necessarily believe in those sort of estimates or deadlines or things like that. But I do think there's an element of looking across the portfolio of your current projects and continuously adjusting. Are we spending too much time on this? Are we spending not enough time on this other thing? And sort of continuously adjusting that. Every month, we do a lightweight sort of set-the-priorities-for-the-month, and here are the things we're trying to accomplish, communicate with all teams, make sure there's nothing that's missing.

    The only correct way to add a new function to the company

    48:26
    Erik Bernhardsson48:26

    Sort of the, let me see, we have everything covered and there's nothing that's missing. But that's like a Notion doc. That's like super, very basic. And then we do track things in Linear, but for most of the sort of day-to-day management, it's like people just come in and write a lot of code.

    Matt Turck48:41

    You had an interesting tweet a few weeks ago: the only correct way to add a new skill/function at a company. I don't know if you recall off the top of your head, but you can comment on it, like from doing the thing yourself, like that evolution?

    Erik Bernhardsson48:58

    Yeah, it was just a kind of reflection of how I've always built functions at startups, is you can skip steps, but it's generally kind of a bad idea to, let's say, you're not doing sales and suddenly you go from that to, let's hire a VP of Sales. Now you're skipping three steps, or maybe more, right? The correct way, I think, to build up sales—and by the way, this is probably the thing I have the least experience with—but, or like HR or finance or whatever, you start doing it yourself, and then eventually you hit a wall and it's taking up way too much time.

    Erik Bernhardsson49:42

    Then you can delegate to someone, or you can sort of find a part-time contractor or something like that. And then eventually that takes up too much time, so you convert that into a full-time person. And then eventually you realize you need more people, and then you bring in someone else, or you bring in additional people. You should manage that team initially too. And then eventually you promote one of those people into a leader, and then you can manage fewer people. But I think skipping steps in that sort of growing teams is kind of dangerous because how are you going to manage, say, a sales team unless you know how to run sales and how to do sales yourself?

    Building company in NYC

    50:07
    Erik Bernhardsson50:07

    To me, that seems very hard. I think you kind of need to do it yourself first. Not for 10 years, but do it for a little bit. And then eventually you hire people and promote people, and then you can manage the people.

    Matt Turck50:24

    We're still, I guess, on the tail end of a period where people thought that building a technical company in New York was sort of impossible, even when proven otherwise by MongoDB and Datadog. But any current thoughts on building in New York versus other places?

    Erik Bernhardsson50:41

    I mean, I only built a company in New York, so I don't really know. I'm sure there's many reasons to build things in San Francisco or elsewhere. I just happen to live in New York, and I think New York has an amazing talent pool. If I lived in, whatever, Jacksonville, I probably would have moved. But I think New York has such a good talent pool that you can totally build a company here and build a very large one and be very successful, as evidenced by the companies you mentioned.

    Erik Bernhardsson51:04

    So I'm very happy to—I just love the city. I have a wife and kids here. I'm not going to move in order to start a company. And luckily, New York is a great place to build these types of companies.

    Matt Turck51:11

    So, newsflash to everyone: two European guys who chose to live in New York and work in tech do think that tech in New York is great.

    Erik Bernhardsson51:32

    I think tech in New York is amazing. I actually think, on that European thing, I grew up in Sweden, right? And I still have a network in Sweden. So one of the things we've done at Modal is to hire a bunch of people in Sweden. That's actually something I think New York is very good for, right? If you're a European founder or have access to European talent, we have many examples of that. I think Datadog has a bunch of people in France, for instance, because they're French founders.

    Erik Bernhardsson52:01

    And with the time zones, it's much easier than SF, right? Because it's like you have a six-hour difference, you can actually have a decent overlap and coordinate across the time. And there's such amazing technical talent in Europe, especially low-level stuff, Linux containers, whatever, the kind of stuff that we need. So for us, that's been a great sort of competitive advantage, is having a team in Stockholm. It's still pretty small. There's six people, but they're very core contributors to the infrastructure.

    52:05
    Matt Turck52:09

    What are you building next in terms of a roadmap you can share?

    Erik Bernhardsson52:30

    I hope to build this for the next 20 years. So I think there's so much stuff we want to build right now. We're focusing a lot on just taking the existing compute platform and improving it in various ways. One thing, for instance, I'm very focused on is how can we get into more real-time, low-latency use cases, which is kind of an annoyingly complex technical problem. We've sort of relied on a centralized control plane running in one single region, us-east-1, but we sort of recognized, in order to get to latencies below 150, 200 milliseconds, we probably need to split that up and run a decentralized control plane.

    Erik Bernhardsson52:47

    So that's going to take a long time too.

    Matt Turck52:49

    What's an example of speech-to-speech?

    Erik Bernhardsson53:18

    Speech-to-speech, I think, is really exciting, like real-time translations or voice manipulations. There's a lot of interesting use cases around that, or video, actually, is another use case. So low-latency stuff is something we're really interested in focusing on. We're interested in kind of pushing the boundaries of very large-scale batch processing. Modal right now, when you start doing millions of function calls or billions, you start to run up against limits. But we're very interested in how can we push that into millions and billions in a reliable way so you can truly run this very large-scale batch.

    Erik Bernhardsson53:46

    That's another thing we're very focused on over the next six months. What else? I think as we go up this, as we sort of mature the compute layer, I think there are many opportunities to go up the stack and build more enterprise features, build more observability into the tool. We're very interested in deepening our enterprise security compliance offerings and the less technical side, like SOC 2 Type 2, and we're going to be available on the AWS Marketplace pretty soon.

    Erik's predictions on AI

    54:04
    Erik Bernhardsson54:04

    Things like that, I think, is kind of exciting too.

    Matt Turck54:22

    Any prediction on AI? Like, you have a front-row seat for what people are building. I mean, do you see any, I don't know, current pattern in terms of, like, is excitement slowing down, plateauing, accelerating? Are we at the beginning of all of this? I think we're so early.

    Erik Bernhardsson54:50

    I don't know. It's, like, weird, right? To me, the best analog would be either, like, fiber optics or, like, mobile or something like that. I still think they're kind of, like, bad analogs. But one thing I, for instance, thought a lot about is the applications that, in the long run, tend to be sticky are the type of things that really truly take advantage of things that couldn't be built previously. And that's why I'm really excited about Suno, for instance. Suno is a product that clearly couldn't have existed a few years ago.

    Erik Bernhardsson55:14

    And I also think music is kind of always, like, a good sort of place where you tend to see new applications pretty early. And similar to mobile, I've been asking myself, what's the sort of Uber or Spotify or Instagram of AI? What's the sort of consumer application that is AI-native in a certain way, right? And I don't know if it's, like, ChatGPT. I think ChatGPT is fantastic in many ways, but it feels like, I don't know if it's, like, long-term, it feels like I'm looking for something that's a little bit more consumer-y and something that's a little bit more vertical.

    Erik Bernhardsson55:38

    And is truly AI-native. And I don't know if we've seen that quite yet. I think something like Suno is probably the closest I've seen in that vein. But that would be what I would be looking forward to.

    Matt Turck55:45

    Thank you so much, Erik. Thanks for coming by and sharing all of this. And very exciting to see what you're building here in New York.

    Erik Bernhardsson55:47

    It's great. Yeah.

    Matt Turck56:06

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you in the next episode.