I'd love to get a little bit of a tour of the different pillars of the product and what it does currently.
So I think the core of it is really just compute and running containers. That's where we started. We always had a vision of building a platform, but the core of it is: we want to make it easy to run code in the cloud. And it turns out that inference is such a killer use case for that, that we effectively spent the last three, four years now entirely focused on just how do we take user code and execute it in the cloud in a way where we handle the scaling and all the infrastructure.
We sort of let everything else wait until we really nail that experience. How that breaks down is under the hood, and we realized in order to deliver that experience, we had to kind of go deep in the layer of the infrastructure and throw out a lot of existing stuff. In particular, one of the things I always cared about is developer experience. I wanted to build a tool I always wanted to have through my own experience building these types of things.
I didn't mention it, but at Spotify I built a music recommendation system. So a big part of my life used to be iterating and shipping and deploying things, or just running various types of batch jobs and debugging things. And at the core of that, one thing I realized, for instance, was if you don't have a good, super-fast feedback loop, you can never make it. So much of developer experience comes from having a super-fast feedback loop where you can take code and execute it in the cloud in a way where it almost feels like it's local.
And if you solve that problem, you solve a lot of different problems. You make engineers more productive because they can iterate and run things quickly. You basically get rid of what I think was the gap between local development and cloud development, which I think with AI and machine learning has always been a big challenge, because you end up having different environments, different toolchains, and then you end up having to take research code and repackage it and rerun it in order to get it out there in the cloud.
So if you get rid of that distinction between local and cloud and just turn it into one environment, you always run things in the cloud, you solve a lot of different problems. But in order to do that, you have to make it fast. You have to make it feel like you're developing locally when you're running things in the cloud. So what we realized, in order to do that, is you have to solve a number of core technical challenges around container cold starting.
So what I mean is you have to take code and you have to ship it to the cloud and you have to start that container. And ideally you want to do that in a couple of seconds at most, because that's sort of a human reaction speed. It kind of feels snappy. And then we started looking at, okay, well, what actually happens typically in a system when you do that is you deploy something to Kubernetes, and then you have to pull down a Docker container image, and that then has to start up.
And all of those steps are super inefficient. So we built our own. We threw Kubernetes out the window, we threw Docker out the window, we built our own file system in order to optimize for how container images are distributed. We built our own scheduler in order to maintain this pool of workers and make it possible to start tasks very quickly. We built our own container image builder because we had our own container image format.
That's a lot of effort. How long did that take?
I mean, it was too long. And this was during the ZIRP era. So I feel like VCs were like, yeah, cool. People were asking, why do you really have to go so deep? But fundamentally, I think we felt like we're going to solve this problem well. So yeah, we spent basically a year and a half, two years, two and a half years, almost like in a cave, just building a product that no one had used, which is maybe a bad idea.
And then when we actually kind of emerged out of that cave, it turned out we didn't really have an obvious use case. So we struggled for six months or a year or so to try to find a use case. And then luckily all this Stable Diffusion came out in summer of 2022 or something like that, and a lot of people started coming to us and were like, actually, this kind of makes sense. You have this serverless GPU thing. You can sort of realize this makes a lot of sense for us to focus on.
And then we started seeing more meaningful traction with users.
The GPUs that I have access to through Modal, where are they? Where do you get the GPUs from?
They're all over. And this is not a secret, by the way. We use AWS, we use GCP, we use Oracle. We have a number of alternative cloud providers now that we're also using. And we run a lot of different regions. We have a bunch of GPUs in the US, we have a bunch in Europe.
And is part of the trick to be intelligent about which one you get where, when?
Yeah, totally. And for a few different reasons. One is just capacity management. From time to time, capacity varies, right? There's this sort of cyclicality. If you look at Europe versus the US, it's actually kind of interesting. You see that when it's early morning, the US is not high utilization, but Europe has high utilization, and vice versa. So you can play all these games kind of just aggregating capacity across different regions. Pricing is the other obvious dimension, right?
We actually have a system that continuously looks at cloud pricing and solves an optimization problem to sort of figure out how do we allocate capacity across 100 different regions in order to deliver the capacity we need to our customers. Because we use a combination of reservations, but also a lot of spot and on-demand capacity, which means we rely on the cloud's ability to scale up and down, because it's a sort of supply-demand matching problem, right? We have volatile demand coming from our customers, which means ideally we also have the ability to sort of match that with on-demand capacity on the other side.
Which is also why I think we found a good product-market fit with inference, because inference is inherently sort of unpredictable, right? You don't know necessarily when you're deploying something what the usage is going to be. There's going to be a lot of daily variations, going to be a lot of weird spikes when something goes viral, et cetera. And a big part of the pitch of Modal is you can use us, we'll just scale automatically. You don't have to worry about making a cloud commitment, especially not a three-year or one-year reservation.
You can just deploy things on Modal, and then if you one day need 1,000 GPUs, we can typically get you 1,000 GPUs pretty quickly, like, talking minutes.
In terms of the model itself, I just bring my own model effectively. I go on Hugging Face and grab some open-source model and deploy it on Modal.
Yeah, you can definitely take a Hugging Face model and deploy it. That's what a lot of people use Modal for. I think where Modal really shines is also people training their own models, like having custom models. One example I always bring up as an amazing use case—I love the product—is Suno, which is AI-generated music. And they have a big cluster; they train their own models outside of Modal, and then they use Modal for the inference side, which means all this sort of generation of AI-generated music happens on Modal at very large scale, right?
Millions and millions of pieces of music generated on Modal.
Sort of full circle with the Spotify experience.
Yeah, it's actually funny because sometimes when I talk to them, I end up talking a lot about licensing and stuff like that, music licensing, because I happen to know a lot about that. But yeah, so what I love about that is they can do what they're good at, which is to build an amazing user experience and a user application and train models. And for all this sort of infrastructure mess, they can effectively outsource that to Modal. And that happens to be something we love and we can do really well.
So I think that's sort of the comparative advantage of specialization in that sense.
The training side would not make sense.
We are interested in training. We're looking at it. I think training at very large scale, people are very price-sensitive; it's somewhat transactional. I think you have very different sort of requirements, typically, like InfiniBand and very high interconnect. It puts a lot of different constraints on the infrastructure that we don't support today. It's something we're interested in, somewhat looking at down the road, but it's not a super high priority for the business right now. I think another use case that we do support is actually a lot of preprocessing for training.
So a lot of users use this for, let's say you have millions and millions of audio files or video files and you want to extract some training data from that. Typically, that requires a sort of batch job that extracts features from it. And that's something that Modal does really well too. It's like very high, bursty batch parallelism, just like fan out, parallelize, run this over millions and millions of audio files or something like that.
And then you're done, right? So it's sort of the ability to be elastic and just grow very quickly.
Totally, exactly. It's sort of pay-as-you-go because everything in Modal is usage-based. So you just pay for the exact time your GPUs or CPUs run.
Fine-tuning, is that emerging as an important—
I think fine-tuning is like—a year ago, I would have said fine-tuning is absolutely critical. I think now the models have gotten so good that this is sort of a question of: is fine-tuning even needed? I think fine-tuning has a lot of different use cases. I think in particular, if you have very domain-specific models, let's say you're focused on financial data or document extraction or something like that, I think fine-tuning absolutely makes a ton of sense. And so we have a lot of users using Modal for that.
I think beyond that, I don't know, fine-tuning is still, to me, a little bit of an unproven thing, but clearly there are certain places where it does make sense.
It's a fascinating comment, actually, in terms of the current state of the AI market and where it's going, right? It sort of feels like we're just coming out of a big phase of training and fine-tuning to some extent, which I guess is a part of training, and going into the world of, okay, we built those things, now let's use them, which is the inference.
Exactly. Yeah. And also, a couple of years ago, a lot of companies tried to train their own LLMs. To me, it makes almost no sense, right? Like MosaicML and stuff like that. So, yeah, I agree with you, right? It's obvious that in the long run, inference spend will dominate just for the obvious reason that that's where you make the money, right? When you're training a big model, you're spending a lot of money training that model. How do you recoup that?
It's typically through inference, right? If you look at OpenAI, any of these businesses. So to me, it's very clear that maybe in the past we've had like 80/20 spend, GPUs training versus inference. I think in the future it might be the other way around. Much more dollars will be spent on inference than training.
You mentioned batch processing, job queues, and all that stuff, which in some way feels more like the sort of more classic data engineering part. Is that a big use case as well?
Yeah, there's a bunch of people using us for data parallelism and pipelines and stuff. There's weird sort of unexpected use cases. I shouldn't say weird—it sounds like I'm dismissing that—but actually fascinating use cases. We've seen in biotech, for instance, companies having—I don't know bio super well, so this might be a gross, generalized, bad sort of characterization—but scanning millions and millions of compounds for certain chemical properties, or medical imaging where customers are running computer vision on millions of medical images.
So there is a lot of sort of batch processing also in those use cases I find fascinating. Not just sort of data pipelines, which is what I think people think about with batch processing, but also all kinds of other interesting applications around biotech or video transcoding or feature extraction and things like that.
And to finish the tour, sandboxed code execution. What does that mean?
Yeah, we started seeing a lot of customers who build LLMs needing, especially people doing LLMs for code, needing safe code execution. As it happens, we had great primitives for that because that's effectively what we built, right? We built a system for containing user code and executing it in a safe way. So we started thinking about, can we expose that as a service? So we also have primitives to take sort of untrusted code and execute that in a safe way. And I would put that in the sort of category of unproven.
It's still sort of early days. We're trying to figure out exactly what that market opportunity looks like. But there's so many startups building LLMs for code, whether it's code generation or migrations or things like that, or debugging or things like that. So to me, that remains an unproven but potentially very large upside market opportunity for us that I'm pretty excited about.
Why did you pick serverless as the right level of abstraction for this space in terms of what developers need?
I think it's funny because I feel like serverless was all that hype. People were really excited about it when Lambda came, and then it didn't really catch on. And 10 years ago, when Lambda came—I actually don't remember what year Lambda came, but I think it's about 10 years ago—I probably would have expected in the future 90% of web development will obviously just be—people don't want to think about containers and processes and things like that. And that didn't really happen.
And I've thought a lot about why. And one of the reasons I've kind of realized maybe the reason it didn't happen is actually the developer ergonomics around building backend services is actually kind of fine. If you're, like, a backend developer just deploying with Kubernetes or running things locally, it's actually not so bad. It turns out moving things to Lambda is actually kind of worse in many ways. If you look at data teams, though, like data, AI, machine learning teams, I think we never figured out developer ergonomics for those companies, for those teams.
And 10 years ago, I was the only data engineer, the only data scientist at Spotify. It was kind of fine because it was such a small fringe part of software engineering that probably didn't really matter. But now there's a lot of people building data, AI, machine learning applications. And I think developer ergonomics and developer experience matter so much more. And as it turns out, actually, serverless kind of makes sense, I think, for a lot of the applications that those people are building.
So serverless always felt to me like, in hindsight, it looks like the right idea but applied to the wrong problems. And I always wondered. Now, looking at what we built at Modal, I'm kind of convinced that serverless actually makes a lot more sense for data, AI, machine learning applications for a lot of different reasons. You have very bursty things, you have a lot of different containers, you don't necessarily care about hyper-low latency. I think there's a lot of different reasons why it actually makes a lot of sense for those types of use cases.