And so that's the kind of software stack that I think really is hard for a lot of the new entrants in the chip space to overcome. I think we're already, like I said, in a world where there are multiple options for silicon. The biggest labs in the world are using multiple different types of chips to do their inferencing and training on.
What would be a plain-English definition? We talked about the chips, but the rest of the stack, the networking and the storage, just walk us through how it works.
When you're running a cloud service, one of the things—you'll train your model or you'll upload your trained model, and you're ready to start doing large-scale inferencing. Well, you're gonna need a place to put your data, whether it's the data that you're using to train with, or whether it's the data that's coming in and streaming in from your end customers. And so having high-speed storage is a really important part of it. And so Lambda offers the AI-optimized file system service that is significantly faster than your standard, let's say, cloud file system, which is maybe more of a traditional NFS-type of thing.
This is a highly optimized parallel file system that's designed for high-performance reads and writes, and mostly high-performance reads. That's kind of most of the workload.
And that's something you built in-house completely?
We have. I mean, it was in-house completely, right? You have to ask the question: what is the definition of in-house completely, right? We've never spun a PCB at this company. We have not authored—for example, we use KVM/QEMU for our virtualization. And so we have both commodity off-the-shelf hardware that has software installed on top of it for some of our storage. We have some storage partners that we work with as well.
But generally speaking, everything that we do on the cloud, I would generally say, is something that we rolled ourselves with the help of the broader ecosystem. Because again, there's no such thing as rolling it yourself unless you're mining ultra-pure silicon from somewhere and then coming up with your own ASML. It's funny.
Yeah, that's the highly optimized storage. What else? The networking part and what other pieces?
Yeah. So I was talking about this one-click cluster product that we've got, and the way for everybody to think about this is, okay, well, look, you've got a bunch of GPUs. Let's say you've got a cluster of 10,000 GPUs. Well, I want to partition that cluster up. And so what it is, is it's a bunch of GPUs, some CPU servers as well, because you need to have an orchestration fleet as well. And then you've got some storage, and all of the CPU servers and the storage servers and the GPU servers are interconnected with the storage so they can quickly read and write from it.
And that communication happens over what's called the in-band network. And then there's the compute fabric, which is where I was talking about all of the model weights and feature activations being shared throughout that compute fabric. And then there's an out-of-band monitoring network where you've got access to whether it's BMC or some of your DPUs. And when you are trying to create a subpartition of a 10,000-GPU cluster, you need to simultaneously partition the in-band, the out-of-band, and the compute fabric.
That complex coordination between, we've got a bunch of bare-metal systems to, hey, we've got a virtualized system that has what's called RDMA, remote direct memory access, that allows them to read and write quickly, not just from the disks, but from each other's memory, the GPU's sort of HBM memory, and allow them to do that sort of direct memory access, allowing it to go directly from a GPU to another GPU without getting copied to the CPU, for example. Having that all work is an immense, immense software undertaking.
And this is going back to the original question: what are people not getting about neoclouds? Well, first of all, the answer is that most neoclouds don't have this kind of technology. Most neoclouds have not made the high tens to hundreds of millions of dollars of software investment that you need to make to build a real cloud system that can partition a high-performance computing environment like this, and then to have it all work with the storage.
Anyways, I guess that sort of summarizes the steps that you need, and you can think about all the different moving parts of a modern—how does an AI data center work? People talk about AI data center, but really you have to go down that one next level down, which is—because if you were to ask an AI data center landlord, a traditional one, what's going on inside of the data center, they'd be like, well, look, we're real estate people and we really outsource this to the GC, but the GC doesn't know.