Very cool. All right, so let's dig into the product itself. And I was trying to find it in my notes before they all fell. You historically, as you said, started with the storage layer. And as I learned more about the company, it's like this amazing story of starting from storage and then adding one thing and the next and the next, which we're going to get into. But starting with the storage layer, so you have this term I was trying to remember, which is what, a unified disaggregated...
The architecture, we call it disaggregated shared everything. And that is basically what we try to do, is break the fundamental trade-offs that have existed in infrastructure between price and performance and scale and resilience and ease of use. And we thought if we can build one system that's faster than the fastest was before, and cheaper than the cheapest was before, and way more scalable, way more resilient, and that manages itself, then that's a solved problem. You don't need to think about that piece anymore, and you can focus on your application.
And so in order to do that, we had to come up with this architecture.
Yeah. So let's unpack this a little bit. And I'm going to go out on a limb because I'm not a storage expert, but the little bit I know, the problem is that you have different tiers, right? And historically, before you guys, so you put the new thing in the first tier and the old stuff in the second tier. But the problem is to process it. So that's the historical thing. And you guys completely broke that.
We tried to, yes, collapse that pyramid because historically it was relatively easy to say, this is two days old, I need fast access to it. Now I'm moving it down to a midrange system. It's two weeks old now. It's two months old. I'm never going to touch it again. I can put it in an archive. But AI doesn't conform to that model. You need fast access to everything in order to build an AI model. And once you have it, you want to infer on everything to generate value out of it.
And it turns into this loop of random reads over and over and over again. And so you want a new system that enables these new workloads.
And very practically, that matters because if you're a machine learning engineer and you're trying to build something, if you have to just run your system on old data and it takes forever to come back, you basically give up.
We don't want a machine learning engineer to think about infrastructure. The three words that we get back most often from customers, and that we love to hear, are, "It just works." And, yeah, infrastructure is like plumbing. If you don't notice it, it means it's doing a good job. And you need to focus on where you bring value and let us worry about everything underneath.
Okay, so maybe to double-click for the more technical people in the audience, how does that work? How were you able to sort of do this no-compromise kind of approach?
Sure. So every scale-out system that I am aware of before VAST is based, loosely speaking, on a concept called shared-nothing sharding. You have a lot of nodes in a cluster, and each one of them has direct access to a bit of the namespace, and each one of them has responsibility for a piece of the pie, for a piece of that namespace. And we realized that that architecture is reaching the end of its rope. It starts to see diminishing returns in performance as you scale beyond a certain limit.
Resilience is very problematic. If a node fails, you need to recover that node's responsibility, and it can take a week. And during that time, you can't have another failure. And so that limits the scale of it. What we needed to do is the opposite. Instead of direct-attached, having drives in the nodes, we disaggregate. We put the drives on one side of the network, we put the logic on the other side, and we leverage a new protocol called NVMe over Fabrics to make it look like all of those drives are directly attached.
That allows us to scale capacity independently from performance. It allows us to scale with dissimilar parts over time. So you never have to migrate your data between systems. More importantly, it allows us to move from shared-nothing sharding to shared everything. Every node can now see the entirety of the dataset: data on low-cost flash, metadata on storage-class memory, all on the other side of the network, such that they become stateless and they don't need to talk to each other anymore.
And so now you have the most resilient, the most performant, the most scalable architecture. And then we added a lot of algorithms and metadata structures on top to solve the cost equation and to make it efficient.
Okay, so that's the storage layer. Have we covered all the things for storage?
Yeah, it turns out that this architecture is really good for storage. But once we solved the storage problem, our customers asked us for more. They said, couldn't you use this same architecture and break the fundamental trade-offs that exist at the database layer and at the compute layer? And can you build one platform that vertically integrates all of the software infrastructure that is required for these new AI workloads?
So that's how you went from DataStore, which is a storage layer, to DataBase.
And the DataEngine is the latest member of the family.
That's exactly right. DataStore is unstructured data. It's files, it's objects, very, very large, without a real understanding of what's in there. And then there's metadata. Every file, every object has metadata about the file itself. It has metadata that's contextual that comes with it. Let's say it's a genome. Who does this genome belong to? Which sequencer sequenced it, at what time? What are the characteristics of that person? And then you run that genome file through an inference function.
You understand more about it: which genes are in there, which mutations. All of that metadata was just sitting there, and people want it right next to the raw information on one hand, but they want the ability to query it on the other hand. And that led to this idea of a mashup between an unstructured data store and a structured database, enabling smart and insightful questions of the underlying raw information vicariously through this metadata layer.
So the database is to query the metadata, is that what you're saying?
And everything expanded into...
Yeah, everything in VAST is multi-protocol. And so, for example, you can write a file and then read an object. Everything is the same underneath the covers. And so the database is the same. You can write a Parquet object and then query it using SQL natively off of our platform, without needing any higher layers on top. And so now people are using it not just to analyze their metadata, but also as a data warehouse. And in the same way that we broke those fundamental trade-offs, it grows to extremely large scale without compromising on performance, without compromising on consistency and ACID requirements.
And so we find that in the database space, there are more trade-offs to be broken. Row-based databases and column-based databases, that all stems from storage. It stems from hard drives needing to be sequential. If we can give different views into the same information, then you don't need that complexity. You can consolidate, very similar to what we did in consolidating those tiers.
All right, so DataStore, DataBase, and DataEngine. That's the compute part.
Yes, that's where we bring it to life. And so both the store and the database are static in nature. You write to them, you read from them, you query from them. Our customers didn't like that. They don't like the fact that their application is written in that way. In fact, they want it to be—I'm going to put a plug for you now—data-driven. They want everything to be based on the information as it flows in. That genome file came into the system, it should trigger that inference function that runs on it.
It should trigger an incremental training job that runs on it. That should run on a low-end GPU. This one should run on a high-end GPU. If we can build that language of triggers and functions where you can then understand more and more about your information and put that in the database, and that triggers more functions, and you have this recursive machine that's all data-driven, then we have the full stack, and that allows us to do things that couldn't be done before.
And it's declarative. Like, I tell the compute engine where it needs to go based on certain characteristics of the data, or does it infer from the data where it should go?
Okay. So the triggers are based on actions that happen. So you can have on new file, and then you can have a filter on that. It's not every file, but only files of type genome and only that belong to people above the age of 50. And then you call that function if that happens. Or as things get updated, you call another function. The functions themselves can be whatever you'd like. And of course, once we understand things about them, like the length or the urgency, then we can schedule them in a much more efficient manner.
The fourth piece of the platform, in addition to the data store, database, and data engine, is what we call the data space. And that allows us to go across geographies. And so now you can run a VAST instance in public cloud, AWS Region East, or in CoreWeave or on-prem or at the edge. And we stitch all of those instances together into one global namespace. And that allows us to schedule close to where the data is. It allows us to move functions rather than moving data across country lines.
In some cases, it's not allowed to move data out of Germany, but you can run a federated training job that's global. And so that again gives us an ability to do things that couldn't be done before. For example, break the speed of light. Data has gravity. Compute is a lot more lightweight.
So hold on, you're breaking the speed of light?
Well, you don't need to move all of that. Not anymore. You don't need to move all of that information across oceans in order to bring it to one central location. Everything is distributed, and everything stays distributed in that sense.
But just to bring that part of the infrastructure to life, maybe for everyone. So what you're saying, so there's this VAST, which is a software layer, but then it sits on top of different types of systems that can be cloud.
Yes. So cloud, there's going to be the AWSs, the S3s of the world, possibly, which sit on top of bare metal, and we try and go as low as we possibly can to generate this advantage. But yes, continue.
So the hyperscalers, but also the more specialized GPU players. So the CoreWeaves, the Lambda.
Okay, so that's one group. Then there's a second group that's going to be more kind of like on-prem, because you have a partnership with HPE for that.
Okay, so we try and find the most data-intensive organizations out there. We don't care how big the company is. We care about how much data they have. And if it's 100 terabytes, that's not interesting to us. If it's 100 petabytes, it's really interesting. And so that's where we gravitate towards. And those companies tend to have their data across multiple sites. Some of it is on-prem, some of it is in the cloud, some of it they're leveraging these new AI clouds.
And so wherever their data is, wherever it leads us, that's where we go. And we're agnostic to the underlying platform.
Yeah. So most of the data that we're talking about is natural information. It went through some type of analog-to-digital converter, whether it's a video camera or a genome sequencer or a microphone. And so we want to be as close to the origin of the data as we possibly can. Once you have a VAST instance at the edge, you can write wherever the data originated, and then you can analyze wherever you have compute resources to do so. And we take care of data movements and scheduling and all of that stuff.
Okay, great. All right, so the software layer, there's like the kind of infra storage layer, edge, on-prem, cloud, and then what's above it? So how can you query the data? Do you integrate with the query engines and all the things? How does that work?
We do. We try to build our own as much as possible in the data path. And so we've found that in order to leverage this architecture, we can't just bolt on a stack on top of us because it will dilute the value. And so everything is our own, or almost everything is our own. We don't use open source nearly as much as other companies do. And again, from the very low level of managing SSDs and running on CPUs and DPUs and GPUs all the way up to the protocol layers, we write ourselves.