Maybe a few words on your story, your background. In particular, you were one of the key people on PyTorch at Meta, or Facebook at the time. Yeah, Facebook at the time, I guess. What is PyTorch, for anybody that doesn't follow these things in great detail? And then how did it all come about? What was your role there?
Yeah, that was an amazing journey. When PyTorch happened, what was the kind of industry context here? So I joined Meta in 2015. At that time, Meta was called Facebook. Facebook was going through the transition of finishing mobile-first to starting AI-first. The fundamental reason why it sequenced that way: after Meta transitioned its application from desktop to mobile-first, and it's widely accessible everywhere using your phone, you can connect with people. And that drives a lot of user engagement and, therefore, a lot of data being created.
And then we all know data fuels AI. So quickly after the mobile-first transition, we started AI-first. At that time, there was no software, no hardware, there were no GPUs, no people building all this infrastructure. So we bootstrapped the whole thing from the ground up. And it was an interesting time. It feels like right now there's so many companies building their own AI framework, a gazillion number of AI frameworks. Even within Meta, there were three different flavors: one for mobile, one for research, one for production.
So it's very confusing. And we decided to unify all of that and have one framework aiming to be the best tool for researchers to create innovative model architectures, but also one framework to deploy and deliver those innovative models into production to power all the product needs. So these two actually have huge tension across them because, for researchers, you want flexibility, you want ease of use, you want them to just think about what's possible, right? And for production, it's a constraint problem-solving, as in you have latency budget, you have cost budget, you want to scale, you want to be reliable.
So it's kind of, hey, you have all these constraints, let's kind of optimize. It's an optimization problem. So how to create one framework to solve both problems with big tension in between is a very, very hard challenge. I would say we had a lot of fun thinking about that. It was deemed as mission impossible at the beginning, and many people felt this was so experimental, probably about to fail, but it's worth a try. And I would say our initial attempt was not successful.
When was that, 2016 or...?
From 2015 to 2016. And because the idea is great. The idea was simple because PyTorch has a great frontend, very simple, Pythonic. And we had another framework internally. It had high-performance kernels backend. And you would think we can zip them together so we get the benefit of both.
How hard can that be, right? So let's just zip them together. Internally, it was actually called Zipper. And it turns out it's really hard to bring two frameworks completely designed with different goals and different interfaces and internal APIs to kind of work seamlessly together. It's like building a bridge with both ends moving. It's going to be an extremely fragile bridge. So then we decided, hey, making both sides happy is not the goal. We need to build an extremely strong product.
And we decided we were going to rebuild the whole entire backend from this beautiful frontend and bring it to production. It took us five years. It took us five years to get to the stage supporting almost all internal needs using deep learning at massive, massive scale. So it also has been deployed to not just data centers—Meta owns its own data center—but also mobile phones. If you use any of Meta's apps, that means our PyTorch infrastructure is running on your mobile phone and also AR/VR devices.
It's kind of ubiquitous everywhere. At that time, PyTorch externally, as an open-source project, had grown from a baby GitHub project into more and more massive adoption. So through that, we got exposed to many other companies using PyTorch. And it's fascinating to see the growth of all different use case adoption, from image recognition to ranking recommendation to robotics, self-driving cars. And GenAI is nascent, but it's literally everywhere.
And it's used by all the research labs, right? Like all the—
It started from research labs. And the fun fact is OpenAI switched to use PyTorch fully.
From TensorFlow. Right. So TensorFlow was dominating at that time.
TensorFlow being the Google framework.
Google framework, yes. Google also put a lot of effort behind open-sourcing this effort, growing the community. But simplicity wins. PyTorch is so simple, and the user interface, the debuggability, and people can easily change the model architecture as they want. It's dynamic and flexible.
To drive home the simplicity point: so it does, because you had the sort of low-level and high-level stuff, so it does anything from whatever the research scientists need to do in terms of, I don't know, like a backprop, whatever, on one hand, but on the other hand, it's going to do GPU usage optimization. Is that fair? So it spans all the things that you need to do to develop faster and simpler.
Right. So I think the API kind of adopts Python as the programming language. That's why it's called PyTorch. The idea is from Torch, LuaTorch time. And Python is a very simple, easy-to-use programming language for many researchers. And they think about deep learning neural networks in code, right? You need to think about that as graphs, as nodes and edges and how to tie them together. When this graph becomes like hundreds of thousands of nodes, then it's how do you even kind of manage that thinking process?
But representing that in code, in Python code, is very easy to manage. And then you become a software engineer to think about a neural network. And it's very easy to debug because all Python tooling, the toolchains, can all work. So that actually unblocks a lot of progress made in modeling innovation. And it also, of course, can run on GPU, it can run on different hardware SKUs. We have many different hardware vendors plug in at a lower level to support PyTorch, and that includes TPUs.
So, and because Google has a compiler that can directly connect with PyTorch and lower the program into kind of the GPU runtime.
So that enabled significant advancement from the model research side. And the interesting part of model research is a lot of research has been published as open-source projects. And Hugging Face is one example of adopting and becoming the centralized repo for new research ideas. And it's very easy for people to take this model architecture. So there's a pre-GenAI and post-GenAI era. Pre-GenAI, the model code is open source, but there's no weights, right? Everyone has to curate their own data and train from scratch.
That just means any company that wants to invest in deep learning, they have to first hire people who understand how to curate data first, and then who understand how to train, how to manage GPU fleet, and so on. So that takes a lot of time to hire those people, which come from a very, very small pool, and takes a lot of capital investment to get this going. So that's why pre-GenAI, not all companies can access this technology, even though it's great.
And usually it concentrates in hyperscalers or whoever has big resources to be able to do it. And then post-GenAI, GenAI is basically built on top of foundation models which can learn from world knowledge, internet knowledge. And these models are good by themselves. So applications can develop directly, run on top of those models. So you don't need a machine learning team to begin with to start imagining what are the new user experiences. Or you can have a small machine learning team who just do fine-tuning or reinforcement learning, which requires much smaller samples to curate, and the post-training process is much simpler.