MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Lightning AI: Build and Deploy AI with Pytorch with CEO William Falcon

    William Falcon is the Founder and CEO at Lightning AI. We cover why collaborative ML projects become unmaintainable when distributed-training changes are duplicated across forks, why enterprise LLM deployments require auditability and human oversight because models hallucinate, and why he expects specialized small language models to replace giant models for deployment on a single GPU.

    05/03/2023

    Hosted by Matt Turck · with William Falcon, Founder and CEO, Lightning AI

    PyTorchLLM deploymentopen source AIAI infrastructuresmall language models
    Listen now
    YouTubeApple PodcastsSpotify
    28 min · 1 chapters
    Contents

    Transcript

    Full episode

    0:00
    Matt Turck0:52

    Welcome, Will. So you are the CEO of Lightning AI, the platform to train, deploy, and build AI with PyTorch lightning fast. You are based right here in New York on 23rd Street.

    William Falcon0:53

    Four blocks away.

    Matt Turck1:30

    Four blocks away. So, local AI. And you have raised about $60 million in venture capital, most recently a $40 million Series B in 2022, at least based on what's announced. And then Lit-LLaMA, which is your own generative model. So I'd love to use this as a framework for discussion: start with Lightning and then go into Lit-LLaMA. So maybe just as a level set, what is PyTorch, and what was your journey from PyTorch to Lightning?

    William Falcon1:56

    Yeah, well, first, thanks for having me. Thank you all for coming. Hopefully you get something interesting out of this. So, who's heard of PyTorch? Raise your hand. Great. Who's heard of Lightning? Anyone? Great, love it. Who hates Lightning? Raise your hand. Ah, one person. Are you raising your hand?

    Matt Turck1:57

    All right.

    William Falcon2:20

    I love Lightning, but maybe I'll convince you not to hate it. Anyways, so Lightning started around 2015. PyTorch happened a few years later. So you're like, but how is this related? So I was an undergrad here at Columbia, and I was working in computational neuroscience. And at the time, we were using GANs and autoencoders, which were all the rage back then, to try to figure out what was encoded into certain neural activity. So we measured neural activity in some animals, and then we used a model to take that neural activity and reconstruct—for us, it was ImageNet—images.

    William Falcon2:52

    We published a paper at NeurIPS about that, which was fantastic. Last time I published at NeurIPS. Didn't happen again. But while I was working on that, I started with Theano and then I moved into TensorFlow and then Keras. What I always found was that everything was overly verbose and very restrictive. At some point, you just needed the flexibility as a researcher to do things, and the framework locked you in. So I basically started building my own library for my own research purposes as an undergrad so I could go between research ideas without changing too much of my code.

    William Falcon3:22

    Today, that's still true. If you're working on deep learning at a company, or ML, what tends to happen is, if you don't have something like Lightning, you will have to write your whole project and then you'll have to clone it, and then whatever you did to that, you have to reimplement in the other one. If you added multi-GPU support, now you have to do that on the other one as well. Very quickly, you have too many forks, and it's really hard to maintain, and nothing ever works after that.

    William Falcon3:54

    If you're doing your own personal project, you probably haven't really noticed this until you start to collaborate with people. So that was a problem that I had because I was collaborating with a lot of people. So basically, over the last, I don't know, seven, eight years, you can think about Lightning as research into how deep learning code should be structured and what should be factored out and what should live where, right? There's no standards like that for deep learning. There are for React and web and everything else.

    William Falcon4:14

    For deep learning, we've kind of set a lot of those standards now. You see them in other libraries that are not Lightning today, in the way that they adopt their APIs, because it does give flexibility, which is great. So Lightning evolved based on that. Now, that was a private project that I had for a very long time. Eventually, I started my PhD at NYU, and I was working there with Kyunghyun Cho, who did a lot of work on seq2seq and GRUs, if you read that paper.

    William Falcon4:41

    And he's at Genentech now doing amazing things. And then Yann LeCun, who everyone has heard of as well. And they were my PhD advisors. And so there, I took this project that I had and I started using it for my PhD again. And I realized that it actually allowed me to try ideas very, very quickly, and it kind of stood the test of time at that time. It had been a few years, and I could actually work on—I actually started working on NLP at first.

    William Falcon5:10

    This is pre-transformers, right? So we were doing a lot of decoding, seq2seq with attention, that kind of thing. Then I worked on audio, so speech synthesis. Then I worked on video, on images. So again, back to how do we represent images? And then I open sourced Lightning at that point. It was around April or May of 2019, mostly so that other people at NYU could access it. I was sending the code over Slack. I was like, okay, how about we just pip install the thing, right?

    William Falcon5:35

    And so, just open source it for that. My goal wasn't to build anything crazy with it. And so I just open sourced it, named it, put it out there, and then I joined Facebook AI Research after that. So at Facebook, we started training on YouTube. So we were trying to train a contrastive learning model, self-supervised on YouTube. And YouTube's very big, and Facebook has a lot of GPUs. So my advisors were like, what if you had all the compute in the world? What would you do?

    William Falcon6:02

    I was like, well, I'm definitely going to train on videos, and I'm going to use as many GPUs as I can. So this is 2019. Lightning, at the time where I sat at Facebook—which is actually, where are we? It's a few blocks from here, actually. Next to me was Sumit and the PyTorch team. Behind me was the people from FairScale. So if you know FSDP and all these distributed strategies, it's these guys. There's another team; most of them are at Character.AI now.

    William Falcon6:29

    So there's a few—like, it was very early. FAIR was only about 100 people back then. So it was early days of distributed training. We'd just gotten DDP to work. And then I started implementing a lot of this into Lightning. And I started running jobs on the Facebook cluster where one model would train on about 1,000 GPUs. The Facebook cluster has a three-day limit where they shut it down because it's a Slurm cluster, so you have to share it. So every three days, the model gets killed and you have to resubmit the job.

    William Falcon6:57

    So I'm not going to sit here and manually resubmit, so I built fault tolerance into it. Lightning would detect the signal, save itself, and then restart again and continue training. We trained these models for about three months at a time continuously. Today, I think the world's starting to do that. They're trying to figure out, how do you train models on 1,000 GPUs over three months? Lightning's been doing that since 2019. Literally, it's why it's built for that.

    William Falcon7:22

    That's how it started. Back then, we were using half precision. APEX had just come out from NVIDIA. A lot of these tricks, we were putting into it. Over the years, Lightning's evolved into a place where really what I wanted was a place where all of the research in the world and all the tricks can live in a single place for the benefit of everyone. And that's happened today, right? So that's kind of the story of it. Lightning's evolved since then to become a platform as well.

    William Falcon7:49

    So we have PyTorch Lightning, which is what I'm talking about here. We have Fabric, which is basically the Lightning internals as standalone pieces that you can actually put into your PyTorch code. And you can build your own trainers, you can build your own LLMs. If you've worked on LLMs and large-scale training, you probably got annoyed at PyTorch Lightning because you had to mess with the trainer, and the trainer does that for you. So we actually give you a way to write your own trainer, and that's what Fabric does.

    William Falcon8:12

    And then we have the Lightning Platform on the cloud, where we actually can help you train on the cloud, deploy, and we'll get you the GPUs, and all of that infrastructure stuff is kind of gone, and you can actually work with your teams and build AI together. So that's kind of the full stack of our offerings today as well.

    Matt Turck8:17

    Great. And what are some examples of what people do with the framework?

    William Falcon8:42

    Yeah, so who's heard of Stability AI and Stable Diffusion? That was trained using Lightning, right? Thousands of GPUs on AWS. You can go to the GitHub repo and see that today. OpenFold, also trained with Lightning. NVIDIA just announced all these NeMo services. Those are all powered by Lightning. There's over 10,000 companies across the world who use Lightning to train and deploy. Facebook, a lot of it is powered today by Lightning. You have other major companies, which I won't talk about here because I don't know if I've gotten permission.

    William Falcon9:08

    But all the way from banks to self-driving car companies to tech, probably if a company's using AI today, there's a team internally who's at least using Lightning for something. So we work with a lot of companies on how do we help them standardize more of their code across the org? And if you're training LLMs or using APIs to do other things, we can help you do a lot of that.

    Matt Turck9:12

    Great. So where does LLaMA fit into that picture?

    William Falcon9:32

    So we've always been kind of at this level of infra. We're kind of like the mechanics in the F1 team. We're always building these really fast cars, and we're like, we know we can do this, but we've never actually ourselves put an F1 car on a track. So this is our way of doing that. So Lit-LLaMA is a very, very high-performance model. So if everyone's familiar with LLaMA from Facebook, that model was released under a GPL license, which means that—

    Matt Turck9:43

    Do you want maybe to go into what LLaMA does for anybody that may not be familiar?

    William Falcon10:07

    Oh, sure. Yeah, so LLaMA is a language model, right? So it's like an open-source alternative to ChatGPT, for example. So you can grab it from open source, and then you can use it like you would with ChatGPT. And there's many technical ways of doing that. Now, when that open-source repo was released by Facebook, they only gave you the code to do inference, meaning to predict with it. And the model weights were also not given out. You actually sign up for them.

    William Falcon10:29

    Now, the code on the repo is GPL, meaning if you do anything with that code, you have to open-source that stuff. That's what the GPL license does. Which means it's kind of not— and they did it because they want to keep it mostly for academics. So it means it's not usable for enterprises. We took the LLaMA paper and implemented it from scratch completely in Apache 2, and we open-sourced it. It's fully usable for enterprises, but we also gave you the training code, not just inference.

    William Falcon10:52

    We also gave you fine-tuning methods as well. They're all in there. Now, the repo has Lightning in it, so it's very simple. There's not a lot of boilerplate. It's very readable. It's like one or two files for most things you want to do. It's not hundreds of files, and we want to keep it super minimal. Now, the weights are not there. Obviously, you have to grab your own weights, but we will be training Lit-LLaMA with the community as well to then open-source the weights.

    William Falcon11:22

    But that repo, we have collaborations with all the big hardware companies where we have engineers there that work with us to make sure that they run really fast on GPUs and TPUs. I don't think you're going to be able to make models more performant than the people who built the hardware themselves. So that's what these repos are going to be. It's really, really the fastest LLMs possible that you could have.

    Matt Turck11:31

    And are you planning, or maybe you are already, to offer it as a service, or is that open source and for people to do whatever they want with it?

    William Falcon11:51

    Yeah, I think probably unlike most other companies that are out there today, we actually don't really care about models. Our business is to help enterprises build and adopt AI, not to power APIs. So actually, our incentives are truly aligned with the community, which is, we just want to give you the fastest models because we know that we can build things like this and you're going to need to run them, right? So we can help you run them and do all that stuff, or you can do it on your own.

    William Falcon12:03

    That's fine. But no, we're not going to be offering services and APIs and all of that. And really, that's so that we can have the same incentives as the open-source community.

    Matt Turck12:25

    So I've heard you in other interviews say something very interesting. So you're as deep in the space as it gets, but I've heard you basically offer words of caution to people, especially enterprises that want to deploy LLMs. Can you explain why?

    William Falcon12:45

    So it's really hard to tell at a glance how useful something's going to be when you go to the ChatGPT UI and you play with it because it's good, right? And I use it all the time, and it is great. And from there, you're like, okay, I can take this immediately and apply it to my company. No, you can't, right? It's a big gap between a UI where you can chat and even an API to a production-ready system because you need to have auditability.

    William Falcon13:10

    You need to be able to trace what happened. Is your data private? I mean, all of this stuff is going to OpenAI, right? So there's so many things that go into this. First is the viability, like how hard it is to actually put this stuff into production. But second, the model lies a lot, right? And unless your product is a creative product where lying is actually a feature, meaning if it hallucinates something, it's a better image, right?

    William Falcon13:38

    Then that's cool. But you don't usually want your models lying. In those cases, you need to be very careful. So even from a research world, it's really unclear as researchers, how do we keep it from doing those things? It's a lot of ad hoc rules and things you have to do to it. So we haven't really figured this out in research. So I wouldn't go all in on this unless you can have a human in the loop who's helping.

    William Falcon14:04

    So before all this stuff happened, I had a startup where we helped low-income students figure out how to pay for college over text message. And what we did, and this is 2016, 2017, before Transformers, we gave you suggestions. We had a whole dashboard for our internal employees where we gave them suggestions on what to say, but we always had a human in the loop. I think that probably one way to adopt AI today, especially these LLMs, is to have this human-in-the-loop situation.

    William Falcon14:34

    I think it's too early to do it otherwise. The problem is it looks so good that you're bought into, oh my God, it's going to work immediately. Everyone's rushing to it. You have to be very careful. Take your time with it and actually evaluate it because if you do it wrong, there was Zillow or someone, they did something very wrong and it really tanked their business. So one rogue model can really have a huge impact very quickly.

    Matt Turck15:06

    Do you have a rough prediction on when that problem might be solved? Because I think we're all experiencing this crazy moment right now where everything feels exponential and compounding, and the problems that seemed intractable recently suddenly are solved. Do you think the hallucination problem is something that's going to stay, or is it going away?

    William Falcon15:26

    So it's funny because if you were doing a PhD or you were in research, we've been feeling this for many years already. There were like 100 papers a week, and you're like, ah, how can I possibly read all the papers? You couldn't, right? Today it's kind of like, okay, the world realized that this is a thing, so everyone's feeling the same way now. But in research, we've been kind of feeling this for a long time. So I think it's still going to be, I don't know, like five years probably.

    William Falcon15:44

    Like, I don't think it's going to be like one day suddenly it's solved. I think it's going to unlock industries sequentially. So, like, year one, maybe we can do this industry and then this one. But healthcare and finance are like the last ones that we'll actually be able to do there.

    Matt Turck15:50

    And why is that? Is that because it's mission-critical and you need to get it right, or?

    William Falcon15:53

    Well, because those are the ones you have to audit, usually.

    Matt Turck15:55

    Because they're regulated industries?

    William Falcon16:11

    Exactly. So we work with a lot of banks, and when a model goes into production at a bank, you need to be able to explain to the SEC why that model went into production and how, and all the things that happened to it, right? So I remember I used to work at Goldman Sachs on the trading floor. This was 2018, probably, or '17, and I wanted to use deep learning on the trading floor to suggest trades. And my MD looked at me. He was like, hey, so Netflix, if you recommend a wrong movie, it's fine. But if we recommend a bad trade and you lose a billion dollars and you can't explain it, that's different, right?

    William Falcon16:28

    I think that's still true today.

    Matt Turck16:52

    There is a little bit of a new political economy that is getting formed around generative AI. You mentioned LLaMA as an open-source project. Where do you think open source falls compared to the OpenAIs of the world? And why does open source matter?

    William Falcon17:13

    So I think all roads lead to open source at the end of the day. No matter what you do or how much money you have, you cannot compete with the world's resources put together to do something. So some companies will have an edge for a bit, yes, but open source will catch up, and it'll catch up very quickly. It will always be open source. AI came from open source, came from academia. Trace it back to the very beginnings, some of the early stuff that Yann was doing was always open source.

    William Falcon17:44

    And in fact, I think that's why FAIR is probably one of the best open-source labs today, because of that DNA. So it will always come back to open source. The role open source plays, for me, is about giving back to the community, it's about sharing knowledge. If you've been in other sciences like neuroscience, where nothing's open source, progress is super slow. Like, you want to get a dataset, you can't do it. You ask a lab and they're like, no, but I grew this monkey for three years, I need to publish on it.

    William Falcon18:09

    And it's like, okay, well, it'll take a long time. AI was exponential. The reason why we're here is because it was open source. So turning your back on the community today, you can do it. It's your business and you do what you do for profits, it's fine. But none of this would exist if Google hadn't published Transformers and put the paper out there. None of this would have existed if the attention stuff hadn't been out there, if Theano hadn't been created, if TensorFlow hadn't been created.

    William Falcon18:28

    All of this is built on open source. Even if you're going to make money on it, it's fine. We all are making money on it, but it's all built on open source and you should honor that.

    Matt Turck18:47

    There is an emerging stack around AI in general and generative AI in particular in the enterprise that you guys are very much a part of. What else do you think is important? And I'm going into vector databases and LangChain and all the things. How does it all fit in?

    William Falcon19:07

    Yeah, I mean, good question. There's a lot of noise. I think there's a lot of interesting projects, and it's early research, so people are tinkering and trying things. But I would think about it that way. I think it's a lot of research. Like, what's going to last from there? I'm not sure. Like, probably 1% of the things. But I find it fascinating to explore and to see, oh, cool, this is interesting. Like, okay, so you can chain together a bunch of prompts and things happen. That's great.

    William Falcon19:37

    How far can that take you? I don't know, but I think it's a research question. So if people, if we were back in academia, no one would be criticizing these things because it's research and you should be trying things out. But I think because there's VC money involved and all these other things now, it gets a little weird. So I think as long as you can maybe separate both, I think we're okay. But there is a lot of noise in the market.

    William Falcon19:57

    We tend to be agnostic to most of that because we've been doing this for a long time, and we kind of know what we want to be doing and we know exactly where we need to be. But we definitely look at what's happening, and if there's a real trend that might last, we will adjust the tooling. In fact, Fabric is exactly that. There was a long time where we heard PyTorch Lightning users complaining, hey, but I don't have flexibility for this or that.

    William Falcon20:22

    And we're like, it's fine. And then inklings of, oh, but we're doing this thing called pipeline parallelism, and how would you implement this or whatever? And we're like, well, it's very easy to do in Lightning. Until eventually, after a few years, we realized, oh, these people are all doing this new thing called LLMs. And actually there's a gap in the market for a tool for this. And so then we realized that the tool that we had was good for 2016 through 2019 deep learning, but as of GPT-3 deep learning, it wasn't good enough.

    William Falcon20:38

    And so we had to upgrade to this kind of new paradigm that we introduced. So the research changes and the tools will change as well, and we change with it.

    Matt Turck21:01

    I mean, obviously there's been this arms race towards ever-bigger models, GPT-3, 4, maybe one day 5, and all the things. On the other hand, there seems to be a bunch of LLMs that are smaller, and perhaps LLaMA is one of them, I don't know. What does the future look like? Is that a race to having ever-bigger models and whoever has the biggest model wins, or is that a sort of polyglot future where you have a bunch of different models for different things and perhaps it all works together?

    William Falcon21:44

    So if you kind of step back to original ML data science, there's a long history of compression, like autoencoders, where you're bringing things in, you're projecting them and expanding, or kind of going the other way. PCA is another good example. So I think if you extrapolate that, there's always going to be a pattern of go big to get some result and then compress it back. And I think that's what's happening right now, right? So we needed to train for a long time so that we have these LLMs.

    William Falcon22:09

    That makes sense. We will now compress them, and that's happening. LLaMA, you can go big or small. You can have 7 billion, you can fit on one GPU. We love for everything to work on a laptop. That's great. I think what people miss is when you have a massive model, it's sexy to train something on 2,000 GPUs, but how are you going to deploy it? What are you going to do? You really want things to fit on one GPU if you can.

    William Falcon22:31

    So there is a race to that. I think it's a matter of time. I kind of think that the whole large-scale model thing is really a gap in our scientific understanding of these models. Instead of looking at the math and coming up with a better loss function, we figured out that we could throw compute at it and get us pretty good results. But I think the math field will catch up, and we'll have some fancy regularizer that will make all this go away.

    William Falcon22:46

    And now we're back to a small model. So I guess I think that the future is small large language models that are specialized. That's just called CNNs, I guess.

    Matt Turck23:00

    Great. Last question from me, and then I'm going to open it up to you all. Maybe zooming back out, what's next for the company? What does Lightning AI look like in three years from now?

    William Falcon23:24

    Yeah, so I think where we're focused right now is on these really high-performance models. We're in that growth stage of the company, so we're helping a lot of enterprises revamp their current toolsets into this kind of new generation. What we found is most companies have some sort of platform that they built internally, and that's not scaling. They're bringing us in to basically say, okay, you guys know what you're doing here. Help us figure out how to scale AI across the org.

    William Falcon23:51

    Whether you're using other tools, we are complementary to all the tools. The platform that we have is one where you can use all the things that you're used to already. The company will be more of that. I really think about it as, if models are the rockets that everyone's going to be going after, we're like NASA. We help you build them, we help you launch them, we help you—or maybe SpaceX is better, I guess. I don't know. That's where we are.

    William Falcon24:15

    We're really agnostic to this. We will provide certain high-performance models because, like Ferrari, we believe that we can build some of the best models around, and we will give them to you and you just do what you're going to do with them, right? So it's more of that. It's more of that platform experience. And yeah, I mean, we'll be doing a lot of things for the community. We'll be training models. We'll be putting things in open source.

    William Falcon24:27

    But probably in five years, Lightning is the standard platform for you to do anything with AI. And it works with all the other tools in the ecosystem that you want to use as well.

    Matt Turck24:35

    Very cool. All right. As promised, I'm going to open up to questions. I can bring up the mic.

    William Falcon24:59

    This is going to be a little—well, anyway, Cockroach Labs gave a talk here, and they were all open source. And then they ran into an issue with Amazon using their open source for for-profit, and it seems like you've said you've put some things in place for that, or do you see a point in time that you might need to leave open source and/or not? What's the difference between you and Cockroach not needing to do that? So we're actually open core, meaning most of our things are open source, and then the platform things are not open source.

    William Falcon25:31

    So when you run on Lightning—there's, like, a PyTorch Lightning Cloud. When you run Lightning Cloud, a lot of the things there are not open source. Our orchestrators, our security, all the things that you probably don't want open sourced as a company are the things that we don't open source. Meaning, if you're a customer, you wouldn't want us to open source that stuff. But the core stuff that you can run on your own, you can. So we are probably—I mean, every cloud provider today can use PyTorch Lightning to train models, and they do.

    William Falcon26:04

    So that's already happened. It doesn't really affect our business too much because our platform—I mean, you go train a model on a cloud provider versus our platform, it's going to be a 10x, 100x difference. It's very, very different because we know how to configure the GPUs and what drivers should you use and how to set up buckets and all this stuff. We really look at models holistically. The other thing that I haven't mentioned is we have an in-house PyTorch team. We hired a lot of the core leads from Meta.

    William Falcon26:29

    The person who wrote a lot of the compile stuff that you're seeing here, the linear algebra libraries, the person who wrote all the profiling stuff, right? So we have an unfair advantage in talent at the company as well, where we can really help you do things that are very, very unlikely to do on your own or at these other providers. But we won't be leaving open source. Our DNA is open source.

    Matt Turck26:31

    Great. One more question here.

    William Falcon26:59

    All right, thanks. So in the realm of research, how much headspace and time do you have to dedicate to sort of the ideas around narrow AI versus general AI? And I don't know, I heard it framed this way—maybe it's, I hope it's helpful—the question of representation versus learning. So my PhD was in basically—I mean, everyone's, I guess, in deep learning is representation learning, right? So how do you encode concepts? Most of the stuff that I did was contrastive learning, self-supervised learning, right?

    William Falcon27:30

    That led to a lot of the things that you're seeing today with vision. Back then it was like CPC and other things that we were doing. As far as AGI, I don't really care about AGI. I don't even think it's a real thing that you can do. We're not AGI ourselves. I'm not even sure that concept actually exists. Like, if I were AGI, I would know everything about the universe, and I don't, right? So to me, it's really, we are specialized intelligence, and that's the case.

    William Falcon27:42

    Like, are we going to have really smart AIs? Sure, but I don't think there's going to be, like, a god AI roaming around.

    Matt Turck27:48

    So we're safe for now? There's no Terminator-like future?

    William Falcon27:49

    We're all good, yeah.

    Matt Turck27:49

    Okay.

    William Falcon27:53

    I mean, I think we would have time traveled back to now to tell us not to do this, right?

    Matt Turck28:03

    All right, cool. On this note, we're going to call it a wrap to keep it going, but this was fantastic. Thank you so much. Really appreciate it.

    William Falcon28:19

    Thank you for having me. Thanks for listening to The MAD Podcast. If you liked this episode, be sure to leave us a review. firstmark.com/events/data-driven.