MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Lamini: Fine-Tuning LLMs for The Enterprise with CEO Sharon Zhou

    Sharon Zhou is the CEO at Lamini. We cover why fine-tuning learns from data while prompt engineering is strongest with little data, how narrow domain boundaries make model guardrails easier, and how parameter-efficient methods cut switching among 1,000 models from an estimated three months to three milliseconds.

    11/08/2023

    Hosted by Matt Turck · with Sharon Zhou, CEO, Lamini

    LLM fine-tuningenterprise AIRLHFmodel inferenceAMD GPUs
    Listen now
    YouTubeApple PodcastsSpotify
    44 min · 1 chapters

    Transcript build · model comparison

    assemblyai+openai/gpt-5.6-terrafree · 66s
    assemblyai+openai/gpt-5.6-terrafree · 62s
    Contents

    Transcript

    Full episode

    0:00
    Matt Turck1:29

    Sharon, welcome to The MAD Podcast. You're the CEO of Lamini, the leading enterprise LLM platform for fine-tuning. And we are going to talk about what that means. So welcome, excited to have you. Been looking forward to the conversation. As tends to be the case with these things, and you and I were just chatting a second ago before starting to record, I'd love to start with your story, including your very impressive educational background. Among other things, I read that at Harvard, you were the first person to do a joint degree in classics and computer science.

    Matt Turck1:42

    So how does something like this happen?

    Sharon Zhou2:16

    Yes. Well, thank you so much for having me. I'm very excited to be here. Yeah, I can give a little bit of history. Actually, most recently, I was computer science faculty at Stanford, leading a research group in generative AI, generative models. I'm also teaching there. I teach actually about a quarter million students and professionals online on Coursera in generative AI, the most recent course being fine-tuning LLMs using Lamini. And I did my PhD at Stanford in generative models, generative AI.

    Sharon Zhou2:52

    I've been doing this for almost a decade. But before my PhD, I was actually a product manager at Google. I was a machine learning product manager on Google Cloud. I was enamored with product and user experience, and how I became that classics and computer science person actually comes from a very different background than most people, which is I started off in literature and loving languages, loving Latin, ancient Greek, French, as we talked about previously. And I'm a people person.

    Sharon Zhou3:22

    I love people. I never thought I'd fall in love with a thing, which is a generative model, for example. So I moved into computer science and found it by way of user experience. I actually felt very intimidated working with technology when I grew up. So user experience was like this holy grail of, oh, we can actually design systems to make them not intimidating for other people to use. I want to do that. I want to make it so much easier for people to use these systems.

    Sharon Zhou3:50

    So that's part of why I'm in technology at all. And of course, I fell in love with generative models completely on a whim. The joke is I was trying to drop out of the Stanford PhD to start a company like Larry and Sergey. And I failed to drop out because I just fell in love with generative models because they were so magical. And that is a joke, but it is genuinely like I told my PhD advisor, Andrew Ng, that day one, that my plan was to drop out.

    Sharon Zhou4:20

    So it was genuinely my true path. So yeah, bringing together user experience, accessibility, making technology easier to use and access, plus my love of generative models is kind of how this company was born, Lamini was born, where I think the future is where everyone should be able to own their own LLMs. And the path to getting there is democratizing access to actually being able to build it, make it possible for people, for enterprises to be able to build it.

    Sharon Zhou4:37

    And this is especially true in this age where LLMs are, I believe, the new IP. So, yeah.

    Matt Turck5:14

    All right. Fascinating. Thanks for sharing. So since Lamini is an enterprise platform for fine-tuning LLMs, maybe a great place to start would be with a little bit of definitions and sort of if you could help us contrast training and pre-training with prompt engineering, and then fine-tuning, and sort of explain what, for a broad audience of people looking to learn about the space, those things mean and how different they are.

    Sharon Zhou5:39

    Right, of course. So fine-tuning is the technology that got us from a research project in 2020 called GPT-3 and turned that into ChatGPT, a billion-dollar app, right? So that was the technology that made that happen. And what it's doing is it's modifying the large language model to take in data, learn new knowledge from that data, be able to have certain guardrails, like you've probably seen: as an AI large language model, I can't answer that question, or I'm a little bit uncertain about that question.

    Sharon Zhou6:13

    And be able to give more consistent outputs. And so that's the technology behind it. Of course, there are different challenges. Happy to chat through what makes it hard. It's very hard. And that's in contrast to prompt engineering, which people have been hearing about. Prompt engineering is actually something that everyone has been doing for a long time, almost a decade before even large language models came about. I think Google is essentially prompt engineering. When you edit your query to get the results that you want, that's prompt engineering.

    Sharon Zhou6:42

    And so prompt engineering is great. And all you want to do is just ask the language model something and be able to actually interact with the language model. So prompt engineering is just a way to interact with a system, whether that be a language model or even a search engine, to be able to get different outputs from the system that you would want. And it's very, very good when you don't have a lot of data. Of course, you just don't have examples to necessarily steer it in a certain direction, and it's very, very easy to get started.

    Sharon Zhou7:16

    And that's in contrast to fine-tuning, where fine-tuning is not as easy to get started. You do need some technical lift, especially on the data side these days. But it does work really well when you have a lot of data. And of course, these are not mutually exclusive techniques. You can do prompt engineering on top of a fine-tuned model. ChatGPT, you're obviously doing prompt engineering on ChatGPT, which is a fine-tuned model. And so, yeah, that's how I would think about it from those terms.

    Sharon Zhou7:23

    Is that helpful? Okay, great.

    Matt Turck7:39

    Yeah, that's very helpful. And just to drive it home completely, training is what you do even before that. That's the multi, possibly billion-dollar effort of just building the model before you do the fine-tuning.

    Sharon Zhou7:40

    Is that right?

    Matt Turck7:40

    Yes.

    Sharon Zhou8:07

    Yes. Oh, yeah. Maybe the stages of what the large language model goes through. So initially, it knows absolutely nothing. It actually can't even produce an English sentence or English words, and having it learn from that so that it gets a basic understanding of English knowledge and certain skills. And I would say language skills, perhaps, maybe not English necessarily. It could be a different language or different domain. And so that's kind of what the pre-training phase is known as.

    Sharon Zhou8:36

    And that is really expensive, costs billions of dollars to develop, especially different iterations of it. And that's kind of what's known as those different base models like GPT-3 initially, LLaMA, LLaMA 2, some of those that have come out. And then fine-tuning makes the model actually be able to interact with different interfaces, essentially. So be able to actually have a chat discussion much more consistently.

    Matt Turck8:58

    Okay, great. All right. So thank you very much for that. So definitions aside, just to take you up on your offer to explain why fine-tuning is difficult, I'd love to—

    Sharon Zhou8:58

    Oh, yeah.

    Matt Turck9:00

    —cover a few of the challenges.

    Sharon Zhou9:19

    Yes, of course. So to get fine-tuning right, it's difficult. And maybe even on the team side, you can tell that it's a big team of top AI engineers at OpenAI, for example, being able to make that actually work effectively for them. And I would kind of break it down into two different categories. One is convergence. So, getting the model to converge is something that's not very easy and automatic right now, and it requires machine learning expertise to make it happen.

    Sharon Zhou9:58

    And that's where you tune these things called hyperparameters to make the model nudge towards a place where it actually improves, as opposed to go off the rails completely. And by off the rails, I mean completely off the rails. It forgets English, et cetera. So that's one path. I think another thing that's very difficult is we're working with probabilistic systems. So what that means is that you might need to do many different iterations. You won't get it on your first shot necessarily, and you might need to train the model several times to be able to get the results that you expect and that you would want.

    Sharon Zhou10:29

    And so, as a result, efficiency is extremely important. So to be able to do this process very efficiently could actually speed it up significantly. And so, to address both of these things at Lamini, we basically offer different auto-convergence algorithms for specific use cases that have been tuned for those use cases. And then for efficiency, we have built-in parameter-efficient fine-tuning with different techniques to train slices of the model to very efficiently be able to adapt it to different use cases.

    Sharon Zhou10:52

    And by efficiency, I mean, instead of something that might take weeks or even months, that's bringing it down to even the millisecond level.

    Matt Turck11:04

    Okay, great. All right. So that's a great segue into the company itself and then the platform, and let us go into all the things. So when did you start the company?

    Sharon Zhou11:08

    Oh, we started the company last year before ChatGPT came out.

    Matt Turck11:35

    Very, very young company. But it seems that you've achieved a lot, and we'll talk about some of this in a minute. So it's an enterprise platform for LLM fine-tuning. So you mentioned a couple of things, but maybe give us an overall product tour. Where does it start? Where does it end? What do you currently have? What are you building? Yeah, a little bit of a tour would be great.

    Sharon Zhou11:54

    Yeah, okay, a little tour. So for the product, exactly what goes into it? Software engineers at different enterprises today use our products, and they're ingesting data. Maybe that's with some of our partners like Snowflake, Databricks, and Nutanix. And they bring either their own models or models on Hugging Face, like any model from Hugging Face, essentially the most popular one being Llama 2, that series, and they put it into Lamini.

    Sharon Zhou12:35

    And the way they put it into Lamini is through different SDKs that we have to easily be able to ingest that data in different formats, whether that be raw docs data, just documents, raw text data, or if that is something where they want an agent flow, different agents, different operators being able to use different tools or be able to route a request from a user to a human or an LLM. And inside of the Lamini product, I would kind of break it down into a couple components.

    Sharon Zhou13:14

    So one component is definitely the fine-tuning piece, where we have this auto-convergence piece, this parameter-efficient fine-tuning efficiency piece. We have retrieval-augmented fine-tuning, which is like merging a bit of the retrieval work out there with fine-tuning. And we have RLHF, so reinforcement learning with human feedback. And then on the inference side of things, we have—and what inference is, is after you fine-tune the model, or even before you fine-tune the model, you want to be able to run the model efficiently. You want it to maybe give out structured output.

    Sharon Zhou13:42

    So we provide interfaces to do that. And then finally, I would say all of this, this whole system, could only run efficiently on compute. So we are compute agnostic. So we actually can run on different VPCs, on NVIDIA chips. And actually, one special thing is we are the only folks who can actually run your language models on top of AMD GPUs. So what that means is that unlocks about 20,000 GPUs readily available today for enterprises to be able to use.

    Sharon Zhou13:54

    And to get a sense of what that means, that means you can train GPT-4.

    Matt Turck14:10

    Okay, super great. So to unpack some of this, you mentioned retrieval-augmented fine-tuning. Is that the same thing, or is that different from retrieval-augmented generation? Is that—

    Sharon Zhou14:11

    Yes.

    Matt Turck14:12

    Moving terms are different.

    Sharon Zhou14:42

    Great question. They are two different things. So we have simple SDKs on top of our system for RAG, retrieval-augmented generation, which is very common these days. Retrieval-augmented fine-tuning is taking that to the next level and incorporating that into the training process. So retrieval-augmented generation, RAG, which is very commonly used today, is a way to actually get information in at inference time during prompt engineering, essentially. But for the model to learn new knowledge, retrieval-augmented fine-tuning is actually incorporating that retrieval technology into the fine-tuning process.

    Sharon Zhou15:16

    And something that I'm very excited about is these two very big communities being moved together. So one community is the AI large language model community, and the other community has decades of research on information retrieval. And I'm very excited to see these two communities really converge to make these models more powerful.

    Matt Turck15:21

    Yeah. And just to play it back in layman's terms.

    Sharon Zhou15:22

    Oh, yeah. Sorry.

    Matt Turck15:48

    No, no, no, it's great. RAG is just basically checking the information, whereas retrieval-augmented fine-tuning would be knowing the information. Is that one way to think about it? Okay, so that's one, but then you can do both. So you can do the fine-tuning, which is a more advanced one, but you can also do RAG. So you integrate with vector databases and that part of the world?

    Sharon Zhou15:55

    Yep, yep. And you absolutely should do both to get the most powerful model. Yeah.

    Matt Turck16:13

    Okay. So that was one thing that caught my attention, what you said. Another thing is RLHF. So do you want to talk about what that means, both in general, but also in the context of the enterprise on a sort of private basis?

    Sharon Zhou16:27

    Yes. RLHF stands for reinforcement learning with human feedback, and that's one of the techniques that was used by OpenAI, and of course now a few different folks, to build ChatGPT and several different models. And what it is at a conceptual level is being able to incorporate human feedback back into the model. Basically, you can think of it as thumbs up, thumbs down, but it really is like a ranking of, hey, the model said these five things. Which one was the best? Which one was the least—

    Sharon Zhou17:11

    Which one was the worst? Excuse me. And how do we actually take in that feedback back into the model as another kind of rich dataset so the model knows how it should be actually behaving? And so this is a great way to actually incorporate what I view as, one, usage data. So being able to, as your users use your product, incorporate that back into the model very easily. And then I also take quite an opinionated approach to this, which is not necessarily—I question the H in RLHF.

    Sharon Zhou17:36

    I don't think it necessarily has to be a human, and I know Anthropic has touched on this a little bit in some of their recent papers over the past year, which is that why can't it be another LLM or a pipeline of LLMs that can help with that feedback? I think manual labeling is very tedious, especially for our target user, which is a software engineer, and I don't think people should necessarily have to do all that if zero-shot large language models are actually doing quite well.

    Sharon Zhou18:01

    So using a large language model prompt engineering, it's actually doing quite well, well enough to be able to actually transform data. Yeah, that's very interesting.

    Matt Turck18:25

    You anticipated my next question, which is that what I know, and I may be wrong, about RLHF in the context of OpenAI or other providers is largely that it's an outsourced operation in overseas places where people spend hours and hours just talking back to the model.

    Sharon Zhou18:25

    Yeah.

    Matt Turck18:46

    So I was wondering what that meant in the context of an enterprise, because it almost feels like there's a human aspect, but there's also kind of the UI/UX aspect of this. Are we all expected as enterprise users or business users to be sitting down in front of the model and keep giving it thumbs up or thumbs down? It sounds like that could be the case, but most importantly, just to play it back, the future is actual models doing this, replacing humans.

    Sharon Zhou19:15

    Yes, it is actually. I think there are models, but where humans come in isn't this weird outsourced task force in the Philippines doing labeling. It's actually just your users using your product. And so there's ways to bake this into a product user experience that I think are actually quite standard. When I think of a product like Amplitude, for example, it's able to look at very standard software engineering, standard products, be able to see how users are engaging with a product and be able to segment those users and also understand all the clicks that are going into what a user is doing.

    Sharon Zhou19:57

    And then you're able to do analysis on, hey, these are users that successfully went down this funnel, et cetera. And I think the same thing can apply here. The absolute same thing can apply here. Of course, I think people are a little bit worried with language models. They don't even grasp it. But I think that's the future. The future is analyzing, hey, based on how users are interacting with this product, maybe it's editing something it generated, maybe it's replacing it, maybe it's accepting it and sending that email, for example, if it generated an email.

    Sharon Zhou20:38

    Those are all user engagements that can then be fed in via RLHF, right? So that is my view of it. It's not going to be this, I hope not at least, sobering view of outsourced labor for this, which, based on my experience for generic tasks, is actually doable. But for very expert tasks, that's actually very, very difficult. And maybe just to give an anecdote, during my PhD at Stanford, when I was working on various different AI models, I worked on an AI pathology project where we had to detect basically a bacteria that indicated stomach cancer inside different pathology slides, biopsies.

    Sharon Zhou21:24

    And we couldn't get labels from our board-certified pathologists at Stanford because their time is very valuable. They were very generous with their time, but I was like, maybe you should go save some people and I will learn. And I did, actually. I spent all my time learning and doing flashcards with Anki flashcards and learning what H. pylori looked like, what this bacteria looked like, and segmenting it on slides and iterating over and over and over again. And that was valuable not just because PhD students' time is essentially indentured servitude time.

    Sharon Zhou22:03

    No, I'm kidding. But essentially, I was training the AI system. I was training the models. So I actually understood how to iterate on what data labeling looked like that dramatically influenced how the model would behave. And when we got the board-certified pathologists to do it, it turned into a model that was actually overconfident because they weren't sure, like, oh, was I supposed to write uncertain over here? I just marked everything.

    Sharon Zhou22:28

    So it actually helped that I knew. And by the end of it, according to them, I was actually better than their residents, the pathology residents, and I got the same agreement rate as two board-certified pathologists. So it got to that level, but that is not something we tried outsourcing actually to Scale AI, et cetera. None of that worked. It had to basically be me. So my belief is that the people building and the people understanding the products will be able to either, through a product kind of like Amplitude but for large language models or something like that, extract those insights and be able to bake them in and further improve the models.

    Matt Turck23:19

    Amazing. And do you think there is a clear path to basically solving the hallucination problem that everybody is thinking about? So between RAG and retrieval-augmented fine-tuning and then, what should we call it, model-assisted RLHF, do you think that you get into the 99% correct kind of model performance, or where do you think this is going?

    Sharon Zhou23:41

    I think we can get to that performance today, but it's based on how you scope out the problem. So if it's a very narrow scope, of course you can get that. Now, the more and more you broaden the scope, the reason why ChatGPT is hard to put guardrails on is because they're trying to go after every possible use case, right? So that's actually quite difficult. That's very general, and it's hard to predict. So as a result, it's really hard to put together an evaluation set before you see what users actually use it for, for creative tasks.

    Sharon Zhou24:03

    And I actually think it's a great thing, almost, that we don't even know what people use it for. I think for a more scoped project that we're seeing with our customers and enterprises, when it's domain-specific, experts do actually understand some of the parameters around it. So I think some of the things to reduce hallucination with fine-tuning are actually drawing these boundaries for the model, these guardrails essentially, for when it should talk about a topic and not talk about a topic.

    Sharon Zhou24:35

    For example, we have demos out there that show a fine-tuned model that has these boundaries on it, and it only talks about Lamini. It won't talk about any other topic. And it's the same thing that makes these models say, "As a large language model, I can't blah, blah, blah." But it's because we can scope it out just to Lamini. And so that makes it actually quite easy to draw those boundaries versus everything, if that makes sense.

    Matt Turck25:17

    Very interesting. So you mentioned as well, as you were describing the product, grabbing models off Hugging Face and Llama 2 as an example. So this discussion around smaller models and open-source models versus commercial language models, very large models. I mean, it sounds, without putting words in your mouth, that you're very much in the camp of smaller models and open source for discrete use cases. Is that fair?

    Sharon Zhou25:40

    I think it depends on how you define small. So we can train up to 100 billion parameters. So that's, I think, sometimes in the large range now. But what we encourage is actually starting with some of the larger models that just start at a better checkpoint and then moving smaller to optimize for efficiency. And so what that usually means, a common path, is not your own custom 100 billion parameter model, which, of course, happy to take that in. It's starting off with Llama 2 70 billion and starting from there.

    Sharon Zhou26:15

    And of course, our customers are playing with all the different Llama series. And I think other constraints are around compute. So how much you want to be spending on compute also matters. And then how much data you have. So the amount of data that you have significantly influences which model you should choose. Because if you don't have a lot of data, sometimes that might mean a smaller model can be better to fully utilize all of that if you're doing pretty heavy training on it.

    Sharon Zhou26:54

    And I would call that domain adaptation training, like the very heavy training on it. But if you're doing something very lightweight on top of something, then actually starting with a bigger model where you want it to have some of those generic use cases is better. So it really depends on what data you have and what use case you have. And I think the industry is starting to learn what that is, is my sense.

    Matt Turck27:14

    As I was prepping for this, I also read about your PEFT framework. I don't know if that's the right term, framework, but PEFT, which stands for parameter-efficient fine-tuning, and then the concept of model switching. So what does that mean?

    Sharon Zhou27:42

    Yes, what does model switching mean? So let's say you have a single server, like a single node with eight GPUs or something like that, or you just have, let's say, a single GPU. Let's make this super easy. And you have a model that can run on that server, right, on that GPU. Great. What if you have, let's say, 1,000 different customers and you fine-tuned a model for every single one of those customers? Okay, you would need 1,000 GPUs to serve each of them on, and that's doing it the normal way.

    Sharon Zhou28:13

    You would need 1,000 GPUs. That is really expensive. And also, they're not super available right now unless you want to go the AMD route with us. So that is not something we recommend necessarily, to scale it out to 1,000 GPUs per customer. So what can you do? What does model switching mean? If you wanted to just use one GPU and switch across those 1,000 customers. So let's say 1,000 customers are hitting their models and you're just switching it on the GPUs.

    Sharon Zhou28:45

    You're changing which model is loaded up on the GPU. The estimate is about three months of switching time, just pure sheer switching time. It would take three months to serve all 1,000 customers. And I don't think that's quite real-time inference from my vantage point. You ask a model something, three months later it comes back. It better be right, by the way. So that's what that is. And then with the technology that we've used with parameter-efficient fine-tuning and different efficiency methods, that time to switch across 1,000 models is three milliseconds.

    Sharon Zhou29:18

    Yeah, so a very, very big difference. That's a drastic difference, which suddenly means you don't need those 1,000 GPUs. You just need that one, right? And that becomes a dramatic shift in how you even think about your product or how you even scale out compute. Like when we talk to customers, this completely changes how they even think about fine-tuning and adapting models to every single one of their users, whether that be 1,000 that we just talked about, 10,000, 1 million, 10 million.

    Sharon Zhou29:39

    And it just becomes a very different type of thinking, essentially. Okay, fascinating.

    Matt Turck30:03

    And you mentioned agents as well. So first of all, maybe let's start with a definition, and then what is the reality in the enterprise of agents and chains? Obviously, it's a topic that a lot of people have been excited about, but I'm wondering how that manifests in reality so far?

    Sharon Zhou30:31

    Yeah, so that is one of our common use cases. So what customers will do is build out essentially these agents, which call out either different language models or the same model, also call out different tools and APIs that they might have internally. And one of those tools could be go call a person at the end of the usual customer service route. So these agents are essentially, I would say, there are a couple different components. One is kind of a routing ability.

    Sharon Zhou31:00

    So being able to route to those different APIs, and that's the language model's ability to be able to plan and execute a task. So planning is actually something that is in the wheelhouse of large language models. This is an ability that has emerged from these large language models that have not been as great before, I would say. So this has influenced not just these agents, but also things like robotics today. And it's very exciting that they can plan.

    Sharon Zhou31:25

    So, like planning a five-step process, for example, to be able to hit API one, API two. If all that fails at the third step, call a human. That's like the planning stage. So it's able to then hit those APIs, and some of those APIs might be itself to ask a question, return some kind of value, and be able to then go on to the next step. So I think that's a very general framework that people are using today and being able to build out to many different use cases.

    Sharon Zhou31:53

    I think some of the common ones are user onboarding, reminders, customer service, of course. But as you can imagine, these can get pretty specific to certain domains. And for those domains, customers are collecting a lot of data from them as they deploy these models, or already have data. And then that is used to then train the model to be more consistent, to be better at planning for their specific use case, to hit the targets that they want it to hit, and to be able to basically return a better result and continually improve it over time.

    Matt Turck32:28

    The model-to-API sort of action seems reasonably straightforward. Is there a difference when you do model to model to model? Are there different levels of complexity?

    Sharon Zhou32:48

    I would say that every model call in a chain results in, because these models are not 100%, some form of error. So error does compound. It also results in latency for every model call. So I would say that those are considerations to take in when doing massive chains of large language models. And the ways to reduce it are: you take the input of the first thing into the chain, you take the output of the last thing, and you fine-tune only one model to do the whole thing, for example.

    Sharon Zhou33:24

    And so that's a common kind of path as well. Yeah, so that's kind of how I would approach that. But first I would prototype it, and it's okay if it's multiple chains long, totally okay. You're getting the prototype up, you want a general sense of what's going on, and then you're like, okay, now we're ready for production, we want to optimize. We want this to improve over time. Let's go down that route.

    Matt Turck33:41

    Maybe a quick word on that exciting partnership with AMD. You alluded to it upfront, but that's actually a product, like a package. I think you call it the SuperStation. Is that correct?

    Sharon Zhou34:07

    Yes, our Lamini Superstation with AMD. Yes, our secret's out. We've actually, in terms of our hosted service, the Lamini hosted service over the past year has been running on AMD GPUs only. We haven't been running on NVIDIA chips. Of course, for our customers, when they're in their VPCs, we run the Lamini software package on top of their NVIDIA GPUs almost exclusively for that. And some of our customers now have both because they have now purchased essentially these Lamini Superstations that include Lamini pre-installed on them with AMD compute.

    Sharon Zhou34:46

    And the reason for this is because of the compute shortage. I think without the compute shortage, it would not be exactly a thing we'd necessarily be doing very explicitly. But with a compute shortage, needing GPUs that are powerful enough to be able to run some of these models, Llama 70 billion, but even Llama 13 billion, and be able to fine-tune it is compute-intensive. So you need that compute. And as a result, to unblock a lot of our customers who can't get them on a tier-one cloud—and by the way, that includes Fortune 500 and very large enterprises who can't do that.

    Sharon Zhou35:23

    So, we've offered this as one solution to doing that. And we offer it in two ways. Of course, we recommend our hosted solution all the time, which is running it hosted in our data center. But we also have an on-premise solution. Essentially, we would ship the box to you, and your DevOps team could handle it. And we do have customers doing that as well.

    Matt Turck35:39

    Great. And part of the realization was that AMD Instinct accelerators are as powerful as NVIDIA CUDA. Was that the analysis?

    Sharon Zhou36:04

    Yes. So maybe a bit of background on that. It does extend a bit beyond a simple realization. It's been in the works for multiple years before the company even started. So my co-founder, Greg Diamos, was one of the original CUDA architects at NVIDIA in 2008 and has built out and shipped very large language models in production to over a billion users, for example, in the Baidu search engine. One of the earliest examples of that. His team has now built out all the major foundation models, including GPT-3, GPT-4, Claude, Llama, Llama 2, et cetera.

    Sharon Zhou36:37

    So, very big team there. He founded MLPerf, MLCommons, basically the standard for machine learning performance in the industry, to make these models very efficient both for training and inference. So with that expertise for many years and with the deep relationships with AMD, we were able to actually really get this working. And this is with the foresight of something called scaling laws that Greg co-invented several years ago, about a decade ago.

    Matt Turck36:53

    Yeah, what is that? I read about this a little bit. I saw it come up as I was reading. What are the LLM scaling laws? What does it mean?

    Sharon Zhou37:11

    It's a crazy thing. It's actually a simple formula, essentially, for intelligence. So what it's saying is that with more compute and data, linearly the model will improve. And that's almost a very crazy thing to realize and to see in these models. So that's what it's saying. So he knew from the onset, from that moment, and everyone on that team clearly knew because they built out now the big foundation models, is compute matters a lot and data matters a lot, right?

    Sharon Zhou37:53

    So public data has gotten us very far with all these great models today. They're amazing. I believe that the next frontier is enterprise data. And so that's why we're building this company as well. And compute matters a ton. And that's why we engaged with AMD before even ChatGPT came out, because it's a very long process of working together and getting this to work. But it's been very fruitful. And I'm very, very excited because we have reached software parity with essentially CUDA.

    Sharon Zhou38:11

    And so what that means is a huge amount of AMD GPUs are now available for use by our customers. And what that means for them is large language models are unlocked for them.

    Matt Turck38:50

    Amazing. Okay. Speaking of customers and switching maybe away from the product into the go-to-market, what are you seeing and learning? I was thinking through this conversation. I mean, obviously you're just at the very edge of this world and just building some amazing stuff. I just put myself in the shoes of a customer, and the level of precision and sophistication and all the things obviously is exciting, but maybe daunting. Do you find yourself doing a lot of handholding? I guess, what is your experience to date in interacting with customers in the real world?

    Sharon Zhou39:28

    Yeah, I would say I'm actually impressed with the rate of learning that customers have had over the past year. We started before ChatGPT, and it was very preliminary. Customers didn't even know the word. Obviously, generative AI was not a thing. Large language models wasn't even a thing. So, just a very different ecosystem. That was obviously a huge amount of knowledge transfer that we would have to do for a customer. And it also narrowed who our initial customers could be.

    Sharon Zhou40:00

    Today, customers are learning very fast. And we see customers at varying levels of the maturity curve, but they are learning very fast. So they'll come back to us after a month or something and be like, "We're ready." So I'm seeing that as really fast. And of course, I've been trying to contribute to that as well through the courses like Fine-Tuning LLMs, which now has reached many customers who have now tried all those techniques and best practices. And I'm really happy about that.

    Sharon Zhou40:28

    It's a scalable way of doing it. But of course, we do have change management as a thing. It is still realistically a thing at different levels of a company. And realistically, having experience and also been around—our team has been around a lot of these deployments in the past—it isn't easy. There's buy-in that needs to happen at every level. And so I think that that's a piece of it. But of course, that's almost a piece of every enterprise sale.

    Matt Turck40:43

    Yeah. Any emerging use case or use cases that seem to be popular, or things that are turning out to be much harder than expected? Any learnings around use cases?

    Sharon Zhou41:08

    So in terms of use cases, I think they're actually very obvious ones, and they should be very obvious. Essentially, I want ChatGPT on my data. So, chatting over some kind of documents. So that's very, very common. Again, the agent one, operator one is very, very common as well. What's not working is getting compute to scale out. Of course, we're unblocking that and also encouraging customers to deploy much faster. So I think this is a very hard thing, to be able to deploy these models quickly, because it's a new piece of technology.

    Sharon Zhou41:42

    But I will say all of the customers who do deploy fast essentially will reap the benefits of it, because even if you launch something that's embarrassing at first, like GitHub Copilot was kind of embarrassing at first, I think there was a lot of—and people didn't say necessarily positive things—but it improves over time. That's the whole point of it. It improves over time, improves your usage, improves your data. And so that's what deploying faster means. And I know it can be very challenging for folks to deploy something that feels like it's a probabilistic system when they're used to deploying deterministic systems.

    Matt Turck42:20

    Cool. Awesome. Well, that feels like a wonderful place to leave it. Thank you so much. I really enjoyed this conversation. Let's do this again. All right. Well, that feels like a wonderful place to leave it. Thank you so much. Really enjoyed this conversation. Where can people find you online and the courses that you mentioned? And how do they learn more about you and the company?

    Sharon Zhou42:42

    Oh, of course. So you can find me on a few different platforms. Twitter and LinkedIn are my most common ones. Lamini.ai will also reroute you there. And in terms of the courses, they're on Coursera, so they're free, or should be free, I hope. And they're free to anyone who wants to just learn about them and play with the technology. It's Lamini Open Core there, so you can actually hit our servers, hit AMD servers essentially, as you run through it.

    Matt Turck42:59

    All right, terrific. Thank you so much, Sharon.

    Sharon Zhou43:30

    Thank you so much, Matt. Really appreciate it. Thanks for joining us for The MAD Podcast. We're back here every Wednesday with new conversations with leaders in the machine learning, AI, and data space. And if you like this show, you can also find a video recording of not only this episode, but many, many more over on the Data Driven NYC YouTube channel. Thanks again, and catch you next week.