Sharon, welcome back to The MAD Podcast. You and I did a great episode just a few months ago, I guess last year.
But in the AI world, a few months is like seven years on Earth or something. So let's do it again this time. We're going to talk about Lamini, we're going to talk about the markets, but let's stop here and there to talk about specific concepts that may be interesting to a broad audience that's interested in AI but may not be super into the details, into the technical stuff. So let's do that like we did last time.
Thank you so much for having me back.
Yeah, absolutely. So maybe let's start at the high level. I'm curious, from the perspective of a founder who's very much in the proverbial trenches of AI, specifically enterprise AI, day in and day out, what's the current mood in July 2024 as we record this? It sort of feels like the AI hype is slowing down a little bit. What are you seeing in the markets? How are customers responding? What are they doing? What are you seeing?
It turns out having a non-negative margin is an interesting thing, isn't it? Who knew? Okay, so a few things. I think one is from the customer perspective, from enterprises. We sell to enterprises. I think what's going on there has been extremely exciting, but enterprises are actually starting to structure their organizations, having centers of excellence who are tackling this problem, being able to onboard different products internally. And I would say the sentiment is, hey, we've tackled the shallow use cases.
We've been able to put some of those in production, but now we're thinking about what's next. Is this technology really only going to help me compose email? Or is it going to do something a little bit more? Is it going to help me leapfrog my industry, or is it something in between that? What's the next step? What are the steps I need to take to get there? And so I'm starting to see deeper use cases.
I'm starting to see people tackle those deeper use cases across both their data science and engineering and even infrastructure teams. Yeah.
What's an example of a deeper use case? So those shallow use cases, that's like email search, I guess, that kind of stuff?
Yes, yes, yes. Things that for every model ping, it might not be necessarily immediate ROI. So I was talking to a large insurance company this morning, and they said, well, we have all of these essentially field reps, and they have extremely high turnover there, and we want to be able to train them up, right? We want to be able to train them up very effectively, and it needs to be grounded in these real facts about our business.
And right now, we have these interesting demos. We have it kind of in demo land, but it's not in a place where we can realistically deploy them.
We're not in a place where it's actually affecting thousands or tens of thousands or even hundreds of thousands of people. And that's like one step deeper. And then another step deeper is thinking about, okay, well, let's just think about drug discovery, right? Drug discovery takes a decade with a 90% failure rate today. But imagine if a company could pull that in and do, I don't know, like only 30% failure rate in one year.
Like, that would be a dramatic improvement.
And that would be very, very interesting. That would be a true leapfrog where this company would completely shatter that field, right? Like, because you would succeed at a completely different timescale and at a completely different failure rate. So I think that future is coming, but it's not necessarily coming as fast as ChatGPT entered everyone's phones and made it possible for us to send a text message and be able to get some form of an intelligent answer back.
Yeah. So what are big enterprises doing currently? Are we still in a stage where they brought in the consultants, they built the center of excellence? I mean, what are they doing, and what would you recommend they do to move as fast as they need to on generative AI?
So one thing that they've already done, I see, is I think compute infrastructure has been laid out more in place. I think budgets are there.
I think the next thing around organizational structure that's extremely important that I think people don't fully realize is relating the technical piece of building out a successful model with the development piece of this model. And what I mean by that is generative AI is famously very hard to evaluate. We have no idea what's good, better, best, unless someone who's an expert in understanding that use case can tell you that, right? Like, I can't tell you a medical generative AI model is actually that good because I'm not a doctor, right?
I can tell you when it's at, like, toddler-teenager stage, but by the time it's getting an MD, I'm out. I'm not there. So being able to have those people sit closely with the development teams actually accelerates those use cases much more quickly. And that becomes an organizational problem. That becomes a very, very big difference I see across some enterprises where those teams are closer together, so those use cases can get out much more quickly, and then other enterprises where those are much more disjoint today.
So they need to reorg to be able to actually get those closer together in order to deliver those applications. And I think that's why realistically we've seen many more code agents and applications around code, maybe text-to-SQL. We've seen that a lot because the developer can also evaluate the outputs, right? They are one and the same person.
And so that makes development much faster because, again, that is a famously difficult problem in generative AI from back in the research days. That was my PhD dissertation, actually, to today for realistic applications.
Yeah, very interesting. And what are you seeing, again, for enterprise customers in terms of, like, oh, let's just bring in OpenAI on top of Azure or whatever to do these use cases, versus let's play with open source, our own models? And how's the thinking evolving?
Yeah, so I think it's pretty clear that it's going to be a hybrid world, right?
And I think the general-purpose models of the world are very good at these quote-unquote shallower use cases that aren't necessarily business-specific, right? Like, they're not necessarily going to be your proprietary model or heavy proprietary differentiation using AI. But they are going to be important to composing email, maybe writing certain types of basic code that aren't relevant to deep codebases that are very custom. So I think that will always be there, and there will always be a best general-purpose model, whether it be OpenAI's or Anthropic's or Google's, et cetera.
And then I'm starting to definitely see a shift of, okay, well, this is great. This is an absolutely amazing technology. We've gotten to develop advanced RAG and prompt engineering solutions on top of it. But it's still not enough. These models, when they're general, they're optimizing for what's known as generalization error, or the average error across all examples it sees on the internet. And as a result, it's pretty good at everything, but it's perfect at nothing. And despite our flaws, Matt, you and I, we're actually perfect at some things.
Yes, I remember your name. You remember mine, presumably, and your own birthday, maybe, and maybe revenue number last quarter. I remember the liability number in my insurance claim, for example, if I were an insurance company. So I want the LLM to remember these facts with crisp accuracy, right? With factual accuracy where there's no alternative at all. And that means being perfect at some things.
That means being perfect at recalling whether it be a column in my SQL table or, again, that revenue number in my earnings report. Like, that is incredibly important for the reliability of these models. And so slightly right is not the same as right for these facts.
And that is where I think the next frontier is of, okay, well, now we've been able, with memory tuning, which is what I've been working on, to remove those hallucinations, to remove that, and actually get these models from not necessarily being general for everything and instead of being pretty good at everything but perfect at nothing, to be actually perfect or near-perfect at some things and still pretty good at everything else. And I think that's where we're going.
And I think that's extremely relevant to the enterprise.