MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Hippocratic AI: A Safety-First LLM for Healthcare with CEO Munjal Shah

    Munjal Shah is the CEO and Co-founder at Hippocratic AI. We cover why healthcare’s biggest opportunity is adherence and support rather than diagnosis, how voice agents costing roughly 18 cents an hour could provide “super staffing,” and why Hippocratic AI relies on clinician-led testing and safety kick-outs before deployment.

    08/23/2023

    Hosted by Matt Turck · with Munjal Shah, CEO and Co-founder, Hippocratic AI

    Healthcare AILLMsPatient supportAI safetyVirtual care
    Listen now
    YouTubeApple PodcastsSpotify
    49 min · 1 chapters
    Contents

    Transcript

    Full episode

    0:00
    Matt Turck1:38

    Munjal, welcome to The MAD Podcast. Today we are going to dive into the fascinating intersection between healthcare and generative AI. So, to set it up, you are the co-founder and CEO of Hippocratic AI. Hippocratic AI is a young startup that is building the first safety-focused LLM for healthcare, and we're going to talk extensively about what that means, with an overall mission to increase healthcare access for millions of people around the world via generative AI. While you are a young company, as I mentioned, you've already raised a couple rounds.

    Matt Turck2:19

    I read $50 million in seed, but then more recently another $15 million tranche, making it a total of $65 million. So, congratulations on your early fundraising success. I'd love to start the conversation maybe with the origin story of the company. As I was prepping for this, you're coming to this very much from the machine learning and AI world, Like.com, which was, I believe, a computer vision company that was acquired by Google. So I'd love to hear your narrative of the journey and how you went from computer vision and AI into the world of healthcare.

    Munjal Shah2:58

    Yeah, I am a serial entrepreneur. I love building companies. I love using them to try to make an impact on the world. I'm a computer scientist at heart. I did my undergraduate and graduate work both in CS and in machine learning. As an undergrad, my senior thesis was building a neural network to predict protein-ligand binding efficacy for 3D-modeled drugs. I wrote the code to parallelize it on the first multi-node supercomputer because back then things didn't work so fast, and you didn't have all of the tools you have today.

    Munjal Shah3:28

    Like.com, where we used computer vision to search inside of photographs, looking at the color, shape, and pattern of aesthetic items like shoes and handbags and find you stuff to shop for that looks similar. And we sold that to Google. That was my journey up until that point. And then the day after I sold the company, what should have been one of the best days of my entrepreneurial career, I ended up with chest pains and ended up in the ER.

    Munjal Shah4:06

    I was 37 at the time. My father had had his first heart attack in his mid-40s. So my genes weren't exactly the best genes, but I ended up losing 30, 40 pounds. I'm not that tall a person, so that kind of weight's actually pretty significant on me. And then really got into healthcare, actually took classes in endocrinology and loved it and said, you know what, the next company I want to build is in healthcare. And then I spent 10 years building a healthcare company only to realize, man, healthcare's hard.

    Munjal Shah4:31

    It gets a lot harder than pure tech is. But honestly, like so many others, I saw ChatGPT come out, and I said, oh my God, AI finally works. Like, we've all been talking about it, we've all been trying to use it. It's been a bit of what I call an idiot savant till now. It could do one narrow thing well, but if you took it anything outside of that, it just didn't work.

    Munjal Shah4:55

    At least traditional classifier AI. And I kind of said, but now there's a real potential. And I said, this is the time. This is kind of what I've been waiting for my whole career: to combine these two passions of mine and build something like Hippocratic AI.

    Matt Turck5:37

    Great. So obviously the healthcare field is very wide, and one can think of various applications of AI in general and generative AI in particular. One is clearly the field of drug discovery. Another one would be diagnosis, helping doctors come up with the right call on what may be happening. And then there is the whole world of healthcare management, operations efficiency, and all the things. You've resolutely decided not to focus on diagnosis. I'd love to hear the rationale why.

    Matt Turck5:48

    So if generative AI changes everything, does it not change that and open up the opportunity for diagnosis?

    Munjal Shah6:10

    I don't think it's safe enough. I mean, it's really simple in some ways. Like, I don't know why people keep—it's like flies to the flame. I mean, they just keep going there. I'm gonna do diagnoses because it's grand or it's challenging, but I mean, they're gonna kill somebody. I'm like, this is really serious. I don't want that to be my legacy. So my view was, and they've not done the math.

    Munjal Shah6:38

    $6.2 trillion. Let's take all the spend that goes to doctors and take that as a proxy for diagnoses, right? It's not a perfect proxy, but it's a decent proxy. That's only $600 billion. Like, let's go solve that. They're like, oh, let's do drug design. I mean, drugs are not that much either, right? So, I mean, they're still a big chunk. Maybe I think there's $400 billion, $500 billion. I forgot the exact number, but it's in that range.

    Munjal Shah7:01

    I'm like, look, there's so much more. There's another $3 trillion that's non-drug, non-diagnosis. Let's go make that part of healthcare work. And it's not like the country is short making the right diagnoses. Our biggest issue in the U.S. and most of the developed world is that people don't follow the directions after they're diagnosed. Like, it's not an issue of diagnosis, it's an issue of adherence, it's an issue of ongoing attention and support, it's an issue of getting the other barriers out of the way that's keeping them from being able to manage their health well.

    Munjal Shah7:42

    Like, let's go do that. And so I think what I found was, I think just most people who build LLMs just did not have enough experience in the healthcare world to understand there are so many other problems to be solved that LLMs can solve. And so they just solve—it's like when you see startups by kids dropping out of college, they're always like, here's how you jump the line at a club, here's how you find a college roommate. Like, every startup idea is in one of those because that's just what their lived experience is, right?

    Munjal Shah8:13

    And here, I think that's most people's lived experience if they didn't work in healthcare, but they just were consumers of healthcare. And so those are the ideas that they think of. Now, it's fine by me. I'm more than happy. I'm like, hey, all of you keep trying to do diagnoses, keep trying to work on that. But I'm going to go do this other thing over here.

    Matt Turck8:47

    Great. So that narrows it down by quite a bit. But obviously, in the $3-plus trillion that's left, there's a lot of things. And bearing in mind for anybody listening to this that Hippocratic AI is pre-launch, just looking at the website, you cover a broad range of things that go from healthcare administration and helping people just process operations faster to having dietitians involved. So what is the space you're operating in and the general problem set that you're looking to solve with LLMs?

    Munjal Shah9:14

    Yeah, so I lay out three different uses of LLMs to impact healthcare, and that's why I call it a healthcare LLM, because healthcare is broader than just diagnosis, right? It's all the things going. But there's three use cases. First, let's call it productivity in workflow. That means you're in the electronic medical record, you're helping the doctor write their answer or communication from a patient, something called the in-basket, which is kind of their equivalent of an inbox, effectively.

    Munjal Shah9:52

    Or a doctor writing a note to an insurance company to escalate. A lot of people have come up with these ideas. They're interesting, they're helpful. They're actually probably better suited for the people who sell those software systems today to build in. But they don't change healthcare that much. They maybe make you 5% more efficient, 10% more efficient. It's not clear if you make somebody 10% more efficient that they see 10% more patients.

    Matt Turck9:54

    Right?

    Munjal Shah10:19

    In fact, if they're right now spending every evening answering their in-basket, which most doctors are, and that's time they would rather have spent with their kids, when they get that 10% back from that time, they spend it with their kids, rightfully so. But that means the system didn't get any efficiency. That means we didn't see more patients. And that's what happens. The second use case is healthcare has a massive staffing problem. Most people don't know this, but in the pandemic, a lot of healthcare workers quit.

    Munjal Shah10:46

    They just got burned out. It was just too much, and they quit. And so in the country today, we're 20%, 30% understaffed in nurses alone. This isn't just true in the U.S. This is true in every single country out there. You look at the wait times in some areas of Montreal to get a primary care physician, they're like three years in Canada. The wait times in the UK are also significantly long. So then we said, hey, look, what if we could use LLMs, add voice to them, and have them call patients in a virtual care setting and do a bunch of tasks in an automated way, right?

    Munjal Shah11:19

    An autonomous, automated way, still with a safety kick-out. If they sense you're saying something, they will need to kick out to a real human being, but you're gonna get way more leverage out of that. So solving the staffing crisis is really where we see a lot of the big ideas. And we're going in and solving some of those things. Because how are you going to get 30% more nurses overnight? You're going to grow them from scratch?

    Munjal Shah11:44

    It's going to take years and years and years and years to do it. So that's kind of one aspect of it. Then there's a third idea, and we call this super staffing. In fact, this is the idea that I think I'm most excited about. And it says, hey, let's not just try to fill the staffing level, fill the gap to get to the level we think is the right level. A large language model speaking at 100 words per minute will cost somewhere around 18 cents an hour.

    Munjal Shah12:14

    For the LLM, plus the ASR cost, the automatic speech recognition cost, the text-to-speech cost, the TTS cost. We kind of put it all in there and we're like, this is 18 cents an hour, maybe a little cheaper even. If we get new inference chips, it'll get even cheaper. There's all kinds of interesting things that go there, but it's almost zero already, right? And that's how you should think about it. What are the things in healthcare we don't do that we would do if it was zero?

    Munjal Shah12:40

    So do we call every patient two days after they start every new medication to see if they're having any side effects and then quickly try to change their dose or change the medication if they are? Because what we don't want is for them to have a really bad experience, like, oh, that thing made me throw up all day, so I just stopped taking it. Because that's really what happens most of the time. And so I talked to the head of a pharmacy who's one of our partners, and I said, what would you do?

    Munjal Shah13:04

    Like, do you call? And they're like, no, we don't call. I'm like, why don't you call them? There's not enough pharmacists even to do the stuff we have to do, let alone something like that. And we don't make enough money per pill bottle to ever justify making a phone call to check in on your side of things.

    Matt Turck13:05

    Mm-hmm.

    Munjal Shah13:28

    And I said, oh, would that lead to better care? Yeah, it would lead to better care. Would that lead to better retention for the pharmacy? Yep, better retention. I'm like, great, why don't we do that? Here's another example. There's 68 million Americans that have two or more chronic diseases today. Only the top 1%, 2%, maybe 3% get chronic care nurses that call them up. Chronic care nurses don't do diagnoses. They just call up and say, hey, did you take your medications today?

    Munjal Shah13:55

    Do you need more refills? Did you make an appointment for your follow-up? Do you need a ride to your follow-up? They're just helping you manage your disease. They're care managers, in a sense. But why don't we have 68 million chronic care nurses? Because we can't afford to. At $90 an hour, your average cost of a nurse in most areas of the country, there's no way—it's actually ROI-negative. Because that uncontrolled diabetes might be bad, but it's not going to kill you tomorrow.

    Munjal Shah14:22

    The system actually won't save money for years and years and years intervening. But it'll pay the cost today. Most people also don't know this. They're like, oh, an ounce of prevention is worth a pound of cure. I'm like, no, the problem is time. It turns out the average person in the U.S. changes healthcare plans every three years because they change jobs every three years. And so the healthcare plan will bear the cost to do the prevention but won't basically see the benefit because actually most of the benefit will accrue to Medicare once you're over 65.

    Munjal Shah14:41

    And so, look, we believe there's a huge ROI in doing what we call super staffing, massively overstaffing the system in a way we've never done it before.

    Matt Turck14:42

    Yep.

    Munjal Shah14:45

    And that's the big idea here. That's what we should use LLMs for.

    Matt Turck15:09

    Fascinating. So as you build the product, are you thinking, or are you building a copilot kind of experience for the staffers? Or are you eventually also creating those super staffers as independent AIs that will be disconnected from any kind of human?

    Munjal Shah15:33

    We think all of the leverage only comes from creating kind of the automated bot, but with the right kick-out. But I don't know that we won't do a transitional strategy with a copilot. But I think of the copilot as different. There's copilot, where you're helping the nurse, right? And telling her, hey, say this, say that, do this, do that. Then there's copilot that's like a three-way call. Why don't we just have it on the line? And when the nurse talks to the patient and says, hey, do you need a ride to your appointment next week?

    Munjal Shah16:00

    Oh yeah, I do. I don't have a way to get there. Which is a major barrier of access for healthcare in the country. And so great. Then the LLM pipes up. Maybe the LLM was introduced at the beginning. The nurse introduced the LLM, said, hey, here's my LLM, Rachel. Or she won't say it that way. Here's my virtual assistant, Rachel. She's here on the line. Rachel goes, hi, how are you? I'm here to help out with anything.

    Munjal Shah16:24

    And then while we're talking, oh, you need a ride? And Rachel pipes up and goes, hey, you need a ride? Why don't you guys keep talking? I'll call a bunch of transportation providers right now and see if I can get you a ride. And then two minutes later, because she could parallel dial five of them, right, she comes back and says, hey, I found one. They can't pick you up at two. They can pick you up at 1:30.

    Munjal Shah16:29

    Is that all right? Oh, yes, it is. Okay, let me let them know.

    Matt Turck16:29

    All right, bye.

    Munjal Shah16:52

    Great. And then Rachel can call on her own the second time. “Hey, you remember me? You talked to me already, right? The nurse asked me to call you and just see how you're doing. And anything you tell me, I'll tell the nurse. Don't worry.” I think we have a transition problem. I don't think we ultimately have a conversation problem. These LLMs are very conversant. They can even do things that—think about it.

    Munjal Shah17:14

    Today, everybody's like, “Hey, when I'm engaging with patients, I don't get good engagement.” I'm like, “Yeah, because you call them, nag them to take their meds, and hang up.” Right? When you run a call center, what's your number one criteria? Average handle time. Get in, get out. But what builds relationships? Talking to people. What drives empathy? Then, by the way, the number one driver of whether you thought you had an empathetic doctor was: did they not cut you off when you started telling your story?

    Munjal Shah17:39

    And man, do people love telling their stories. “Well, there I was skiing, and it was a cloudy day, and I had a new pair of skis.” And then they give a lot of detail that the doctor's sitting there going, “This is irrelevant. I don't really need these details.” But in that detail is the relationship building.

    Matt Turck17:40

    Yeah.

    Munjal Shah18:05

    Well, at 18 cents an hour, why don't we let it talk to you for a half hour, an hour, about anything, right? “Oh, I was in this battle in Desert Storm, and that's when I injured my knee originally, and now I'm in rehab for it again.” And, “Oh yeah, oh yeah, that battle.” This thing's read every page of Wikipedia, probably knows every major battle. There is a whole new paradigm here that can be deployed. And I think people haven't even begun to realize what happens when the incremental cost goes to zero and you have infinite capacity and infinite time and really infinite attention.

    Munjal Shah18:19

    What's his name, the author of Sapiens, Harari?

    Matt Turck18:19

    Yes.

    Munjal Shah18:45

    Was being interviewed by Lex Fridman, and he said something so interesting. They're like, “Are you worried about AI taking over the world?” And I'm going to paraphrase him. And he said, “No, I'm not worried about the maniacal AI, this Terminator or Matrix sort of thing. I'm more worried about the fact that the AI is so seductive because it gives you undivided attention. And nowhere else in life do we get undivided attention.”

    Munjal Shah19:12

    And I thought about that when he said that. I was like, “Oh my God, he's so right.” We never get undivided attention. And now in the iPhone world, you really can't even be at a dinner table and get undivided attention when you're having dinner with friends. But undivided attention is incredibly seductive, and people want it badly. And so I think there's some really interesting new conversational paradigms to be had.

    Munjal Shah19:22

    I mean, by the way, there's a real technical issue here.

    Matt Turck19:23

    Yeah.

    Munjal Shah19:31

    Go look at the Llama 2 paper. Go look at the instruction tuning they did. What you'll notice: the average number of instructions on—

    Matt Turck19:32

    There's a—

    Munjal Shah19:36

    They have like seven datasets they showed for their instruction tuning. The average number of instructions: 1.

    Matt Turck19:37

    Mm-hmm.

    Munjal Shah19:44

    Nine. What's the average back-and-forth conversation in a voice-based thing going to be like?

    Matt Turck19:50

    Forty, 100? Yeah.

    Munjal Shah20:12

    Nobody's instruction-tuned for that. They've all instruction-tuned for what I call a query, not a conversation. A query is like, I put in ChatGPT, I get an answer, I modify it slightly, I put it in, I get an answer. And they're long-winded answers. I mean, they should have never called it ChatGPT. They should have called it BlogGPT because it's super long-winded, and you can't even get it to be short-winded. You can try to prompt it and be like, “Please be short-winded. Please say this concisely.”

    Munjal Shah20:39

    And it'll just give you this—or say, “In one sentence.” Try it. It'll give you like a five-line run-on sentence. It's crazy. And I think you need a whole new conversational paradigm. You need to instruction-tune these models very differently to have this paradigm. In voice, you do things you don't do. We repeat words again and again. We say incomplete sentences. We give a lot of active listening clues—“uh-huh, uh-huh”—and those are really important.

    Munjal Shah21:02

    And so, “Oh, that's too bad. Oh yeah.” When you're telling your story of your woes around your health, you need these little affordances and clues. And what we figured out is this is a different model that needs to be built. And so that's why we're building our own model, and we're building it from scratch.

    Matt Turck21:40

    Yeah, so I'd love to dive into this part of the conversation about the model itself. So, it's presented as a state-of-the-art model that has outperformed GPT-4 on 105 out of the 114 healthcare exams and certifications. So I'd love to maybe start with that bit, talk about that process. And then after that, we can get into the more technical details. But maybe walk us through that certification process and how that worked and anything that you learned there. Yeah.

    Munjal Shah22:09

    So we're building multiple versions of our model. The first version we built was, “Hey, let's just teach it to pass every medical certification we could find.” So we literally went and found every last board licensing exam for OB-GYNs, the geriatric nurse certification exam, the ICU nurse certification exam, the NAPLEX for pharmacists. That's the main test that pharmacists take. The NCLEX for nurses. We just found everything. And then we took it, and then we had GPT-4 take it, and we had all the other language models take it, and we beat them all.

    Munjal Shah22:39

    And we beat them on 105 of 114 for GPT-4, for example. And so what we did is we proved that, hey, if you focus on a vertical, you can make a thing that can take tests better. I'm not gonna overstate it. I don't think you built a better healthcare LLM yet, because you really just built a version that could take tests better. But it's a start. And it wasn't that hard to do, which actually is a good thing, because you were like, “Hey, you know what?”

    Munjal Shah23:13

    There's a big gap here. By the way, one of the reasons I think there's a big gap is healthcare data is very different than other data. Imagine an iceberg floating in the water, and everything below the line is stuff that's behind a firewall, and everything above the waterline is on the internet. Healthcare looks like an iceberg, right? Thanks to HIPAA, most of the data's below the waterline. Travel is like an iceberg flipped upside down.

    Matt Turck23:14

    Yeah.

    Munjal Shah23:37

    Most of the content's on the internet. And math is that way too. And I think coding is that way too because of GitHub. People are like, “Oh, horizontal LLMs are much better than vertical LLMs.” I'm like, “When most of the content's on the internet.” Otherwise, how the heck? It hasn't even seen content. So we've done some really interesting things on building our language model from scratch to focus on this.

    Munjal Shah23:50

    We've instruction-tuned it better, as I talked about before. Five trillion healthcare tokens that—

    Matt Turck23:54

    So maybe let's dive into that.

    Munjal Shah24:25

    Obviously, that's how you build a better—yes, yes, yes. So the first version took tests better. The second version we're now building is designed to outperform. It's better, it's instruction-tuned differently, it's fundamentally pre-trained differently on healthcare content we got. We actually just announced 10 health systems that partnered with us, and they're going to be helping us to make it safe and working with us to ensure. So that's part of it. We then continued to—we're doing RLHF differently.

    Munjal Shah24:37

    We're actually using healthcare professionals to do the RLHF, which we can talk about when we talk about safety a little bit later.

    Matt Turck24:49

    Meaning, just to double-click, meaning that you have a, I don't know, a dietitian, a psychiatrist, or whoever that sits down and plays—

    Munjal Shah24:53

    Whatever role it's supposed to play, we're getting—yeah, exactly.

    Matt Turck24:56

    And then they give a like, or thumbs up or thumbs down, to the answer.

    Munjal Shah24:58

    Okay.

    Matt Turck24:59

    Exactly. Okay.

    Munjal Shah25:25

    So we've just gone through—as we're building the model, as we're building the next generation of our model, we basically said, pre-train it better. We're setting up the embedding space differently. We're tokenizing differently. Think of healthcare words. Think of drug names. Drug names are not real words. They're completely made up. They're not the sum of their tokens, right? Like, they're just a completely made-up word. And so we said, oh my God, you have to treat these differently.

    Munjal Shah25:33

    How do you treat these differently? That's what we figured out.

    Matt Turck25:34

    Yeah.

    Munjal Shah26:02

    Most of the other domains don't have their own dialect. Healthcare has its own dialect. It's its own thing. It needs its own vocabulary. So that's part of what we've been able to do. So that's us now working on building that version of the model. We'll then take it and further launch not only an API that other people can use to build, but we'll launch specific roles. We'll say, okay, here is our chronic care nurse.

    Munjal Shah26:24

    We've now instruction-tuned it for every single thing a chronic care nurse would do. Because most people keep thinking an LLM is an application. It's not. It's a rough-draft engine, right? It's better to think of it like an OS. Somebody's got to build Excel on top of it. And that's kind of our view, is that some health systems out there, some insurance companies, can probably take an API from us and build something, but most don't.

    Munjal Shah26:51

    Most need an application. They need to buy an application. And so we've come to realize that most of the health systems in the U.S. are going to need a ready, finished product/application, not just an API.

    Matt Turck27:21

    Yeah. Can you talk about the 80% of the iceberg that you mentioned a minute ago, and around data acquisition strategies? And, I mean, to the extent that there's stuff that you cannot talk about, maybe talk about it indirectly. For anyone that might be listening to this that's thinking about building a vertical LLM, how do people—you just go around and strike partnerships, you pay for the data—how does it all work?

    Munjal Shah27:50

    Look, I think the acquisition of training data is one of the most proprietary elements, so we don't really talk about a lot of it. But I would just say, I'll give you one example. If you're training an LLM, you do need to go and figure out what data—a healthcare LLM needs a broad set of data, and you need to go and figure out where that data is and what it might be useful for. For example, how can you have an LLM that might be asked to function as a dietitian and not know the menu of every restaurant in the country?

    Munjal Shah28:01

    Right.

    Matt Turck28:05

    Interesting. So not all data is health data.

    Munjal Shah28:09

    Well, I don't know. The menu of every restaurant in the country is certainly health data by my definition.

    Matt Turck28:11

    Yeah, exactly right.

    Munjal Shah28:33

    I mean, because what does a person really want in their healthcare dietitian? They don't want a dietitian that educates you on how to eat. Like, nobody wanted that even ever. What they want is, here, I'm at this restaurant, these are the medications I'm on, these are the procedures I've had done, this is my health situation. What should I not order? Or what three choices should I make? And what ingredients should I tell them to take out?

    Matt Turck28:36

    Hmm.

    Munjal Shah28:54

    Well, that's what you want, right? Well, guess what? That world is here, but it's not so easy to get. You have to kind of go and get it. Some of it's in the Common Crawl, some of it's not in the Common Crawl. A lot of it's been done digitally, at least thanks to DoorDash. And things are out there, but you really do have to find and think through content like that and say, what content is necessary and available that will enhance my language model, my healthcare language model's ability to do its full job?

    Munjal Shah29:30

    We've actually taken the time to go and get every single 200-page PDF of every healthcare plan in the country. You need to know what's covered, not covered. It should have read all that, but you're like, oh, it's on the internet. No, it's not all on the internet. It's not hard to get, but you got to go get it. And there's a lot. I think somebody on our team was just telling me there was like 68,000 from one carrier alone the other day that we got of these documents.

    Munjal Shah29:55

    I was like, oh my God. So, again, how could you have a language model that doesn't know what insurance covers? In fact, you would be so much happier if your doctor knew what every single plan covered and memorized it in his or her head. If they knew every formulary, every drug cost—oh, you go to Walgreens, I'm going to put you on this thing rather than this one because Walgreens' formulary under the Humana XYZ plan will have you paying this much copay versus you only pay this much on this deal.

    Munjal Shah30:31

    I mean, it should know all that. It doesn't. No doctor does, but that's part of the real opportunity. So we've really thought broadly, and there's just a ton of content that you just have to think through and really include. And that's part of what we're building. Yep.

    Matt Turck30:43

    So we talked about data acquisitions, we talked about RLHF. The model itself, did you guys start from open source or did you start from scratch?

    Munjal Shah30:48

    We're starting from all sources. We've tried everything.

    Matt Turck30:49

    From open source?

    Munjal Shah31:03

    We've tried—I mean, you don't just sit there and try it one way, right? You try every single way. Every time a new one comes out, you try it. But I will tell you that we are finding that we're having to build almost everything from scratch to get it to work right.

    Matt Turck31:04

    No free lunch.

    Munjal Shah31:18

    There's some benefits, some places you can use one piece for here or there. It's not clear the answer is one model either. It's different models for different things. Speed is a big issue if you're speaking over the phone.

    Matt Turck31:19

    Yeah.

    Munjal Shah31:37

    You got to get the speed right. Voice synthesis. We're not doing our own ASR and TTS. We're leveraging other people's, but we are building our own tone classifiers and things like that to figure out, like, if I'm speaking to you, I should be able to tell the difference between "my back hurts" and "my back hurts," and I should respond differently.

    Matt Turck31:38

    Mm-hmm.

    Munjal Shah32:00

    And so we have built tone classifiers and things like that to be able to do features. Again, we just thought through the domain, right, and thought through the use case. Like, you wouldn't need a tone classifier for most interactions with an LLM, but you do in healthcare because a lot of the information is in the anger, the pain, the frustration. That's where you get a sense of how big a problem this is to the person.

    Munjal Shah32:07

    Yeah, versus it's just a slight issue.

    Matt Turck32:24

    Yeah. How do you think about the hallucination problem, or more generally the fact that AI is a predictive technology that gets it right like 90% of the time, in a context where the outcomes can be pretty high?

    Munjal Shah32:49

    Yeah, that's why I don't do diagnoses. Don't do diagnoses because the thing hallucinates, and you're going to kill somebody. Like, I mean, I think it's like a very—so the number one way to deal with hallucinations is pick the right applications. If I build you a scheduling agent for scheduling your MRI and it hallucinates, you know what? Not the end of the world. Your appointment might get booked a little wrong. We'll fix it. Like, that'll get fixed at some point.

    Matt Turck32:52

    Right.

    Munjal Shah33:12

    So, pick the right applications there. If it makes up a dish that's not at a restaurant when it's recommending something, like, okay, not the end of the world. So, the right applications is one of the keys. Second, actually the LLaMA paper showed something pretty interesting, right, where you overtrain on your high-quality sources. So since I've licensed some of the sources, I'm like, I know this is a good source, so massively overtrain on that source.

    Munjal Shah33:43

    That's exactly what they did, and they showed it brought down some of their hallucination rate. Great, that's what you do. And then there's other ways to kind of cross-check it and do it, but I think the number one element of hallucination elimination is some of these techniques. And so my hope is that as we build bigger and bigger models, we will also see a naturally declining hallucination rate. But I think—

    Matt Turck33:51

    Yeah.

    Munjal Shah34:04

    The scaling laws on that remain to be seen because not enough people have published hallucination rate by size above 70 billion parameters. Like, there isn't a lot of data that looks like that.

    Matt Turck34:26

    And as I understand it, you're going to deploy this on a sort of B2B basis, working with healthcare systems. Would you enable them to bring their own data through a vector database, and they would connect with Hippocratic AI, and Hippocratic AI, the LLM, would be able to check against their data? Is that part of the idea, or—

    Munjal Shah34:35

    There might be some document retrieval solutions as well in there, but they're harder to make work than most people think. I mean, there's like this, oh, I can now just cross—

    Matt Turck34:39

    I think that's probably true of generative AI in general. It's just a lot harder than people think.

    Munjal Shah34:44

    You're much better off trying to train the core model to do it if you can. But we'll see.

    Matt Turck34:45

    Okay.

    Munjal Shah34:58

    But there are some cases where they have their own data and they don't want it commingled and they don't want it in the pre-training. And so then what do you do? You have to do something like what you just said. So we're working through some of that. And speaking of—

    Matt Turck34:58

    Sorry, go ahead.

    Munjal Shah35:24

    That's okay. The nice thing about healthcare is a lot of stuff's standardized. The care plan for dealing with certain health issues is put out by the Board of Cardiology for dealing with XYZ issue in your heart. And that is the standard of care for most places. But some places like a Mayo or a Cleveland Clinic will have their own very specific, better protocol that they've developed. But if you look at the 7,000 hospitals in the country, like, 6,900 will probably use the same protocol.

    Matt Turck36:08

    So let's talk about ethics and regulation a little bit. Obviously a very key topic in healthcare. You have made a very clear claim about being an ethics-first, I believe privacy-first as well, model. I guess the name of the company itself. But can you maybe go into this, including that concept of also bedside manner that you sort of alluded to earlier?

    Munjal Shah36:36

    So first, we're safety-first. That's what we're committed to. And inherent in that is some of the other things. I mean, they're just kind of like, you can't be safety-first without doing some of the other things you mentioned. But I have some—so there's a lot of talk about regulating AI. And the question is, what's the best—like, regulating to what end? Well, regulating to make it safe, right? That's our primary goal, at least in healthcare. What's the best way to do that?

    Munjal Shah36:39

    Is it a bunch of top-down rules?

    Matt Turck36:40

    Maybe.

    Munjal Shah37:01

    I'm open to it. I think it might be necessary. I think we should have that conversation. We should work on those answers. But I think there's an even better answer. I call it bottoms-up regulation. Tell me why. Like, if I told you, hey, you have two choices. You can use LLM A or LLM B. LLM A follows the standard thing put out by some governmental body that says it'll make it safe, as long as it follows these rules.

    Munjal Shah37:40

    LLM B had 1,000 chronic care nurses QA the chronic care nurse functionality and was only released—and they gave it RLHF feedback—and then it was only released when a majority of them said it was safe in a blind taste test where they didn't know it was human or not. I'd go with B. Now, the best is both, which is better, and we'll do—I think that's probably the right answer. But everybody's been talking about A. I've almost heard nobody talk about bottoms-up.

    Munjal Shah38:10

    Like, why aren't we using the experts who do the job today, who know the job best, as our way of determining safety? It seems so obvious to me. It's like, we weren't even that smart in coming up with it. It's a pretty obvious idea, but it feels so much more right. Like, I would want 1,000 genetic counselors telling me my genetic counselor is safe. Like, who better to judge the safety of a language model than the medical practitioners who do it today?

    Munjal Shah38:47

    And so that's our approach. And then in all those other areas, look, I think we're all trying our darndest to do the right thing to get ahead of the technology, but we're all still trying to forecast what something will be when we don't yet know, right? So I think we'll have to learn and adapt along the way. Like, nobody has the prescience to see it all up front. And even some of the dangers of prior technologies, like, okay, if I had sat you down in 2000, and when did Facebook launch?

    Munjal Shah38:57

    I don't know, 2004 or '05?

    Matt Turck38:58

    Or something like that.

    Munjal Shah39:23

    Something like that, right? And I'd say, okay, let's think through all the regulation we could do for this platform. Would you have said, hey, I'm worried about election interference, no matter how much thinking we all did in a room together? No way. I mean, nobody could say that, right? Like, it was a downstream implication we couldn't have even—like, so many things had to happen for us to even realize that. And I think that some of this will emerge.

    Munjal Shah39:52

    We're just going to have to be responsible actors that are constantly looking and trying to address it quickly. And that's what we're doing. I named the company Hippocratic for a reason. I made the tagline "Do No Harm." I don't know what greater way to say we're going to try our darndest than the name and the darn tagline. But I think that it's still going to be a process. There's a lot of things we don't know yet that will happen, and we're just going to have to do it.

    Munjal Shah40:21

    And that's why we're partnering with the health systems first. We did not just release this. We did not just build it and then take it to them. We are literally having them work with us in joint development sessions every minute. And we believe they understand better than anybody how to deliver healthcare safely. And so that's what we're doing, but we are doing it in a vertical.

    Matt Turck40:22

    Yeah.

    Munjal Shah40:22

    Yeah.

    Matt Turck40:32

    I heard you say somewhere that you don't have a precise release date. The release date is when it's good enough to be released, given the importance of it.

    Munjal Shah40:56

    You can't say it's safety first and have a release date. Like, because then that means you'll release when you hit the date because you promised a bunch of guys that you'll release it, rather than releasing it when it's safe. And so I'm like, look, we'll release it when it's safe. I mean, or when they tell us it's safe, really. And by they, I mean both their managerial teams, but really I mean their bottoms-up teams. When all the nurses at that health system say this is safe, then great, safe.

    Matt Turck41:26

    In a world where everybody's imagination has been caught by ChatGPT, are you finding that the healthcare regulators, broadly speaking, are excited by AI, scared about AI? Are they knowledgeable about it? Do you have an opportunity to engage in constructive conversations? What's the general vibe, I guess, for lack of a better term?

    Munjal Shah42:00

    I think it's—my observation so far is they're on it. They're focused on it. They're thinking about it, but rightfully so, from what I've seen. They're far more on it for diagnoses, right? I mean, almost all your AI regulation to date from the FDA has been on diagnostic products. And so, again, yet another reason not to do diagnoses, because there is going to be regulation and it is going to be important. That probably should be slower. So, I mean, that's what I'm saying.

    Munjal Shah42:31

    I think there are areas where we can add value and we can change healthcare and we can lead to better outcomes, and we should do those. And there's areas that are just going to take more regulation, need to go slower. We need to maybe even have advances in core model development to take hallucination rates much lower before we utilize it for that. And I think people have to just accept that there are better and worse applications for each technology, even within a vertical.

    Munjal Shah42:43

    Great.

    Matt Turck43:08

    So as one of the people building one of those new species of generative AI companies, any early lessons learned, just about anything? But one aspect that would be interesting is, how do you build a team to do something like this? Do you have mostly engineers? Do they all need to be super PhDs in AI? How many people do you have on the data acquisition front?

    Matt Turck43:19

    What does the overall sort of structure of the company look like, and what have you learned there?

    Munjal Shah43:48

    Yeah, look, AI is about having a few people with really amazing breakthroughs. Engineering concepts make a very big difference. It's not a volume game. It is really a talent game of, you want kind of the most talented folks. And that's been our hiring strategy to a large extent. We just have some of the smartest folks you've ever seen because it doesn't take a lot of people to build an LLM company. It just takes a very specific level of ability, intellect, and skill.

    Munjal Shah44:16

    And so we have very much focused on that. I mean, we got deluged after we launched. We got some like 4,000 or 5,000 engineering resumes submitted to us. It was insane. I had to hire three full-time recruiters just to go through them all and set up screening calls and go through and figure it out. Actually, four, I think at some point we had. And just very talented people. And then we hired a subset of them, and we hired folks from some of the best places out there.

    Munjal Shah44:45

    I think we have like 11 people from Stanford, and we have like four people from Berkeley, and we have five people from IIT, and we have just a host of very, very talented folks. But that's what it looks like. And it's mostly all engineering. But ours is a weird mix. If you come into our company, you'll be like, wow, there's a whole bunch of NLP, machine learning, generative AI, and general integration folks as well, because we also have to integrate with health systems' IT systems.

    Munjal Shah45:28

    And so we have kind of two teams. We have our engineering team that's doing all the integrations. We have our LLM team that's building the LLM. And then we have a group of clinicians. So we have two doctors full-time. We have two full-time nurses on staff—actually, three doctors on staff—all giving constant feedback into the LLM and giving constant domain expertise. So you do have this kind of duality of, you need both types of expertise, and we have it.

    Munjal Shah46:03

    But we are doing something different. We're 100% in person. Our company is literally five days a week. This pandemic remote work experiment, I don't know. It was an experiment. I think startups that move really fast and need to move fast are better off in person. And it's been fun. Honestly, it's been really fun. You come into our office, you could hear people talking.

    Munjal Shah46:27

    It's like there's a buzz where everybody's heads down, kind of cranking. We have lunch brought in, we have happy hour on Thursdays. I forgot, I don't know about you, but my last company, we worked from home starting the pandemic and then never came back.

    Matt Turck46:28

    Yeah.

    Munjal Shah46:54

    And it is so much more energizing. It's so much more fun. I think the younger folks are getting way better mentorship than they're ever getting in a remote environment. And I think our collaboration is so much faster, our time to respond. You can just grab somebody on the shoulder, tap them on the shoulder, and get an answer in 10 seconds. Then in remote work, everything's a 30-minute minimum Zoom block. It didn't even need to be. You're like, why is my calendar totally full?

    Munjal Shah47:23

    So it's been a joy, honestly, just having such a talented team and having them all in one place and just all with a fixed mission. And having kind of the, like, no engineers in our teams are trying to decide, does the LLM say it that way or that way? Boom, there's a nurse sitting right there, and they're like, does this seem right to you? We just have instant feedback because of how we've structured the team. And then you do have people on the content team acquiring content, and that's another whole exercise that seems to never end.

    Munjal Shah47:46

    There's a never-ending amount of content. Just when you think you've thought of all the places there could be healthcare content, you do a brainstorm and you come up with like 200 more. There's a lot of deep content in places nobody even realizes.

    Matt Turck48:17

    Awesome. Well, it's been a wonderful conversation and such an interesting area and such a gigantic opportunity. So I'm personally very excited to see you guys progress over the next few months and see the product when it comes out and all the things. So I want to say, first and foremost, a major thank you for spending time with us, sharing your thoughts and your experience, your lessons learned. And I very much hope that you'll come back when the product is out in the wild with what you're seeing and how things are progressing.

    Matt Turck48:37

    So thank you so much. Really enjoyed it. Appreciate it, and we'll see you next time. Thank you.

    Munjal Shah49:06

    Awesome. Take care. Thanks for joining us for The MAD Podcast. We're back here every Wednesday with new conversations with leaders in the machine learning, AI, and data space. And if you like this show, you can also find a video recording of not only this episode, but many, many more over on the Data Driven NYC YouTube channel. Thanks again, and catch you next week.

    Matt Turck49:06

    Bye.