MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Building the Easy Button for Generative AI | May Habib, CEO, Writer

    May Habib is the CEO and co-founder at Writer. We cover why graph-based retrieval handles complex enterprise documents without preprocessing, why production-grade AI requires business logic beyond prompting, and why autonomous action starts with the most knowledge-intensive nodes in a workflow.

    11/21/2024

    Hosted by Matt Turck · with May Habib, CEO and co-founder, Writer

    enterprise AIRAGAI guardrailsAI agentsgenerative AI
    Listen now
    YouTubeApple PodcastsSpotify
    36 min · 12 chapters
    Contents

    Transcript

    What is Writer?

    1:47
    Matt Turck1:35

    May, welcome.

    May Habib1:37

    Hi, Matt.

    Matt Turck1:51

    Thanks for doing this and joining our evening and community tonight. I'm so excited for the chat. And let's start from the top. What would be the two-minute version of what Writer does?

    May Habib2:24

    Awesome. Two minutes. So Writer is a full-stack generative AI platform. So we help enterprises quickly get to value with generative AI by combining LLMs with zero-engineering RAG, AI guardrails, and an AI studio. And so a lot of the problems in enterprise right now around generative AI are being able to get to quality with efficiency and with adoption. And we've solved a lot of those problems in one all-in-one solution.

    Writer's founding story

    2:52
    Matt Turck2:52

    And for further context, you are a venture-backed, San Francisco-based startup. Officially, the last round of VC was a $100 million Series B in 2023. But there's a wide rumor, which I'm not going to ask you to confirm unless it has been announced, but I don't believe so, that you may or may not have raised $200 million at close to a $2 billion valuation, in which case I may or may not congratulate you. This is the information I have. And in terms of history, before you started Writer, you founded a company called Qordoba that was doing machine learning and, I believe, focused on machine translation, sort of a Grammarly kind of focus.

    Matt Turck3:07

    Can you talk about that story?

    May Habib3:42

    It's been really one continuous journey for me and Waseem, truly. And we have been in the realm of automating language for the entirety of the 10 years we have been working together. In the first company, Qordoba, we started with statistical machine translation. So this was years ago. And we started using transformers in that business, encoder-decoders at the time, really to solve NLP problems associated with translation. But it was very easy to see how powerful this tooling was. And we, in a way, felt like after years of working on a very mission-driven business around making the language you were born speaking just a non-issue, especially in the workplace.

    May Habib4:16

    To go work on source language really felt like abandoning that initial mission, but the technology was just too exciting. So in a lot of ways, the story of Writer is the story of the transformer, and it was in the first business that we really started to use the technology.

    Matt Turck4:19

    And Writer officially started in 2020.

    May Habib4:51

    Yes, we got our seed-stage term sheet March 3rd. Of course, I remember that date because San Francisco went into lockdown March 14th, 2020. Never closed money faster in my entire life. We closed on the 29th and really restructured the team of Qordoba down to really just the EPD team and that core team. So our core team has been working together for years, six years, seven years, some of our oldest employees across two companies. So a lot of the cultural values of Writer today came out of together choosing to start over with technology that we had developed.

    May Habib5:28

    And the initial pitch, I'm not ashamed to say, is not nearly as ambitious as it is now. The initial pitch for Writer was we were going to go after the AI writing assistant business, but with transformers instead of 900 linguists writing rules. And it was kind of hard to explain to our investors in that it's technology that can impute the laws of English. So cool, right? What I don't think any of us really foresaw was just how fast and how powerful the transformers would become.

    May Habib6:05

    In the seed-stage deck for Writer, kind of an early slide where it's that just exponential growth slide, and it's like AI writing assistants, we are here. And then AI writing, like, I don't know, a decade away. And very quickly we got to transformers that were good at generating, transformers that were good at insights when you connected them to data, transformers that could translate. And when we exposed a lot of just the scaffolding we put around LLMs to build really useful, reliable, consistent, accurate applications to the customer, we got AI Studio.

    May Habib6:42

    And this is just a few weeks ago, we had a customer basically, in a weekend, rebuild Qordoba using Writer. And it's just so cool that we went from transformers good at editing stuff to software that can create software, to software that can build software that uses software. And I think the thing that we have been really good at is productizing on every step-function change in the capabilities of the models really faster than the market.

    Writer is a full-stack company. Why?

    6:54
    Matt Turck7:22

    You mentioned some of the history, and at the very beginning you mentioned that you're full-stack. Were you always full-stack? How did you start? Because in my mind, when I first started hearing about Writer, it was lumped, perhaps unfairly, in the category of what was then known as thin wrappers on top of somebody else's LLM. Was it ever the case, and then you became full-stack, or were you full-stack from the beginning?

    May Habib7:40

    So we've always had our own technology down to the LLM. There was a very brief period, literally like three months, where we, for some of our apps, used GPT-3 because it was just so much better. But we've been able to be weeks ahead or weeks behind state of the art ever since. For us, though, full-stack did come over time in terms of the guardrails and the knowledge graph and the orchestration and observability and everything now that we offer the enterprise, including flexible deployment, et cetera.

    Writer's enterprise use cases

    7:57
    May Habib7:57

    So that all definitely came over time.

    Matt Turck8:07

    What are some of the use cases? If I'm an enterprise potential customer today, what do I use Writer for? And then we'll get into how it works behind the scenes.

    May Habib8:46

    So if you are a major investment bank, let's say you are the top investment bank in America, you are using us for everything from earnings call summaries to deep dives on specific companies and sectors. You're using us for search over merger proxies and M&A docs. If you are Salesforce, they're a major customer, you're using us for dozens of different kinds of custom applications plugged right into Slack to review content for compliance, to automatically rewrite for SEO, to produce marketing and email. If you are an insurance company, CSAA uses us to build knowledge assistants for agents who are answering the phone.

    May Habib9:34

    So our focus areas, and the verticals that we focus on, are financial services, healthcare, and retail CPG. And the types of use cases are everything from your mission-critical, like reviewing insurance policies and doing claim adjudication on them, to helping salespeople sell more. Now, when we are in call centers and support, it's not chatbots for your website or helping exchange mediums into smalls. There's lots of companies doing that. It is really in supporting a high-end kind of knowledge task. So if you are a CPG company and somebody calls into the call center to ask whether there's phthalates in the Aveeno product, right?

    May Habib10:02

    You are, as somebody answering that call, putting them on hold and going through a bunch of documents to figure that out. And that's where we like to play. So it's a very, very crowded space, and being very specific when we talk to a customer about where we're going to be 100x better than any other solution they're going to see, and way better time to value and way better ROI than them doing it themselves on top of the LLMs and building all the services they need around the LLM.

    May Habib10:27

    That's where I like to play.

    Matt Turck10:38

    Presumably, the whole idea behind the full-stack approach is precisely to enable you to be 100x better than other solutions because you control all the elements of the stack.

    May Habib10:50

    Yeah, absolutely. So you take our knowledge graph solution, for example, and we didn't get to graph-based RAG. You got the AI Studio, you got the knowledge graph, you have the guardrails.

    Knowledge Graph

    10:51
    Matt Turck10:55

    So let's start with the graph that you just mentioned. What is it? What does it do?

    May Habib11:23

    For so many use cases. Let's say you are a major CPG company. This is a lot of what we do at L'Oréal, at Kendra Scott, et cetera. Regulatory affairs and compliance are using Writer to go through federal guidelines, federal regulations, stuff that's getting updated in real time, to produce arguments for why we've got the packaging that we do or why we can make the claims that we do. When you are doing this desktop-based research using LLMs, both to go in and digest the regulations and then actually build and write the report, these are great use cases for us that produce a ton of value.

    May Habib12:03

    Are not going to be asking the LLM directly those questions. You are going to be using an app that uses a RAG pipeline to be able to really get that data into the LLM. We do have domain-specific LLMs, but for a use case like that, you need to set up RAG. A lot of times, and we almost always—we're not as big as Microsoft or OpenAI yet, right? So we're the brand that is coming in and the company that's coming in when there's already a DIY project, right?

    May Habib12:26

    Like, these companies have all been trying to crack the nut for the last couple of years. There is something that just hasn't scaled when we are talking to customers. And so we just have to beat the benchmark. And starting a couple years ago, we realized that we were just like the customer, building these pretty brittle, hard-to-maintain, hard-to-scale applications by using—and I won't name names—but we are using folks that everybody uses as partners.

    May Habib13:15

    And it was just very hard. And we were some of those partners' biggest customers in the early days and helped them raise money and all the things. But it's just really been hard to scale that. And so our graph-based approach is radically different. I mean, we trained a separate LLM—this is our brilliant AI team—to build basically the triples. I mean, it is a flat JSON file. This is stored in a Postgres database. It's not even a graph database.

    May Habib13:49

    You don't need an ontology. But because the LLM is so good at building the graph, paired with things we're doing on the index, we do something called retrieval-aware compression. So we're making most use of that context window by just how we compress the data that goes in with a query. And then we actually use the memory layer of the LLM. So it's a technique called Fusion-in-Decoder. It came out of FAIR that we integrated into our approach. So we built a whole product to solve a really, really important part of getting things right, and it's why it's magical out of the box.

    May Habib14:28

    It's why you plug in your data; you're not doing any preprocessing literally when you build a graph in Writer. And the words RAG appear nowhere in our product, by the way. It's just literally like, upload your file, connect via data connector. Here's the API if you're going to send files, and it just works. And that's really what you need to be able to then go solve all the other problems, right? Accuracy is just one.

    Matt Turck14:39

    So graph-based RAG is sort of the cool term du jour, the cool approach du jour compared to vector-based. Can you maybe compare and contrast?

    May Habib15:08

    In a vector-based approach, what you're doing—and I'm grossly oversimplifying here—let's say we are looking at an insurance policy. And in your head, imagine that you've got an insurance policy that's actually scanned, and you've got tables, and you've got two or three columns on every page, and there's tables nested within that insurance policy. That is like good data when you're talking to the enterprise, right? And if you are taking a vector-based approach to answer a question like, I'm talking to a client, and here is their policy, and they want to know what's the coverage maximum if the trailer gets rear-ended or whatever.

    May Habib15:58

    When you are doing a vector-based approach, you're going through and you've chunked that policy. And against the prompt, you are trying to find the most relevant chunks in that policy, which means it does really badly with tables because you've flattened out the document. You have no context for where this nested table is in an insurance policy. It does really poorly with numbers, and it does really poorly when you've got lots of policies that are all talking about trailers getting rear-ended. And there's a ton of techniques to be able to get the signal from the noise in these chunks, but it requires a lot of processing.

    May Habib16:31

    And when that policy, even one policy, gets updated, you're throwing away the entire embedding store and starting over. Compare that to a graph-based approach where we've automatically created the nodes and edges that build the relationships and help us understand where these concepts land and how they're related to each other. So when we're going in and coming up with the relevant chunks—it's a different concept, but comparable—we're sending much more relevant data to the model, and so when you pair that with other techniques around the retrieval index and the Fusion-in-Decoder for limiting hallucinations, you just get much better accuracy out of the box.

    May Habib17:29

    And that's where it starts. And being able to really visualize literally the nodes and edges of the graph helps us help the customer understand where there's actually way too much density here and we've got to put in a light ontology or some light curation, or where is it pretty thin? Questions are coming in and there's just no information on this in the knowledge graph. Then the maintenance of it is just way, way cheaper. We were very validated by a Microsoft paper that came out earlier this year, and they called their approach GraphRAG.

    Guardrails

    17:59
    May Habib17:59

    Except they're using GPT-4 to build the graphs. And literally just adding that approach to our benchmark costs like $55,000 or something. So it's just not a very scalable approach versus our ability to build and train an LLM to do this and solve a lot of our problems with LLMs, frankly. Just very customer-friendly.

    Matt Turck18:06

    Continuing our product tour. So this was the knowledge graph. There is the guardrails.

    May Habib18:19

    So if you are somebody who wants to ensure that no PII comes into any prompt, so that is a governance and guardrail that you are putting on.

    Matt Turck18:20

    Okay, so how does that work?

    May Habib18:52

    Yeah, so there are a lot of different guardrails in Writer, and the approach is different depending on what the guardrail is. So if it is PII, you're literally going into Writer and you click a button under compliance settings that says no PII. It's that simple. And unlike Bedrock, this is not just like a simple regex filter, because that doesn't work. You really need, again, an LLM-based kind of understanding of what is PII. On brand, we're doing post-processing to introduce a rewrite, an LLM-based rewrite based on a brand's guidelines.

    May Habib19:30

    And that is also, think of it as a fine-tuned LLM that rewrites everything that comes out. And again, that's like a multi-select. I'm a pharma company and I have 21 brands, and I'm completely standardizing how I build call scripts for my field sales reps. So you're going in to see an oncologist, you literally have got this injection handling script for a drug. This is stuff that we work on. But it's got to sound like, in addition to all of the compliance around that, it's also got to sound like whatever the drug is.

    May Habib20:00

    And so coming out of Writer, you're able to put that as a guardrail to make sure that whatever is produced aligns with those guidelines. Specific types of questions you can't ask. And again, you need an LLM-based approach to this because regex simply won't work. You're not gonna be able to put in all of the topic areas without really annoying or causing all sorts of inference problems in the end consumer application by having all of these microservices that need to run in a regex-based way.

    AI Studio

    20:17
    Matt Turck20:22

    The AI Studio, is that where people can build their applications? Is that what it is?

    May Habib20:53

    Yes. And there, the insight really came first from—and this was starting not that long ago. I mean, this space has moved so, so fast—but starting about two and a half years ago, and we started in marketing, but for highly compliant industries, folks started to need just way more customization. Hey, Writer, the stuff coming out of your generative model, it's fine, but I really got to squint to think that it's even good. I'm not using it. And so for us to really get amazing required a ton of scaffolding around the LLM.

    May Habib21:24

    And prompting is not even the right word for it. We are really—let's say you are a Franklin Templeton or a Vanguard, and you're producing market commentaries based on fund fact sheets. We need to read the charts, we need to understand the graphics, we need to understand the commentary, and then compare cross-fund to be able to write the commentary on it. And so it started very early for us that the type of content that passed muster and got to production grade and really wowed people just required a ton of business logic around what you ask the model to consume and produce.

    May Habib22:09

    And this was L'Oréal about a year ago, literally like 180 apps, and we needed help, right? And you really want to, especially like the big enterprise, they want to learn these skills themselves, right? So it's almost like a build-operate-transfer on the most mission-critical stuff. And then for a lot of the use cases that we couldn't even predict that they would want, we have to be able to really democratize these skills. And so we built AI Studio as a way, and we're of course the first consumer of it, so it really accelerated our ability to build and ship for the customer too.

    May Habib22:44

    But it allowed them to come on board and help. And we've since built an AI Studio certification and academy, and that has been really fun, and we're just starting that process. And this next step-function change in what the models can do, and really being able to give every app in AI Studio tool skills, right, so that they can actually use the tools associated with the upstream and downstream tasks of the core application, is where AI Studio is going.

    Palmyra X 004

    23:16
    May Habib23:16

    And yeah, that is very exciting. It's not for everybody. We're not democratizing, and everybody's suddenly going to build their own tooling. But the sophisticated business users, we've really been able to get on board. And it's been very exciting for us and them to see the acceleration of that roadmap.

    Matt Turck23:30

    As we mentioned earlier, you also develop your own models. And the latest one, the latest LLM, is the Palmyra X-004. I read that you only spent about $700,000 to train it.

    May Habib23:32

    Just on GPUs, but yeah.

    Matt Turck23:46

    But directionally, that feels a lot lower than what you hear in the press for large models. Maybe walk us through how you thought about it, how you built it, and how you achieved that cost efficiency.

    May Habib24:11

    The composition of our family of models is the kind of core X-Series, where you've got really good reasoning, and then the domain-specific models that are cuts from the X but have domain-specific understanding. So, healthcare model, CS model, creative model, financial services model, et cetera. And for us, synthetic data has been our friend for a long time. Not from the very beginning, but we have seen the use of synthetic data have a lot of advantages that I think other folks are already starting to catch up to and copy.

    May Habib24:56

    And chief among them is efficiency in the training. And if you think about it just logically, right, the training data, the internet-scale training data that everybody uses, that's information that's created for human consumption. Think of synthetic data not as garbage gobbledygook, but as data that's constructed for the consumption of LLMs. And we are able to be so much more efficient on the size of these models. These are still large models, like 100 billion, 200 billion. These are not like SLMs or whatever.

    May Habib25:28

    LMs that folks are talking about. But the training is just much more efficient. We've developed techniques around early stopping so we don't overtrain the model. A lot of models are just simply overtrained. We have innovated on just the vanilla transformer architecture, and that is dozens of algorithmic improvements. And so part of it is, when you're GPU-poor—and we've never been GPU-poor poor—we've got great partners, and NVIDIA is a great customer and partner, and Jensen's awesome.

    May Habib25:58

    But we haven't been swimming in cash. No one ever gave us a $200 million cluster or things like that. So that has been great because the second huge benefit is that our customers, when they're paying for the inference, they don't have to be GPU-rich either. And so that has been really helpful. We're able to indemnify them. This is copyright-free information. We give folks—this is another reason we build our own models—a billion randomized tokens that they can query themselves to validate that we use copyright-free information and that our bias and toxicity distribution is what we say it is and all of this stuff.

    May Habib26:52

    So there are a lot of reasons, and I think even if in the future we offer third-party models in AI Studio, we've got a big corporate announcement happening next week and some partnerships we're announcing alongside that on the multimodal side, like, even if we do that, we will always have to be a state-of-the-art LLM company because we build LLMs to solve so many fundamental problems and challenges that enterprises face building mission-critical applications. This is AI that in three to five years is going to be able to run companies.

    Current state of the AI adoption in enterprises

    27:18
    May Habib27:18

    So stakes are high. And when companies are buying software, these questions are now very sophisticated: provenance of training data and how you filter and clean it, and what is the maintenance schedule, and all these things that, if you're not building your own models, you simply can't answer.

    Matt Turck27:28

    What is your sense of where we are in the adoption curve for enterprises? What is the positive? What is the negative?

    May Habib27:51

    I think we're very, very early, which is, of course, so exciting because we see so much opportunity already. Right now, we can go to the market of truly early adopters and talk tech all day long, right? Literally, the CEO knows what we're talking about when we say graph-based versus vector. And I'm not, like—this is Fortune 10. The CEO knows. The mainstream market, we're going to have to skew all of this stuff up as we go to the mainstream market.

    May Habib28:25

    And so I think that's very exciting for startups. I see a lot of startups go to the enterprise, and you can definitely get demos. I mean, literally, I was talking to a CIO yesterday. Fortune 50 had just personally taken a demo from an eight-person startup because people are so curious. They want to know, they need to be the smartest. But you can't ask the company to build the last 50% of your product. You have to—the stakes are really, really high because everybody else is there, right?

    Writer's sales approach

    28:57
    May Habib28:57

    The consultancies and the strategic advisory firms and every hyperscaler can literally throw more innovation dollars at the enterprise than you will raise in three years from venture capital. So stakes are very high. And I do think that we've been very fortunate that we've just been enterprise from day one and really got to kind of build in total obscurity with the enterprise as our research lab.

    Matt Turck29:16

    How do you sell currently, precisely given the spectrum of knowledge of AI within the enterprise and the fact that we are very early? Do you have a consultative approach where you sit down with the management and walk them through the possibilities? How does that work?

    May Habib29:47

    We have an amazing CRO, Andy, who the first few meetings we had together, he was like, "May, I don't know. I mean, it's just like Writer, Writer, Writer from the first second. What are you doing? We don't really sell enterprise software this way." And I'm like, "GenAI is different, trust me." And he gets it now, right? Because it is everything everywhere all at once, 40 vendor pitches a week, and you have to very quickly cut above the noise. And for us, that is around the sweet spot of the transparency and composability, security, power of DIY, with the consumer-grade application experience, inference, and just ease of use of something off the shelf or Copilot, et cetera.

    May Habib30:26

    And so when you can very quickly get to the heart of why every layer is so powerful and why together, right, it's the Tesla versus the General Motors, that cuts above the noise very quickly. And then there's a true lived and scaled experience across the very specific applications. Like, you look at one of our pitch decks for a pharma company, you would think it's a completely different business that goes and pitches to CPG or goes to pitch to retail.

    May Habib31:09

    Like, that's how specific we get, other than kind of the first one or two slides. But otherwise, it's a very product and solution sale. I think we are really good at bringing in partners to help with the change management. We've got an exciting announcement coming around a managed service around the Writer platform. So we want people—the whole point is that it's the easy button for generative AI once you get going. And we've got this very fast orchestration to value in the first 10 weeks where we just want this explosion of capability that shows up after two years of sort of like ho-hum results.

    What May Habib is excited about in AI

    31:25
    May Habib31:25

    And we were really able to pull that off.

    Matt Turck31:38

    What are you excited about, I guess, in AI in terms of where things are going, whether that's, I don't know, additional modalities or new models coming online? What do you find yourself naturally getting curious and excited about?

    May Habib32:10

    We got some fun, exciting news today from a major analyst that we're included in another big quadrant around agentic for this next quarter. And I don't like that word because there are millions of people who have that job. They're agents for all sorts of things, from insurance to sales, et cetera. We, for that functionality, are using autonomous action as the word. And I'm really excited for what's possible. And you look at a lot of the demos and what folks are trying to do today, and it really just amplifies the weaknesses of LLMs.

    May Habib32:48

    Our approach on autonomous action is really to start with the most knowledge-intensive, the most important nodes of a workflow. And then autonomous action can really do some incredible work in moving from system to system and really being able to reason in novel situations, to take action autonomously and proactively, with all of the right observability and controls, of course. But especially when we have seen our early customers—so, customers who are early on our autonomous action functionality but have been customers for two, three years—just get to square 10 because they've gotten to square one.

    May Habib33:13

    It's so, so exciting. And so I'm most excited about that. I do think 2025 is going to be about AI workloads.

    Autonomous Action use cases

    33:14
    Matt Turck33:22

    What kind of use cases or problems do you think lend themselves best to autonomous actions?

    May Habib33:55

    We've kind of worked backwards from the use cases that we think lend themselves well to that. And we want to be responsible for hundreds of millions of dollars of value and realized value inside of a company. And we're almost there at a couple of our largest customers. And what we really want to be able to do is help customers get new products developed faster, those products get into market faster, and customers serviced better and faster. And none of this is about productivity, right?

    May Habib34:27

    You being able to go through three decades of consumer research on Listerine to be able to develop a new flavor. You being able to develop all of the go-to-market around that new product in literally a third of the time for a third of the cost. None of that is about 10%, 15%, 20% productivity. It's about just 300%, 400%, 500% kind of change in capacity and capability. And for those types of workflows, the nodes, the business logic, and the data of the nodes of those workflows being generative AI-native and reinvented with generative AI, you're really able to then piece together, with agents in between, real acceleration between and around systems.

    May Habib35:21

    So think of kind of the next generation of Writer apps as almost super apps that can both use Writer autonomously on behalf of the customer, things that they have been more declarative about from a business logic perspective, as well as the systems that are both upstream and downstream of those nodes on behalf of a user executing the work. So we're taking LLMs from doing tasks to orchestrating work, and that is really exciting.

    Matt Turck35:23

    Wonderful. May, thank you so much for joining us.

    May Habib35:24

    Thank you, Matt.

    Matt Turck35:45

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.