MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Empowering Millions of Creators with AI Video Editing | Gaurav Misra, CEO, Captions

    Gaurav Misra is the CEO at Captions. We cover why Captions targets people who do not want to learn professional editing, how it uses AI to generate and edit talking-head footage, and why the company breaks videos into overlapping one-second segments to cut generation times to two or three minutes.

    02/20/2025

    Hosted by Matt Turck · with Gaurav Misra, CEO, Captions

    AI videovideo editingAI avatarscreator toolsgenerative AI
    Listen now
    YouTubeApple PodcastsSpotify
    1h 13m · 23 chapters
    Contents

    Transcript

    What is Captions?

    1:30
    Matt Turck1:28

    Hey, Gaurav, welcome.

    Gaurav Misra1:30

    Thank you. Yeah, thanks for having me.

    Matt Turck1:35

    I like to start those conversations from the very top. What is Captions?

    Gaurav Misra1:56

    Yeah, I mean, so in short, Captions is a production studio in your pocket. So it's an app, a product that basically lets anybody make video as simple as possible, right? So think about all the process that you have to go through to create a video, right, as an individual creator. So you have to, like, record the video, and obviously you're gonna get something wrong, so you do retakes and, like, get that word right and mispronounce something. Sometimes you notice something after the fact. It's like, oh my God, that's so weird, right?

    Gaurav Misra2:15

    You wanna redo it. So just the recording is hard enough. There's the editing afterwards, where now you have to go figure out what's a keyframe. And in order to get a good output, you need to know what a keyframe is, right? There's a huge learning curve to these things, and we help people skip past this.

    Matt Turck2:15

    Yeah.

    Gaurav Misra2:25

    So an easy way to—a framework to think about how and when to use Captions is: if you want to do video recording and video editing, you should use something else. If you don't want to do it, you should use Captions.

    Matt Turck2:44

    It's a pretty magical experience, right? So you have your video, you import it, you pick a style, right? And then you basically have, like, a magic button, and on the other side of it comes out a video which is edited, where you can automatically cut the hesitation sounds and pauses.

    Gaurav Misra3:06

    Exactly. Yep. So not just the recording part, you can say a passage, we can generate a video of you yourself or anybody you have the license for. So if you have a license for somebody else's likeness, or we also have actors sort of available in the app that you can pick from, we can generate a video of that person saying whatever your pitch is. Maybe you're pitching a product, maybe it's, like, an ad, maybe it's a social media post, right? We can generate that, and you can get it perfect.

    Gaurav Misra3:20

    You can say exactly what you wanna say, test different messages, whatever it might be, and then we can edit it also, right? So that it—because just footage is footage. Footage isn't exactly usable most of the time as it is.

    Matt Turck3:22

    And when you say actors, are those avatars?

    Gaurav Misra3:29

    Exactly, yeah. I mean, so these are basically real people who have licensed their likeness to be used. Okay.

    Matt Turck3:34

    So it's either you, or you can just completely create the video just typing.

    Gaurav Misra3:36

    Exactly. Okay.

    Matt Turck3:40

    All right. Very, very cool. And tell us about the company. When did you start?

    How did Captions start?

    3:43
    Gaurav Misra3:54

    Yeah. So at this point, we started four years ago. So it was actually February 19th. Yeah, that's right. Or January 19th, sorry. It was January 19th, four years ago, just past that. It was my co-founder's birthday, actually, when we founded the company.

    Matt Turck3:56

    And what was that story like? How did you meet your co-founder?

    Gaurav Misra4:04

    My co-founder and I met about 10-plus years ago. We were at this company called Localytics, which was a mobile—

    Matt Turck4:05

    Right, right. In Boston, right?

    Gaurav Misra4:18

    Exactly. Yep, exactly. It was a mobile analytics startup, and they were one of the first to be doing that in that space. There were many companies that came around. I remember companies like Appboy, which became Braze. Yeah, that was Braze, right, right.

    Matt Turck4:20

    That was a super hot space in the early 2010s.

    Gaurav Misra4:39

    Exactly right. And that's where we met, and we only overlapped for maybe three or four months, actually, which is crazy, by the way. And we kept in touch for almost 10 years after that because Dwight, that's his name, moved to New York after that three-month overlap. He left Localytics, he moved to New York. I ended up moving to New York later to join Snapchat.

    Matt Turck4:42

    But what did you do at Snapchat?

    Gaurav Misra4:43

    That was 2016.

    Matt Turck4:44

    Yeah. And what did you do there?

    Gaurav Misra5:01

    I joined originally as a machine learning engineer. I transitioned to product design, which is a rare transition, I think. But it actually really helped me in my path to starting this company. Because think about it: product design plus machine learning, hard to beat as a combination.

    Matt Turck5:06

    Yes. Yes. And then after your stint at Snap, how did that come about?

    Gaurav Misra5:13

    Yeah. So, I mean, we had actually been meeting up in New York. We would meet at this place called Rye House. It was on 18th Street, actually, not too far from here.

    Matt Turck5:13

    Right.

    Gaurav Misra5:31

    We would meet every six months or so and discuss startup ideas and what we wanted to do. And we both knew we wanted to start a company, but I think we were both of the opinion that, oh, it's not the right time, because we were doing well in our careers and stuff, and things were generally going well in the world, the economy, and the companies we were at. It never felt like, oh my God, I can quit my job right now and start this thing.

    Gaurav Misra5:48

    And there was always this lingering, like, do I know enough? Clearly, I'm not experienced enough to start a company, right? Obviously, that's completely false, but it feels like that, right? And so we were both in that mindset until—and by the way, this went on for like 10 years, right?

    Matt Turck5:49

    Mm-hmm.

    Gaurav Misra6:09

    So until 2021, basically, right? So just post-pandemic, what happened was things were very inflated, just heavily. Even the companies we were at, right? I mean, Dwight was at Goldman, so I don't know what to say about that, but I was at Snap, and Snap was worth $100 billion, right? Which is five times more than that on the current valuation, right?

    Matt Turck6:10

    Yep. Absurd times.

    Gaurav Misra6:31

    Yes, exactly right. And so it felt to me like this is unsustainable, right? It just didn't feel justified. And sure, we were getting paid a bunch of money and stuff because the valuations were so high, but it felt like this is going to dip, and then we're going to have a 10-year road to recovery. And I didn't want to be a part of that journey. And I was like, this is the perfect time to start a company, actually.

    Matt Turck6:40

    Smart. And so that was the two of you. How did you guys get started? So, 2021, clearly pre-ChatGPT moment.

    Gaurav Misra6:41

    Right.

    Matt Turck6:47

    And how did you get started, including on the tech and product front?

    Gaurav Misra7:11

    We actually started with the idea of revolutionizing video. I think the long-term plan was AI plus video, but we knew that the AI part would take like 10 years, or at least that's what we thought at that time, right? But the main driver and our framework for evaluating the initial ideas especially—and I think there's a unique framework for figuring out truly groundbreaking and generational companies. Like, how do you build those companies? To start a company is much easier, I think, than to start a truly groundbreaking company, right?

    Gaurav Misra7:44

    And a lot of luck is involved, a lot of hard work is involved, of course, all the classic obvious ingredients. But our framework was actually to look at society and to think about what are the inevitable changes that are happening in society, right? And these could be things like 5G networks are available, right? That changes how people can now upload and download video on their phones, right? Or maybe it's something behavioral in society that's changing, right? And there's a lot of behavioral changes actually happening in society right now, right?

    Gaurav Misra8:15

    Which are inevitable, unstoppable, right? They will complete over a 10-, 15-year period. A lot of times it's driven by technology. A lot of times it's driven by just social behavior of the masses, right? And if you identify one of these, you can see a world 15 years later that's going to be post-change and pre-change, right? And our framework was to think about what are the companies that are going to be needed and thriving in this post-change world? And can we identify one of those companies and start building it today.

    The strategy behind launching Captions

    8:25
    Matt Turck8:31

    It's a very strategic way of thinking about what company to found. You have different heuristics for starting a company, but it sounds like you guys were very strategic and deliberate in terms of macro tailwinds.

    Gaurav Misra8:46

    Totally right. Especially because the normal advice is, like, do what you enjoy doing, right? But I think we both enjoyed doing a lot of different things, or solve a problem for yourself. And I think we thought about those things, but we didn't really have a lot of problems, to be honest.

    Matt Turck8:52

    Fantastic.

    Gaurav Misra9:16

    So it really came down to solving problems for other people. And that's really what's more exciting to me. The most exciting thing to me is to solve a problem for somebody else and help someone, right? So with that in mind, we came up with this approach of looking at macro trends, right, and trying to figure out where we can have a change. So, I mean, the biggest macro trend at that time was video. Video was the fastest-growing medium of consumption.

    Gaurav Misra9:43

    It still is, right? Continues to be. And it was essentially like, I mean, TikTok came out in 2019, right? In the U.S., at least. It was crushing basically every other type of social media. And I was at Snap, so I saw some of this firsthand. I actually was running the team that was building the response to TikTok internally at Snapchat. And by the way, TikTok spread in the U.S. mostly through Snapchat. And this is not well known out there today, but a lot of it was they were spending a very, very large amount of money on advertising on Snap, right?

    Gaurav Misra9:58

    And there were discussions internally often about, like, should we be selling this advertising to TikTok, right?

    Matt Turck10:00

    That's a wild story. I never heard that.

    Gaurav Misra10:20

    Yeah. And we ran A/B tests, actually, within the advertising groups to see, like, let's segment some users out, not show them TikTok ads, see if they behave differently in the long run, right, in terms of engagement on Snapchat and stuff like that. And there was no difference. And so it was decided that, well, we should just sell them the ads, right? It's making no difference. It was very wrong because it had a big difference, but the difference just took a long time to be realized.

    Matt Turck10:26

    Yeah.

    Gaurav Misra10:51

    And it actually, like, yeah, it affected some of the core markets. Snapchat has very engaged markets in the Middle East and in parts of Europe and stuff, especially Norway and a bunch of other countries where almost all age groups are engaged in Snapchat. And it even depressed those markets. So video was eating all of this stuff, right? And we knew that that means that creating video is going to be a huge industry in the future, right?

    Matt Turck10:52

    Mm-hmm.

    Gaurav Misra11:13

    And bringing all these tools that are only in production studios' hands, essentially, right? Like, you have to spend $10 million to make an artificial shot of something, right, or a movie or a small video or whatever, an ad, right? What if that was $10 now, right? And everybody had access to it. It almost felt a little bit like coding, like what software engineering is to civil engineering, for example. Civil engineering, you can imagine all you want, but you can't actually execute anything because you need a lot of resources to do it, right?

    Gaurav Misra11:31

    Software engineering, someone sitting in their basement can make something that disrupts an entire massive multibillion-dollar company, right? So it's that energy brought to video, right?

    Matt Turck11:45

    Fascinating. And fast forward to today, to give people a sense of scale, I've read somewhere that you've now worked with or empowered over 10 million creators around the world. Is that an accurate number?

    Gaurav Misra11:59

    Yeah, and growing. And that was a little bit of an older number too, but the one that we've announced so far. And it's paying customers in over 180 countries, basically everything except the obvious ones, and growth in all the different countries as well.

    Matt Turck12:08

    And you've raised over $100 million in venture capital. The team is 70 people, you were saying before we started recording?

    Gaurav Misra12:09

    70 people, yep.

    Matt Turck12:11

    Okay, very good. All based in New York?

    Gaurav Misra12:12

    All based in New York, fully in person.

    How is Captions different from other editing tools?

    12:32
    Matt Turck12:38

    In person. Okay, really interesting. Okay, cool. So let's start diving into the meat of the product itself, maybe starting with the positioning of the product. So clearly, one obvious question is: there are creative tools in the world. There's Adobe, Adobe Premiere Pro, like all the things. So how do you differentiate, and how do you think about those? Is it just a completely different target?

    Gaurav Misra13:00

    I think it is a completely different target because I think a lot of, if you look at Premiere, right, some of that level of professional tool, right, these are truly designed for the people who are making careers out of it, right? Like, they are literally working on this. That's all they do, right? And their skill set they've built over maybe decades working with these pieces of software, right? And I think that's what makes it actually hard to innovate in these spaces, is because people have so much built-in muscle memory about how things have been done that if you come in and change a bunch of stuff and be like, "Here's a simpler video editor," people actually don't like it, right?

    Gaurav Misra13:21

    They're like, "Well, why?" Right? Like, I already know how to use this. You messed things up, right?

    Matt Turck13:21

    Yeah.

    Gaurav Misra13:37

    You moved everything around. It's not where I need it to be, right? I think that's where a lot of this fails often, is you go and sort of sell it to the person who's a professional. But I think really for us, the goal from the beginning was not to sell it to the professional. It was actually to enable a new set of people who actually wouldn't be able to, or wouldn't have the time or energy to invest into learning a professional tool, to let them actually now be able to dabble in this space, right?

    Gaurav Misra13:51

    Whether it be editing—

    Matt Turck14:00

    So it's almost like software engineers, to use some of the same route of the analogy that you were using. It's a little bit like software developers on the one hand versus no-code kind of people.

    Gaurav Misra14:04

    Totally, yeah. Very true. Very, yeah, very spot on.

    How does it compare to CapCut?

    14:13
    Matt Turck14:18

    And for that group of people, there are also solutions on the market. We were talking about TikTok. So famously, they have CapCut as well. So how do you think about that set of companies?

    Gaurav Misra14:39

    Totally. So I think there was actually a lot of where those companies came from. Like CapCut specifically, if you look at it, right? I mean, CapCut is—the main goal of CapCut is to grow TikTok, and it's a virtuous cycle there. That's kind of why they have it. But they actually came up with their own set of goals because it is a separate company under ByteDance, right? There's TikTok, which is a different company than CapCut. As different companies, they have different goals, different CEOs, et cetera.

    Gaurav Misra15:04

    CapCut's goal actually became inspired by Canva, actually. And they're aspiring to be like a Canva, maybe a Canva for video. I think this idea has been floated around a lot because people looked at the last decade and they saw basically Figma and Canva as two sort of major design companies that took off. Figma went after the professional, and Canva went after the non-professional, basically, right? And it's actually a very similar story to what I'm telling, right, in a way, because Canva is explicitly used by people who do not design, right?

    Gaurav Misra15:31

    In fact, designers may hate Canva. They might think, "It's not good. I could easily make something better," right? I could easily see a lot of designers saying that, right? But for the person who doesn't know how to design, it's amazing, right? And the main reason is it wipes away the blank canvas. There's no blank canvas on Canva, right? Which is the hardest part of designing anything.

    Matt Turck15:32

    Mm-hmm.

    Gaurav Misra15:49

    You open this thing, you see a blank canvas, and there's like six buttons to make squares and circles, and now you have to come up with something that looks good, right? Really hard. No matter how simple you make the UI, by the way, right? It doesn't matter how simple the UI is. I mean, design, there's only six buttons, right? There's nothing more than that, right? It's not complicated. It's just creativity that has to be applied on top of that, right?

    Gaurav Misra16:05

    And there's some level of skill that needs to be built up of how do you convert these six buttons into UI that looks good, or into design that looks good, right? So that's what Canva solved, and they solved it with templates. You just apply a template and you get something nice, and then you have some opinions. You're like, "Oh, maybe this orange could be like red," or, "Here, I'm going to change a font here," or, "Maybe the text I'm going to update," and you've made something, right?

    Gaurav Misra16:31

    It's almost like Easy-Bake, right? Like, just add an egg and you're done. And I think that's what they did, and did really well. But with that analogy, I think obviously VCs and a lot of founders immediately were like, "Wait, Canva for video, right? Like, of course it's gonna be a thing." But it actually didn't become a thing for a long time because it wasn't analogous in the same way, right? People were making simple UI for video editing, but simple UI doesn't matter, basically.

    Gaurav Misra16:39

    You still get a blank canvas, and now you have to figure out, like, how do I make a good video?

    Matt Turck16:39

    Mm-hmm.

    Gaurav Misra17:05

    Right? And that's the hard part. And that wasn't solved, right? The interesting part that we figured out is, even with CapCut and with many of these other tools, they tried to apply templates, but the problem is that video is just very different than design. Like, you actually can't templatize video because the nature of it is completely unknown, right? It could be saying anything, it could be doing anything, anything could be happening. The template just can't fit with—it can't understand what's happening in the video until now, obviously, right?

    Gaurav Misra17:24

    And I think that was the big change. The reason that AI video editing can even exist is because AI has transformed so much, right? And that actually is the key that unlocks suddenly what templates had unlocked in the design world, right?

    Matt Turck17:25

    Mm-hmm.

    Gaurav Misra17:50

    Is suddenly machines can understand exactly what's happening in the video. They can know who's there, saying what to whom, when, why, how, what's happening, when cuts should be made, when B-roll should be inserted, what types of images would go well with the background over here. Like, maybe we want to have a purple-themed image, right? So all that can be understood and done automatically. The footage can be completely generated, so you can get the perfect footage, right, that says the things at the right time.

    Gaurav Misra18:14

    All the things that actually Canva enabled for design can be enabled for video now, right? And I think that is the big difference. Whereas I think CapCut's approach is much more traditional. They're actually just trying to do exactly what Canva did, which is, okay, let's make the UI simpler and add a few more buttons and essentially bootstrap it to TikTok and distribute it like crazy.

    Who is the typical Captions user?

    18:22
    Matt Turck18:28

    So the people that you empower with those tools are not designers, are not video editors. So that's your TikTok creator crowd, Reels. Is that—

    Gaurav Misra18:56

    A lot of creators, small businesses. So those are, I would say, the two largest sections. I would say small businesses is probably the biggest subsection where I think what you end up realizing is, like, social media and marketing are just hand in hand now, including ads, right? Like, when you run ads, it's just social media, but you're paying for the reach, right? And that's a massive industry with a lot of footprint, especially among small businesses. So e-commerce, think individual businesses, like, I would say, personal trainers, for example, real estate agents, things like that where you need to make video as part of your job for some reason or the other, and you don't want to learn all of the skill set, right?

    Matt Turck19:20

    Yeah, and that can be a little bit of a tough crowd, right? They're very demanding, and it can be fickle, though. Like, what's—yes or no?

    Gaurav Misra19:29

    I mean, maybe a little bit here and there, but nothing more than normal. I think one thing that we did well sort of in the first couple of years of the company is we actually kept the application paid only. And that actually—

    Matt Turck19:31

    That filtered out the—

    Gaurav Misra19:46

    It just filtered out anybody who's not, right? Exactly, right. So it really left us with the people who are willing to pay and who have a strong enough use case, basically, right? So everything they asked for was very reasonable. And so it was a clear way to build a roadmap and filter out the noise.

    Matt Turck20:03

    That's really interesting, right? So you chose to monetize first, as you were saying, not necessarily to monetize, but to filter out people, which is sort of the counter—the other way around from what you see a lot of people do, which is to open the floodgates and do it free. Totally.

    Gaurav Misra20:04

    Yep, totally.

    Matt Turck20:07

    All right. So on the product, so you started with captions.

    Gaurav Misra20:07

    Of course.

    Why ‘Captions’?

    20:13
    Matt Turck20:13

    Which was, aptly enough, the name of the company. So why that at the beginning?

    Gaurav Misra20:38

    Initially, we knew that we wanted to do video creation with AI. Now, at that time, this is pre-GPT, pre-transformers. I mean, transformers were there, but no one cared about them at this point. Essentially, we looked at all the different ways that we could apply AI to video, and we wanted something robust enough that it actually works reliably and works really well. And one of the things that was at that point was transcription, right? Speech-to-text. But interestingly, what we found at that point is, even though it was actually a part of daily life in many ways, Siri and Alexa existed at that time.

    Gaurav Misra21:09

    And these were consumer products. They were out there, and people understood that speech can be understood by machines. But I think people were surprisingly shocked. I think outside of tech circles, the everyday person was kind of surprised to see how good it was in reality. They could understand all this obscure terminology and really perfectly, word for word, get something as it was said.

    Matt Turck21:12

    So did you use that at the time? Was that Whisper already?

    Gaurav Misra21:22

    Well, we used something else pre-Whisper. I don't remember what it was originally, but we were one of the first companies to implement Whisper when it came out. That was—

    Matt Turck21:28

    Whisper being the open-source speech-to-text model developed and launched by OpenAI.

    Gaurav Misra21:45

    Yeah, exactly. So that's when the OpenAI hype started, basically. We did switch over to Whisper originally, but initially we had some model. I think it was some other model by Google. So the original app was made in a weekend, literally in two days, and launched. It wasn't marketed at all. It was just put on the App Store, and the next day I woke up and I just saw all the numbers blowing up completely, and we're, like, top 20 or 30 on the App Store.

    Gaurav Misra22:13

    And I just called Dwight, and I'm like, "Hey, looks like a lot of people are using the app, like 600 videos per minute being made right now." It was like, "Oh, nice." Because I think we just didn't expect it to just—I mean, we had literally just started the company two days ago. We had just incorporated.

    Matt Turck22:18

    That is wild. Yeah. I mean, that's every founder's dream, which actually never happens.

    Gaurav Misra22:21

    Yeah. It was almost—

    Matt Turck22:25

    And what happened? Like, somebody posted about it pretty quickly on social overnight?

    Gaurav Misra22:44

    I mean, so we had done a little bit of research when we chose this. So we were browsing TikTok. And by the way, I think if it doesn't go away, TikTok is a great way to discover what is trending. And I think at the end of the day, finding a good idea to start with very tactically, not in a strategic way but a very tactical way, ends up just being finding something that can go viral, right? That's really all it is.

    Gaurav Misra23:04

    And that means that it resonates with people, right? And we did a lot of searching on TikTok. We were spending 10 hours a day on TikTok just browsing videos. And one of the things that we saw is there was a huge accessibility movement. People cared about reaching all audiences and people who are hard of hearing and all those things. And so people were spending a lot of time literally manually typing out text, right?

    Gaurav Misra23:29

    And then on the other side, we saw sort of more in the Asian market, especially in Japan, they always put these big, big words. This is also from TV over there. They'll always put these really big words that show up in the middle of TV, and it's funny and just entertaining. And we're like, wait, that plus transcription in the U.S., maybe that's the combination. And we can do all of it automatically. APIs exist, like you can build this in less than two days, basically.

    Gaurav Misra23:39

    So we built it and put it out there, and it just instantly took off. I think someone found it and posted it on TikTok.

    Captions’ product suite for production and editing

    23:47
    Matt Turck24:01

    Amazing. What an amazing story. So, fast forward to today, you've got a whole suite of products. I was just looking at some notes as I was prepping for this. So you've got AI Edit, AI Creator, AI Ads, AI Twin, Lipdub. Maybe the quick version of what those various components do.

    Gaurav Misra24:23

    Yeah, so the short version is that we have two core products, actually: AI Edit and AI Creator. So AI Creator mirrors the camera, right? So it is actually the AI version of the camera. So as AI Creator develops and becomes perfect, you will not need to use the camera anymore, right? You'll just be able to generate all the footage that you want. Today, it does talking videos, right? So videos where one of the actors that we have, or you yourself, are saying something.

    Gaurav Misra24:49

    AI Twin is just a specialized version of AI Creator where it's yourself, basically. Now, the other is AI Edit, and this mirrors the editing, right? So you can edit it yourself, or AI Edit can do it for you, right? And that's totally up to the user, how they want to do it. Or they can wait for AI Edit to do it, and then they can make some changes. So combined, this ends up in an experience where you can go from a quick description of what you want as a video to a completely usable output that you can go and just post or use in whatever context you were hoping for.

    Matt Turck25:17

    And from a product-building perspective, those sound like two fairly different kinds of problems. First of all, is that right? And second, maybe walk us through some of the unique challenges of each.

    Gaurav Misra25:42

    Yeah, I mean, there are definitely slightly different types of problems because one is sort of a classic video generation problem, right? Most of it comes down to just the model itself, right, and how the model is designed to generate the outputs that we need. And then, obviously, the UI around it to make it easy to use. Now, on the editing side, though, it's a different type of model, right? And same idea, right? It's like a large model with outputs that essentially act more in an agent sort of behavior as opposed to a media generation behavior.

    Gaurav Misra26:12

    But it's the same flow of data that's training that, where you kind of have to go end to end, right? In fact, the real journey actually starts at the script, right? You write the script, then you record the video, and then you edit the video. The script generation part we don't take on because we think the LLMs and stuff will solve this eventually. I think they're still not very good at video scripts, and they're getting better and better, but we do think that as these models improve, this will be a solved problem.

    Gaurav Misra26:29

    We focus on the ones that we are well positioned to take on, right? Which is the video generation part, so recording and the editing.

    AI models powering Captions

    26:37
    Matt Turck26:44

    You alluded to some of the models, so let's get into that. What do you run on in terms of core models? What is third-party? What is maybe open source? What is proprietary?

    Gaurav Misra27:06

    Yeah, so essentially, internally, what we work on is just video generation, right? That's kind of where our core area is, like what we can really excel at because we have the data for it, we have the talent for it, and we have the unique ability to build that specific type of model. Right now, for everything else, we use some sort of provider, and we've tried basically all of them at this point. So we've been able to figure out which ones work best, but it's an evolving space.

    Gaurav Misra27:26

    And so every two or three months, we have to retest everything to just understand where everybody is. There's often winners that come up over and over again that we see over and over again, like ElevenLabs, for example, comes up over and over again in the audio space as—

    Matt Turck27:27

    And that's who you use for—

    Gaurav Misra27:30

    We have been using them for a long time. Okay.

    Matt Turck27:41

    What other providers, or anything that could be helpful to people who think about building AI products and tips? I don't know, what's worked, what's not worked as well.

    Gaurav Misra27:59

    Right. I mean, so we were on OpenAI for a while. They had a lot of downtime and other issues that were happening. I think that was an intermittent problem. Maybe it's solved now. So we switched over actually to Microsoft OpenAI for a while. And that actually did a lot better.

    Matt Turck28:01

    And that's OpenAI for what part?

    Gaurav Misra28:28

    So this is for script generation. Actually, it's like pre-video, basically. But also within the video here and there, when we need to figure out certain editing things, we might use one of these models. But for the most part, Microsoft OpenAI did much better in terms of reliability. It just worked perfectly for latency and never went down, which is what you want. It's great. And then, but we ended up switching to Claude. So now we're on Claude.

    Gaurav Misra28:31

    And that was just because of—

    Matt Turck28:32

    Which Claude?

    Gaurav Misra28:32

    3.5.

    Matt Turck28:34

    Okay.

    Gaurav Misra28:34

    Yep.

    Matt Turck28:36

    3.5 Sonnet.

    Gaurav Misra28:37

    Yeah, exactly.

    Matt Turck28:43

    Okay. And how did you make the decision? Why did you switch? What are the pros and cons?

    Gaurav Misra29:07

    Yeah, so we actually run an evaluation every few weeks to a month, basically, across all the different models. And we actually have a very unique way of doing this, and probably an evaluation that almost no other company can do, not just on the LLM side, but also on the audio gen side and other things like music and sound effects. And I mean, all these things come together to make a video, right? At the end of the day, and this is an important thing to think about, is if you look at all the media generation foundation models, right?

    Gaurav Misra29:39

    Like, there's music generation and image generation and even stock video generation, and there's voiceovers and, you name it, sound effects, right? Like, where do you think this is all going? Like, what do you do with a sound effect, right? Like, you generated a sound effect, now what? Like, you don't share a sound effect with a friend, you don't post a sound effect on social media.

    Matt Turck29:42

    So you created your own evaluation benchmarks for—

    Gaurav Misra30:05

    Yeah, you put it in a video, right? That's where it's all going. Everything is going in videos, right? All the stuff is going into videos, getting posted on social media. So we can actually evaluate when the user in their workflow is like, "I generated this voiceover and I don't like it, delete," right? And to a different user, we can show a different model, right? And they might be like, "Yep, I do like it," right? And so we can compare from the user's perspective, perceived when they're not being paid for this, they're paying us for this, right?

    Matt Turck30:12

    Yeah.

    Gaurav Misra30:26

    They have all the incentive to pick the best one for their use case, right? We can tell you exactly which one performs the best, right? And exactly which ones users prefer when they're actually paying for it and it's right in front of you, right?

    Matt Turck30:28

    So it's a form of RLHF?

    Gaurav Misra30:30

    It is, yeah. It is exactly right. Yep.

    Matt Turck30:33

    So then you got that flywheel.

    Gaurav Misra30:39

    Yep. So this helps us figure out which ones are the best, basically. And we always surface the best models for the user.

    Matt Turck30:47

    Can you infer from if a person makes one choice that they're likely to make the next choice, and then you make the recommendation based on that?

    Gaurav Misra30:51

    Is that something that's regular? We can. This is all possible in the future, but not something we're doing right now.

    Matt Turck31:00

    All right. So we talked about third-party providers, OpenAI and Claude 3.5. So it seems like there's something in the air.

    Gaurav Misra31:18

    Yeah, they're doing something right. And I'm curious to see how some of these new models like o3 and stuff perform.

    Matt Turck31:28

    Yeah, it's very impressive. We're recording this right after the release of, I guess, o3-mini and Deep Research last night.

    Gaurav Misra31:30

    Which is super cool.

    Matt Turck31:33

    Yes, and the pace of release continues to be—

    Gaurav Misra31:35

    And DeepSeek and all that, right?

    Matt Turck31:44

    Yes, yes, yes. So those are the third-party providers. And then, so you're saying you're building your own models as well for video generation?

    Gaurav Misra31:45

    Right. Yep.

    Matt Turck31:48

    So that's the avatar, that's the edits, or both?

    Gaurav Misra31:58

    It's both. So actually, I didn't mention one thing, which is, on the third-party providers, we also have image generation in the application, right? So people can include images in their video.

    Matt Turck31:59

    For the background?

    Gaurav Misra32:05

    Exactly. Yep. So we are using Ideogram for that one because it generates really good text.

    Matt Turck32:09

    And, sorry, I keep interrupting you. Your models.

    Gaurav Misra32:32

    So the two models that we sort of work on ourselves are video generation and video editing, basically. And on the video generation side, we have a very unique view, is that, I mean, if you look at all the video generation models today that are doing text-to-video, they're all silent videos, right? There's never anybody talking. It's all B-roll, right? It's all essentially shots of New York, right? Shots of the beach, right?

    Matt Turck32:34

    B-roll is the background, right? Just to get the lingo.

    Gaurav Misra32:35

    Exactly.

    Matt Turck32:37

    B-roll is the background and A-roll is—

    Gaurav Misra32:56

    Is primary footage, basically, like what actually tells the story. So to give you an example, imagine that we were filming a movie, right? And when the movie opens, right, it might cut to like three shots of New York City from different angles, like a taxi cab going on the street, right? Like New York here, busy things happening, and then it enters this room, right? Boom. That was B-roll. A-roll begins here, right? And then the dialogue begins, right?

    Matt Turck32:58

    Mm-hmm.

    Gaurav Misra33:21

    And that's just to set the tone of where we are, who's involved, right? And then you get into the real footage that's actually telling the story, right? Which is always like dialogues and monologues and people talking, right? That's what storytelling is at the end of the day, right? And sure, you can make a silent movie, but do you want to? I mean, is that what all movies should be? So I think that's kind of what we realized, and we really focused in on talking videos, is what we call it internally, as what we want to focus on.

    Matt Turck33:35

    And talking videos meaning avatar generation or anything?

    Gaurav Misra33:36

    Yeah, I mean, anything.

    Matt Turck33:44

    So both the editing use case and the talking-head avatar use case are powered by the same model?

    Gaurav Misra33:45

    They're different models.

    Matt Turck33:46

    They're different models?

    Gaurav Misra33:46

    Yes.

    Matt Turck33:47

    Okay.

    Gaurav Misra33:55

    Yep. But for the video generation, it's just diffusion model, but it's conditioned on audio.

    Matt Turck33:56

    The diffusion models?

    Gaurav Misra33:56

    Yeah.

    Matt Turck33:59

    Okay. Based on open source or your own?

    Gaurav Misra34:02

    Based on stuff we're sort of building in-house, essentially.

    Matt Turck34:06

    Okay. And what data do you use?

    Gaurav Misra34:31

    Yeah, so all proprietary data. The data that we're training on internally is all fully licensed, completely proprietary. Obviously, part of it is data we get from operation of the app itself, but we also purchase data. And in the future, we're also going to be generating custom data, as we started to realize that there are certain niches and certain areas where you just can't get data very easily.

    Matt Turck34:41

    Mm-hmm. Is it really hard to get video data these days? Like, you still can't crawl YouTube, right? Only Google can do that.

    Gaurav Misra34:52

    Yeah. I mean, you can do anything you want, but should you is the question. I do think if the goal is to sell to enterprise, it's really hard to justify it if you're crawling YouTube, and how do you—

    Matt Turck34:55

    It's a fundamental flaw in your legality.

    Gaurav Misra34:56

    Yes, exactly.

    Matt Turck35:05

    Okay. And what is working very well and not working just perfectly just yet?

    Gaurav Misra35:27

    So what's working well is what you would expect, which is that you can put more and more data into these models, increase the number of parameters, and they will get better, basically, right? Of course, save any bugs or anything else that might be happening. Now, what's difficult, as always, is to solve—because we are in a unique space from that sense, right? We're doing this A-roll video. And so there are unique challenges in that. Like, audio conditioning is just not a thing that really has been solved at scale, and nobody else has done it, right?

    Gaurav Misra35:51

    And so that means that we have to be the first ones to run into these problems, try a lot of different things until we figure out what is the best way to solve them, right? And then scale that up, right? Which is also all very expensive, by the way, right?

    Matt Turck35:52

    Mm-hmm.

    Gaurav Misra36:10

    So there's only so many mistakes you're allowed to make, right? Because every mistake has a price. So running into those types of problems is where the challenges are. But I think, to be honest, that is also where the fun is, really, right? Like, it's like saying we're climbing Mount Everest, but it was all really easy, right?

    Matt Turck36:15

    Yeah. Was lip-syncing, for example, a particularly challenging area?

    AI lipsync

    36:22
    Gaurav Misra36:38

    It was, yeah. I mean, I think it wasn't obvious which encoder to use, for example, for audio, right? And there's a lot of pros and cons of using a lot of different ones. Like, you could use ones that are derived from text-to-speech, for example, or sorry, speech-to-text, right? So the Whisper encoder comes to mind, right, and things of that nature. And what you have to balance in there is obviously if it's encoding for transcription, it might lose all the emotion.

    Matt Turck36:46

    Mm-hmm.

    Gaurav Misra37:07

    Right? It doesn't need that information, right? So it could be discarding all of that for you, right? It could be losing timing information. I mean, Whisper particularly, like, it doesn't output timestamps that well, and it's kind of like okay at timing information. Like, it does 30-second chunks and stuff. It could be losing a lot of timing information. So if you lose that, then the lips will definitely not be in sync, right? So these are the types of things that you have to list through and figure out which one could work and which one doesn't work.

    Gaurav Misra37:15

    In different situations, right?

    Matt Turck37:22

    And for the captions themselves, so you're well known for the quality of the captions.

    Gaurav Misra37:22

    Yes.

    Matt Turck37:38

    How much of this is Whisper, and how much of this is stuff that you had to innovate on? Especially, I would assume if you have a bunch of creators recording, there might be, like, a noisy environment or that kind of stuff. How do you solve that?

    Gaurav Misra37:57

    Yeah, I mean, to be honest, Whisper is already trained on that type of data, right? So it does work really well for those types of things. But what Whisper isn't good at is, like, it has a core set of languages that it works well on, and there's other things it doesn't work well on. And I think figuring out the limits of Whisper in different places. So to be clear, we don't train transcription models, but we do work on that product area very carefully, right?

    Gaurav Misra38:28

    And a lot of it is figuring out what works well and for which use case in which condition, right? And then essentially using a series of different models, correcting each other's mistakes, using it like, oh, what do you use for, like, I mean, for example, Whisper doesn't work in Arabic that well, right? So we have to figure out how to make something work in Arabic, but it can't be bad because almost all transcription is bad in Arabic, right? There's just not enough training that's been done on that language.

    Gaurav Misra38:38

    So things like that is what we figure out. So it's a lot more like product and engineering work rather than model training.

    Personalized fine-tuned models for creators?

    38:49
    Matt Turck38:55

    Is there a world where you would do fine-tuning on a per-creator basis? So, like, for the bigger creators that may have certain ways of doing things, you would fine-tune the model to just go directly to their style?

    Gaurav Misra39:17

    Potentially, yeah. I mean, I think on video generation, it could be very interesting, right? But I think at the end of the day, if the conditioning is expressive enough, right, and we're talking, like, text, image, audio, right, at least these three, you should be able to basically—I mean, that's the point of a foundation model, right? Is to be able to mold it into whatever, is able to work on unexpected, untrained things it wasn't trained on, right? And it can be molded into any particular situation without any fine-tuning.

    Gaurav Misra39:28

    So I do think that's very possible and very reasonable to expect from a foundation model that's large enough and trained on enough data.

    Building models vs. building wrappers

    39:38
    Matt Turck39:41

    It's fascinating what you said about focusing mostly on engineering rather than the models themselves, which seems to be the recipe for success of a lot of AI application startups.

    Gaurav Misra39:48

    100%. Training a model is expensive. I mean, you have to pick your battles on that one, right? You can't be going around training everything. Yeah.

    Matt Turck40:01

    And it was somewhat controversial a year ago, the famous thin wrapper thing. And now everybody seems to have actually accepted the idea that there's tons to build on top of models that actually creates real value and differentiation.

    Gaurav Misra40:19

    Totally. And I think the unique thing that you end up realizing is actually the amount of complexity in a UI that you actually have to build, right? I mean, for us, building a video editing application, by the way, even putting the AI aside for a second, right? Let's say that was a solved problem, right? We figured it out. It's all solved. It's done, right? Just building the ability to render all this stuff on video on your phone, right?

    Gaurav Misra40:39

    Like, the technical challenges involved, this is probably one of the hardest things you can do besides maybe building a game, right? Which may be the hardest thing you can do. This is one of the hardest things out there. You're using the GPU for that too, right? Like rendering all the video encoding, decoding, all these different formats. Like there's OpenGL running, there's shaders running on top of that that are placing elements and making things disappear and appear and rotate in 3D, right?

    Gaurav Misra40:48

    This is hard stuff, right?

    Matt Turck40:48

    Yeah.

    Gaurav Misra41:05

    And doing that on a browser on top of that, or on an iPhone, right? The camera's on at the same time, all these different sensors are running at the same time. The GPU's running for both models and for video encoding and rendering. Like, this is as hard as it gets in terms of technically something you can do, right? Like, on one end, you'll find applications that are just like, oh, this is just a button on the screen, right?

    Gaurav Misra41:30

    Like, easy to design, easy to build for the most part, right? And a lot of apps actually fall in that category, right? Like, we actually are hard mode in a way, right? Like, there's no buttons on screens, right? It is all complicated UI, really hard to design. How do you make something that has 1,000 different possibilities simple for a user, right? You have to really understand the user. I mean, you just have to be really good at design, really good at product.

    Gaurav Misra41:38

    Really good at engineering, right? Across the board, everything has to be good.

    Matt Turck41:51

    How hard can it be? Yeah, because there's presumably a constant trade-off between providing power all in one and then maintaining the ability to edit and say no to certain content situations.

    Gaurav Misra42:10

    Which is really important. I think for it to be usable, that is actually the key, right? Like, imagine if Canva just had templates, but you couldn't edit them. You can't change anything. Import one, you're done, right? You can import a second one, but you can't change anything, right? That doesn't work. It breaks the whole thing. So I think it's a key element, actually.

    Matt Turck42:24

    So to unpack some of the things you were saying, what happens if OpenAI or Claude comes up with a fantastic video generation model that creates perfect avatars? That's a good thing or a bad thing?

    Gaurav Misra42:37

    I think it's a good thing. Right. I think we're working on, at the end of the day, we own the customer relationship, and if they build a model, we'll include it as another option in the app. Right. No big deal.

    Matt Turck42:39

    And everything you were describing still comes on top anyway.

    Gaurav Misra42:46

    Totally. But we are making the unique bet to say that we can actually build a better model than they can because of our unique datasets.

    Matt Turck43:02

    Yeah. And a lot of the app, going into some of the engineering challenges that you were describing, is real-time and very responsive, including on mobile. So how do you think about local AI versus sort of cloud AI?

    Cloud AI vs. Local AI

    43:09
    Gaurav Misra43:22

    I think, obviously, we prefer moving as much as possible to local just because that means it gets removed from COGS and, sort of from the financial perspective, it's awesome, right? I think not all models are there yet to be able to do that. I think a lot of the easier models, I would say, like transcription, things like that you can run locally very easily, and that cost can be eliminated. But to be honest, even running that on the server is not very expensive.

    Gaurav Misra43:41

    So it doesn't really matter that much, to be honest, except maybe offline access for users is more like a user benefit. But for the video generation models, we're not there yet exactly, especially for the type of users that we're going after. They might have older iPhones and stuff.

    Matt Turck43:42

    Mm-hmm.

    Gaurav Misra44:07

    So they may not have the latest and greatest. They might have two generations before, which may not be able to run a large model essentially on-device, right? Even if it's distilled quite a bit. And people do care about output. And I think one of the main criteria for the ability to actually use something, for it to be actually practically useful versus entertainment and just interesting, is: how photorealistic is it, right? And it has to be past the level that the average person can tell that this was generated.

    Gaurav Misra44:37

    If that's not the case, then it's not actually usable, right? And we kind of saw that boundary being crossed about, I would say, maybe a year ago or something of that nature. And we ourselves run a lot of ads and social media posts and things like that for marketing purposes. And we started using generated sort of videos for that in that timeframe. And initially, people were like, "Oh, this is, like, so fake. Oh my God," right?

    Gaurav Misra45:02

    And very quickly, it became nobody knew any better, right? Nobody commented on it at all, right? People were just like, "It's just another video with a person saying something," and they can't tell. And very quickly for us, it became, "Wait, this actually performs better than if we were to hire someone to make this video," basically, right? And the reason is because we can customize it to an unlimited ability, right? Like, we can get all the variations, test everything, and find the perfect optimum performer, right?

    Gaurav Misra45:08

    Okay.

    Optimizing for low latency

    45:19
    Matt Turck45:34

    Okay, let's finish on engineering. This is super interesting. I want to go into that as well. But just to finish on engineering, another thought that came to mind was latency versus—because this seems to be working very fast. So I guess, how do you optimize for latency, especially across mobile and desktop, which have different kinds of requirements, I guess?

    Gaurav Misra45:53

    Yeah, so we actually start off by thinking about how much a user will actually wait. And this is something that we can actually measure. We've done a bunch of work in the UI and stuff to make the experience of waiting as nice as possible. So we'll notify you when it's done—the whole push notification. We even do the little thing that shows up, the Dynamic Island, the thing that Apple marketed so much, right?

    Matt Turck45:53

    Yeah.

    Gaurav Misra46:13

    Just like you see an Uber delivery coming close, you see your video being generated on the screen. So we've done all the things you would expect from a UX perspective to make the process of waiting as nice as possible for the user. But even with that being said, people don't want to wait, right? Especially the consumer-type customer. They don't want to wait. And we can measure the different drop-off rates, and we kind of end up with, like, okay, two minutes, three minutes—that's kind of where we want to be.

    Gaurav Misra46:43

    Right. And so after that, it's just an engineering challenge. So we started off by, okay, let's divide the video into pieces and generate in parallel, right? So this is a significant effort, and we designed this entire system of how we're going to do it. We were splitting the video into, like, 25-frame segments, basically, and 25 FPS, so that's basically a second. And then we were generating overlaps on both sides, four-frame overlaps with the next segment.

    Gaurav Misra47:09

    And then we would check for differences in the overlap and pick the right one to make sure it's a consistent and smooth sort of flow from one second to the next. And then it would just send it out to this entire cluster, which would just eat the entire thing in one go, right? So now the video takes as long as 25 frames, basically—as long as that takes. And yeah, this is very spiky, by the way.

    Gaurav Misra47:36

    Like, if you're sending a lot of videos through, these are explosions, right? Like, a 60-second video means 60 requests, 60 machines occupied in one second, right? But it did mean that for the average user, the experience became really fast, right? But the whole process of doing all this stuff and stitching it, it is a complicated system, right? It is transferring, and these are frames of video, right? At HD or 4K resolution sometimes, right? Because we generate 4K.

    Gaurav Misra47:55

    So each video is a lot. It's maybe 10 gigabytes of frames being transferred around constantly across all this, stitched back together, unstitched again, then restitched back, and then generated and processed and upscaled, and all this happening seamlessly behind the scenes for, like, in two minutes your video's ready, basically.

    AI/ML stack at Captions

    48:07
    Matt Turck48:13

    Amazing. What's the stack, the AI/ML stack you mentioned that's important, like OpenAI running on Azure? Are you an Azure shop? And I don't know, what do you use for model orchestration or any tools or any part of the stack you can talk about?

    Gaurav Misra48:37

    So we're primarily on Google. Google Cloud is what we use mostly for our backend systems. Obviously, we're using non-Google providers for different select niche services that we're doing, like image generation or various types of, like, ElevenLabs, for example, with voice generation. But for the core services and compute, we're using Google Cloud, including a lot of the GPU work, which, by the way, we get GPUs for two reasons, right? One, for media processing, which is a different type of GPU, right?

    Gaurav Misra49:04

    There's GPUs that are designed for media processing. And, by the way, the H100s and A100s kind of suck at encoding and decoding video, right? They're actually slower than a T4, for example, which has media drivers and stuff that can actually do it really quickly. So we have to have these two different types of GPUs. And that also means that work needs to be passed between the GPUs oftentimes. So an A100 and H100 might do the generation of the video because it needs 80 gigs of memory and all this fast AI processing.

    Gaurav Misra49:26

    H.264 or whatever the format is. So all this has to be orchestrated together. And a lot of that's happening on Google, but now for the H100s, we're using niche providers that we're stringing together with T4s that Google is providing at really good cost.

    Matt Turck49:39

    And the orchestration, you use some, I don't know, LangChain or whatever, or, like, who is that?

    Gaurav Misra49:43

    We're actually just building it. We're just rolling it on our own.

    Matt Turck49:43

    Yeah.

    Gaurav Misra50:00

    So it is very niche, almost in a way, because we have to build these pipelines of video processing because you might imagine, for example, you recorded a video and then let's say you were like, "Oh, actually, I mispronounced a word. So I'm going to change that word and correct it to what it actually is," right? And then, "That sentence over there, my eye contact isn't very good, so I'm gonna look at the camera in this sentence."

    Gaurav Misra50:21

    I'm gonna change that, right? So you're actually taking the same video and applying layers and layers of AI processing on top of it. One changed your lips and made you say the right thing. The other one changed your eyes to look at the camera, right? The third one might be doing something else. So we have to, obviously, not have a system where we're uploading all these frames and then it does the thing and then we download 10 gigabytes again, and the next thing, and then you send it back. You want to do all of it in one streamlined flow.

    Gaurav Misra50:49

    So that's the type of stuff that we have to work on to make sure that all of it is as efficient as possible, because every time you're doing it, there's encoding and decoding involved as well, right? The video encoding and decoding. And so that's just another sort of layer of complexity and of time consumption. Oftentimes, we've found that the video encoding and decoding takes longer than the AI processing, which is like, "Oh my God, we're waiting a minute for the video to encode?"

    Gaurav Misra50:58

    Back to the tutorials.

    Matt Turck51:04

    Yes. Hallucinations. Is that a problem? What do you do about it?

    Gaurav Misra51:06

    I think it's a feature, actually, right?

    Matt Turck51:06

    It's a feature.

    “Hallucinations are a feature, not a bug”

    51:10
    Gaurav Misra51:27

    Yep. So, I mean, I think hallucinations in text obviously can be a bad thing, potentially, because they're saying stuff that isn't true. But in a video, like, you are delegating a lot of decisions to the model, right? At the end of the day, the purpose of having the model and the main benefit—okay, let me make a quick comparison here, right? You can always open up Maya and dream up any scene you want and design it in 3D and render it with ray tracing.

    Gaurav Misra51:54

    It will take maybe more time than the AI processing will take, but you will get a really nice output and it will look great. At the end of the day, I think AI is awesome, especially for media generation, but it's not solving a new problem that hasn't already been solved. It's the same problem, and it's actually a solved problem. It's just making it a lot easier.

    Matt Turck51:54

    Right.

    Gaurav Misra52:16

    So I think with that in mind, it just makes it much more obvious, sort of, what to focus on, right? Yeah. And hallucinations and stuff, these are like features, essentially, in this world. Like, we're saying that the model will make the decision for me, right? And sure, I didn't specify that I wanted a white shirt or a green shirt, but it just gave me a good shirt and it looked great, right?

    Matt Turck52:20

    Interesting, because, like, in your world, more options are better.

    Gaurav Misra52:37

    Exactly. And I don't want to make all the choices, right? Part of the reason I'm coming here—otherwise, I would just go to Maya and design everything exactly as I want, right? Part of the appeal and the reason that anybody can access this technology—there's not a skill gap, basically—is because we're taking away a lot of the choices, right? The model is making these choices for you, where the model knows what good lighting looks like and three-point lighting.

    Gaurav Misra52:49

    And the model just knows what type of shirts you should be wearing or, I don't know, whatever, right? So if you don't specify it—

    Matt Turck52:54

    And you don't have cases where people say, "No, no, I do want a green shirt, and why do you keep showing me?" Totally.

    Gaurav Misra53:08

    Then if you write it down, then it will do it, right? What it will hallucinate is what is not specified, right? So I think prompt adherence is really important, right? What you're saying must get done, but everything else, it's all hallucinations. It's good, right? We should hallucinate everything.

    Prompt engineering

    53:19
    Matt Turck53:25

    And prompt engineering is a big part of what you do. We talked about fine-tuning a little while ago, but most of what you do in terms of passing on instructions to the model is, I guess, system-level prompt engineering.

    Gaurav Misra53:48

    Prompt engineering is actually really important, even for the models that we train ourselves, right? Because every model is trained on prompts in a certain way. And certain models are just better at certain things than other things, right? So a lot of it comes down to, like, what is the model the best at, and how can we massage the words to fit the distribution of the originally trained text that it was given, right? And that's prompt engineering, basically.

    Gaurav Misra53:55

    So we do think about that a lot, not just for our own models, but also for other models that we're using.

    Matt Turck54:07

    So I want to go back to what you were starting to talk about a few minutes ago. So do you think we passed the uncanny valley for those videos and avatars?

    Have we passed the uncanny valley for AI avatars?

    54:12
    Gaurav Misra54:28

    I think definitely, but it is kind of a moving target. I'll give you an example, right? And this is an example not related to us, so it'll be more clear maybe. So I remember when I first discovered ElevenLabs, which was like a while back, I heard the voice and I was like, oh my God, this is going to be transformational. We need to work with these guys. And we reached out to them immediately. They didn't respond immediately, but they were a very small shop at the time.

    Gaurav Misra54:50

    It was truly like nothing that was there at the time could compare to what they had done. Now, some time has passed, obviously, and even a year later after that, the most common complaint that I got from people was like, oh, it sounds so robotic. It sounds so robotic. And it's like, wait, what? Like, this is the same thing. Nothing has changed. But I think people just got used to hearing it so much that they just figure out its nuances, right?

    Gaurav Misra55:12

    And it's the same thing with video. Like a year ago, video just wasn't good enough to even be presented to users potentially. As soon as it got good enough, I think it very quickly crossed the boundary of like, wow, it's very good. And we're actually past the uncanny valley. It became completely usable, right? I think very quickly people will start realizing now that, oh, there's like certain nuances, it moves certain ways, it does certain things that you can just tell if you just know and you've seen enough of them. You can tell that this is generated, for example, right?

    Gaurav Misra55:35

    And then our job is gonna be moving past all of that, right? And then people will come up with new ways to tell, and then we'll move past that until there's no way to tell, right?

    Matt Turck55:35

    Mm-hmm.

    Gaurav Misra55:59

    But I do think the point of true complete perfection in terms of generation, where there is absolutely no way to tell, is still a couple of years out, especially in terms of this type of A-roll video, because a lot of the core problems haven't even been fully solved yet. But I think today we already are at a point where a vast, vast majority of people cannot tell the difference, right? They can look at it. Like, we run these things in ads at scale.

    Gaurav Misra56:19

    We also run these things on social media. Big creators use these models and they get tens of millions, sometimes many, many tens of millions of views on their videos, and nobody can tell that it was generated. There's no comment or anything that says this was generated.

    Matt Turck56:30

    Do you think there's an element of social acceptance as well, where people may be able to tell the difference but may not care?

    Gaurav Misra56:53

    Yeah, I think that's definitely a big part of it. What we actually ended up realizing is that people actually don't care. They just wanna be entertained, right? And actually, even the reason that they might comment that it's AI-generated is just to get clout, right? Like, they're just trying to get likes on their comment. Like, they're like, I figured it out, this is AI-generated, right? And everybody's like, oh yeah, it is. And they like it, right? So that's basically why they're doing it, right?

    Gaurav Misra57:13

    Like, it's not that they hate it or something, right? It's just that they want to show that they're informed and they figured it out, and other people to like their comment. That's all, right? But we're definitely past that point too. So now people can't, most people can't even identify it, right? So there's not even that comment anymore. And so the really cool stuff that we started seeing was actually in the last six months, where we realized that especially with ads, you can actually outperform a traditional content production workflow and you can actually beat what can be done with just sort of human creators because you can create an infinite number of variations, run a lot of these in parallel, find the winning output, which you just can't do with a full person-based production, and just get the maximum value essentially, right?

    Gaurav Misra58:03

    But the most interesting beyond that also gets into sort of the localization aspect of it, right? Because finding a good creative in English is hard enough. Like, finding something that just resonates with people is hard enough. But once you get that to work, now you gotta figure out how to make it work in every other language. I gotta make it work in Spanish and French and German, right? And rediscover it, redo it from the beginning, right? But what we figured is like actually just dubbing it literally with Lipdub, and now it's like the same exact script just in a different language, delivered the exact same way, right?

    Gaurav Misra58:19

    Actually performs just as well as the English version, right? Which was a crazy discovery for us, completely unexpected.

    Matt Turck58:22

    Is there a world where all of this becomes completely real-time as well?

    Gaurav Misra58:28

    It's possible, yeah. I mean, I think it's not something we're chasing after, but there's definitely companies that are going after that.

    Matt Turck58:40

    So what does that mean for the future of creation and storytelling? And also, I guess as a related question, what does that mean for the future of professional video editors?

    Gaurav Misra59:03

    Yeah, I actually think that this is really big for the future of professional video editors. I mean, so our goal actually isn't to disrupt professional video editing, right? It's actually to enable, and it actually probably creates more professional video editors because a lot more people can get into the profession, right? Because they may have been too scared of opening Premiere Pro before, and now they can actually make cool stuff, right? With a little bit of dabbling. And of course, the next question they're gonna ask is like, oh, how do I move this?

    Gaurav Misra59:26

    And how do I change that? And that gets them into it, right? And they get better and better. And sure, maybe one day they graduate to Premiere Pro, and there's nothing wrong with that, right? That's completely fine with us, right? But I do think that the craft of video editing will change. We're not gonna be the ones maybe to change it, or maybe we will, I don't know, right? The future will tell. And this is, by the way, not the first time that craft has evolved, right?

    Gaurav Misra59:40

    Like, craft evolves all the time, right? A great example is just like music. Like 50 years ago, you needed to play an instrument, right? To play music. If you didn't play guitar, you're not gonna be a musician, right?

    Matt Turck59:40

    Mm-hmm.

    Gaurav Misra59:55

    And then digital music came along, right? And suddenly anybody can go on a computer and write some music, and it plays and it sounds really good. And a lot of people thought this is bad for music. Like, now there's gonna be all these bad musicians who are gonna be making music, right? But I think the reality is there was just more music and more musicians and maybe even better music and new types of music that didn't exist before, right?

    Matt Turck1:00:04

    Mm-hmm.

    Gaurav Misra1:00:25

    So I think it's the same thing. Craft is evolving, right? And people evolve with the craft. I think for video editors, there's a small chance that it actually becomes the most important profession of them all. It actually is possibly the profession that rules them all because suddenly, one video editor sitting in their basement can make a movie that goes to theater, right? Because they don't need anything. They just generate all the footage, put it all together, and they're done, right?

    Gaurav Misra1:00:31

    That might be a real possibility.

    Matt Turck1:00:40

    How far are we from that? What does that depend on? Does that depend on Runway or Sora doing the B-roll while you do the A-roll?

    Gaurav Misra1:01:05

    Yeah, I mean, it definitely could be. I think, obviously, for a movie, all the cinematic shots and stuff, all the B-roll is going to become really important, all the A-roll. And then, besides that, the editing, right? So as soon as we can put all that together, there's a lot of steps along the way, like everything from character consistency, which everybody talks about quite a bit, but really what it is today is just face consistency, right? But face is not enough, right?

    Gaurav Misra1:01:30

    Everybody has one body. It doesn't change shot to shot, right? It has to be the same exact body, the same hand size, height, or whatever. Everything has to be exactly the same, but even location consistency, right? Like, we're sitting here, the next shot better be looking exactly the same, right? Not slightly different. So these are the types of things that will have to be developed over time in order for all this to be enabled. But I think it's all very solvable because we've already solved it.

    The impact of deepfakes

    1:01:47
    Matt Turck1:01:55

    Not to kill the mood, but I feel like any conversation about AI video would not be complete without the inevitable conversation about deepfakes and taking people's appearance and making them say stuff that they don't want to say. So how do you think about it, and what do you do about it?

    Gaurav Misra1:02:18

    Yeah, it's a really good question. Something we spend a lot of time thinking about, to be honest, right? Because I think of all the possible sort of AI unlocks, even compared to LLMs and stuff, there's probably nothing more dangerous than video generation in a way, right? And we have actually a framework of how we think about this, right? So if you think about video, it actually splits into two categories for us, right? There's what we think of as documentation.

    Gaurav Misra1:02:47

    This is like, from a personal level, this might be something like you taking a picture with your family, with your friends, or video for that matter, right? For documenting, like, oh, we were here, we were at this restaurant, we had fun, right? Like, it's literally capturing the moment that actually occurred, right? But on a non-personal level, it could be a journalist documenting a war happening, right? What happened? Who was involved? Where did people go, right? What were the good and bad things that happened in some situation that they were documenting literally, right?

    Gaurav Misra1:03:19

    That's sort of one side of things. On the other side, you have what we think of as storytelling, right? This is everything from ads, movies, TV shows, like literally all the things that entertain us, right? Nobody has the expectation that these will be completely true, right? Like when you watch an ad for GEICO, you're not thinking that there's a gecko selling insurance somewhere out there in the world and they captured it finally, right? And even reality TV and social media, it's all designed for entertainment and fun.

    Gaurav Misra1:03:48

    Like, it's designed to make you laugh, have fun, and maybe imagine something that you would never have imagined before, right? But on the documentation side, there's no good thing that AI video can do. Not one good thing, right? Like, it's all bad. There's nothing good. So if we can find a way where we can prevent any impact to the documentation side, there's only positive, basically, that AI video can do, right? Everything is just about more people having access to be able to tell better stories, entertain more people, have fun, right?

    Gaurav Misra1:04:00

    There's absolutely all positive. So that's the framework that we use.

    Matt Turck1:04:04

    What's on the roadmap for the next year?

    Gaurav Misra1:04:20

    A lot of good stuff. So a lot of it is going to be bigger and better models, basically, right? So I think in our current biggest model, we're only using less than 1% of our data that we have, our proprietary data. So we are scaling up those models, and they're going to be coming to market over the rest of this year.

    Matt Turck1:04:29

    And if there was one thing in the AI world, from a technology standpoint, that you would want to happen that would really help you, what would it be?

    CapCut ban and its effects

    1:04:33
    Gaurav Misra1:04:47

    I think it's already happening. CapCut's getting banned. I think you may not be able to give an example of a situation where a company's competitor has been banned. It's just unheard of in the U.S., at least.

    Evolving from paid to freemium

    1:05:05
    Matt Turck1:05:15

    Okay. Fantastic. A few words on the business itself. So you started with what mostly looked like a PLG-type motion, product-led growth. And you were saying at the beginning of this conversation how, in the beginning, you charged people, but now you made the product free. At least there's a free tier. Is that correct?

    Gaurav Misra1:05:15

    Yes.

    Matt Turck1:05:19

    So maybe walk us through that evolution, and how do you sell currently?

    Gaurav Misra1:05:40

    Yeah, so it is a consumer application first. So a lot of our distribution happens through consumer channels. The app started off as a completely paid product. So it was a rare sort of premium-only app. I think it's almost nonexistent at the time, or people would call you crazy for it. It actually helped us develop the best possible product by just collecting the best possible feedback from the users who were willing to pay. And then more recently, we actually switched to the freemium model, which is more recognized, because we realized that we can offer a lot of the classic video editing stuff, right?

    Gaurav Misra1:06:10

    Like literally the old way, right? What people used to do, which is manually record your video, manually edit your video, that should all be free, right? Because there's no COGS in it anyways, right? It doesn't cost us anything, right? So why should it cost the user? And we think that's the old way anyways, right? We should just give it up. And then our goal is to convince those users to try the new way, right? And hopefully we can convince them that they don't need to spend an hour recording and editing a video.

    Matt Turck1:06:27

    Has your go-to-market motion evolved to include some outbound sales now? As presumably you go into larger accounts, you mentioned some e-commerce companies.

    Gaurav Misra1:06:44

    Definitely, yep. So one of the big things that we noticed from the very beginnings of the company is that the consumer growth actually wasn't all truly consumer. It was a lot of small businesses, but it was actually a lot of medium and large businesses as well, including some enterprise. And a lot of the use—I mean, people were using it as a consumer app, so they didn't have consent from the company to use the product, but they were still using it.

    Gaurav Misra1:07:16

    And there were many people in each company, each of these companies, using it. So I can't give any examples, obviously, for obvious reasons, but these are very, very large organizations, right? Like, some are trillion-dollar organizations. With that being said, it made it quite obvious that there had to be a similar sort of growth vector to how many PLG products have grown into the enterprise sector. So for us, a lot of our broad growth comes from the consumer, but we do have a B2B team now that's working on essentially converting some of those users into enterprise contracts, especially the ones that are in enterprise already.

    Building a company on foundation models

    1:07:42
    Matt Turck1:07:50

    And from a gross margin slash unit economics kind of perspective, it's always super interesting, right? We, as an industry, are experimenting with what it is to build a company on top of foundation models from a business standpoint. What is your experience so far?

    Gaurav Misra1:08:19

    Yeah, so it started off being more expensive, and the price has only gone down, right? It's literally only gone down. It's gone down more than 10 times every year that we continue down this journey. So the price is decreasing not just at the GPU level, which is obvious, right? Those prices are coming down, but also the models are getting more efficient, and the software is getting more efficient, and the use cases are getting more efficient. So with all of those things, suddenly there's a huge compounding effect of cost reduction on these COGS.

    Gaurav Misra1:08:47

    Today, for us, COGS is, like, the least of our concerns, basically. I mean, to give you an idea, if we were to remove a lot of the capital investment in model training and things that are long-term value for us, we would be cash flow positive or something. It basically doesn't matter.

    Matt Turck1:08:52

    To close, you and I are both in New York.

    Gaurav Misra1:08:52

    Yes.

    Running an AI company in New York

    1:09:01
    Matt Turck1:09:09

    And proudly so, which is a little bit still, if you listen to certain circles, against the dominant narrative that you can actually build top-notch AI companies in New York. So this is our New York rah-rah moment.

    Gaurav Misra1:09:10

    Exactly.

    Matt Turck1:09:21

    So, jokes aside, what has your experience been? Why did you decide to build a company here? Why did you decide to have everyone here versus distributed?

    Gaurav Misra1:09:35

    Yeah, I mean, to be honest, from the very early days, we started during peak COVID, right? Like 2021, right? And it was a time where you couldn't even—San Francisco had their hotels shut down, right? I remember you couldn't even go there and get a hotel room, right?

    Matt Turck1:09:36

    Yeah.

    Gaurav Misra1:09:58

    And so we were a fully remote company, but as soon as COVID started exiting, we realized that to start a company, there's just nothing like being in the same room. And so we started doing the in-person thing. First started with one day, then two days, then three days, and it became five days very quickly. I think New York is uniquely positioned because of a couple of reasons, right? One is that in-person companies do have an advantage over not-in-person companies. No one is willing to say—I mean, now many more people are—but it is very true that there is an advantage there.

    Gaurav Misra1:10:22

    And the advantage is not just in terms of, like, oh, the meetings are much more efficient or something like that, right? It's actually even on a personal and relationship level. People actually get to know each other. Think about me and my co-founder Dwight: we only overlapped for three months, and we kept in touch for 10 years. That would never have happened in a Zoom meeting, right?

    Matt Turck1:10:22

    Yeah.

    Gaurav Misra1:10:41

    It just wouldn't. And those are the types of—you need trust. You need people to be compatible with each other, to be able to work together closely and build something amazing. There's just nothing like being in person for that, right? And New York is actually well positioned to be the best place to be in person, right? Because a lot of people live within 20, 30 minutes, very easy commutes, trains everywhere, things generally work, right?

    Gaurav Misra1:11:06

    As compared to other cities where it might take a long time to commute and stuff, right? New York, everybody has these small, tiny apartments. Everyone wants to get out. It is actually designed for an in-person environment, right? And so I think that's a big advantage, to be honest, right? It actually gives a unique advantage to New York companies that choose to do the in-person thing around how they can have a permanent edge over any other company that's based in any other place, right?

    Gaurav Misra1:11:20

    Because they all tend to be generally remote or hybrid.

    Matt Turck1:11:24

    And what has your experience been in terms of technical talent?

    Gaurav Misra1:11:47

    Really good. I mean, honestly, a lot of bigger companies also moved to New York over the last decade and stuff, and have built up enough sort of muscle that there's a lot of great talent already here. I also think, interestingly, post-COVID, there was a huge migration from San Francisco to New York because I think people suddenly had the ability to be remote, and they were like, wait, why am I in San Francisco? Where should I be going? And I think the choices were basically New York or L.A.

    Gaurav Misra1:12:13

    And a lot of people ended up in New York, which actually I think was, in my opinion, what raised a lot of the rent prices and stuff in New York, is everybody from SF sort of migrating over here. But I think it created a lot of talent that is sort of hidden talent. They don't exactly work in New York. They actually work in SF, right? But they live here. And that's sort of there underlying the overall trend.

    Matt Turck1:12:31

    All right. So if anybody listens to this and is not based in New York and thinks that's just an insufferable conversation, first of all, we love San Francisco as well. And second, I think the broader point is that great companies, including very much great AI companies, can be built in a lot of different places.

    Gaurav Misra1:12:32

    Definitely.

    Matt Turck1:12:37

    One region doesn't have the complete ownership of great companies.

    Gaurav Misra1:12:38

    Definitely.

    Matt Turck1:12:52

    Great. Well, thank you. That feels like a wonderful place to leave it. Congratulations on everything you've built. Very impressive, fantastic story, and excited to see what's next in the next year or two. So thank you very much for doing this.

    Gaurav Misra1:12:54

    No, thank you for having me. It's great.

    Matt Turck1:13:14

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.