You alluded to some of the models, so let's get into that. What do you run on in terms of core models? What is third-party? What is maybe open source? What is proprietary?
Yeah, so essentially, internally, what we work on is just video generation, right? That's kind of where our core area is, like what we can really excel at because we have the data for it, we have the talent for it, and we have the unique ability to build that specific type of model. Right now, for everything else, we use some sort of provider, and we've tried basically all of them at this point. So we've been able to figure out which ones work best, but it's an evolving space.
And so every two or three months, we have to retest everything to just understand where everybody is. There's often winners that come up over and over again that we see over and over again, like ElevenLabs, for example, comes up over and over again in the audio space as—
And that's who you use for—
We have been using them for a long time. Okay.
What other providers, or anything that could be helpful to people who think about building AI products and tips? I don't know, what's worked, what's not worked as well.
Right. I mean, so we were on OpenAI for a while. They had a lot of downtime and other issues that were happening. I think that was an intermittent problem. Maybe it's solved now. So we switched over actually to Microsoft OpenAI for a while. And that actually did a lot better.
And that's OpenAI for what part?
So this is for script generation. Actually, it's like pre-video, basically. But also within the video here and there, when we need to figure out certain editing things, we might use one of these models. But for the most part, Microsoft OpenAI did much better in terms of reliability. It just worked perfectly for latency and never went down, which is what you want. It's great. And then, but we ended up switching to Claude. So now we're on Claude.
And that was just because of—
Okay. And how did you make the decision? Why did you switch? What are the pros and cons?
Yeah, so we actually run an evaluation every few weeks to a month, basically, across all the different models. And we actually have a very unique way of doing this, and probably an evaluation that almost no other company can do, not just on the LLM side, but also on the audio gen side and other things like music and sound effects. And I mean, all these things come together to make a video, right? At the end of the day, and this is an important thing to think about, is if you look at all the media generation foundation models, right?
Like, there's music generation and image generation and even stock video generation, and there's voiceovers and, you name it, sound effects, right? Like, where do you think this is all going? Like, what do you do with a sound effect, right? Like, you generated a sound effect, now what? Like, you don't share a sound effect with a friend, you don't post a sound effect on social media.
So you created your own evaluation benchmarks for—
Yeah, you put it in a video, right? That's where it's all going. Everything is going in videos, right? All the stuff is going into videos, getting posted on social media. So we can actually evaluate when the user in their workflow is like, "I generated this voiceover and I don't like it, delete," right? And to a different user, we can show a different model, right? And they might be like, "Yep, I do like it," right? And so we can compare from the user's perspective, perceived when they're not being paid for this, they're paying us for this, right?
They have all the incentive to pick the best one for their use case, right? We can tell you exactly which one performs the best, right? And exactly which ones users prefer when they're actually paying for it and it's right in front of you, right?
It is, yeah. It is exactly right. Yep.
So then you got that flywheel.
Yep. So this helps us figure out which ones are the best, basically. And we always surface the best models for the user.
Can you infer from if a person makes one choice that they're likely to make the next choice, and then you make the recommendation based on that?
Is that something that's regular? We can. This is all possible in the future, but not something we're doing right now.
All right. So we talked about third-party providers, OpenAI and Claude 3.5. So it seems like there's something in the air.
Yeah, they're doing something right. And I'm curious to see how some of these new models like o3 and stuff perform.
Yeah, it's very impressive. We're recording this right after the release of, I guess, o3-mini and Deep Research last night.
Yes, and the pace of release continues to be—
And DeepSeek and all that, right?
Yes, yes, yes. So those are the third-party providers. And then, so you're saying you're building your own models as well for video generation?
So that's the avatar, that's the edits, or both?
It's both. So actually, I didn't mention one thing, which is, on the third-party providers, we also have image generation in the application, right? So people can include images in their video.
Exactly. Yep. So we are using Ideogram for that one because it generates really good text.
And, sorry, I keep interrupting you. Your models.
So the two models that we sort of work on ourselves are video generation and video editing, basically. And on the video generation side, we have a very unique view, is that, I mean, if you look at all the video generation models today that are doing text-to-video, they're all silent videos, right? There's never anybody talking. It's all B-roll, right? It's all essentially shots of New York, right? Shots of the beach, right?
B-roll is the background, right? Just to get the lingo.
B-roll is the background and A-roll is—
Is primary footage, basically, like what actually tells the story. So to give you an example, imagine that we were filming a movie, right? And when the movie opens, right, it might cut to like three shots of New York City from different angles, like a taxi cab going on the street, right? Like New York here, busy things happening, and then it enters this room, right? Boom. That was B-roll. A-roll begins here, right? And then the dialogue begins, right?
And that's just to set the tone of where we are, who's involved, right? And then you get into the real footage that's actually telling the story, right? Which is always like dialogues and monologues and people talking, right? That's what storytelling is at the end of the day, right? And sure, you can make a silent movie, but do you want to? I mean, is that what all movies should be? So I think that's kind of what we realized, and we really focused in on talking videos, is what we call it internally, as what we want to focus on.
And talking videos meaning avatar generation or anything?
So both the editing use case and the talking-head avatar use case are powered by the same model?
They're different models.
They're different models?
Yep. But for the video generation, it's just diffusion model, but it's conditioned on audio.
Okay. Based on open source or your own?
Based on stuff we're sort of building in-house, essentially.
Okay. And what data do you use?
Yeah, so all proprietary data. The data that we're training on internally is all fully licensed, completely proprietary. Obviously, part of it is data we get from operation of the app itself, but we also purchase data. And in the future, we're also going to be generating custom data, as we started to realize that there are certain niches and certain areas where you just can't get data very easily.
Mm-hmm. Is it really hard to get video data these days? Like, you still can't crawl YouTube, right? Only Google can do that.
Yeah. I mean, you can do anything you want, but should you is the question. I do think if the goal is to sell to enterprise, it's really hard to justify it if you're crawling YouTube, and how do you—
It's a fundamental flaw in your legality.
Okay. And what is working very well and not working just perfectly just yet?
So what's working well is what you would expect, which is that you can put more and more data into these models, increase the number of parameters, and they will get better, basically, right? Of course, save any bugs or anything else that might be happening. Now, what's difficult, as always, is to solve—because we are in a unique space from that sense, right? We're doing this A-roll video. And so there are unique challenges in that. Like, audio conditioning is just not a thing that really has been solved at scale, and nobody else has done it, right?
And so that means that we have to be the first ones to run into these problems, try a lot of different things until we figure out what is the best way to solve them, right? And then scale that up, right? Which is also all very expensive, by the way, right?
So there's only so many mistakes you're allowed to make, right? Because every mistake has a price. So running into those types of problems is where the challenges are. But I think, to be honest, that is also where the fun is, really, right? Like, it's like saying we're climbing Mount Everest, but it was all really easy, right?
Yeah. Was lip-syncing, for example, a particularly challenging area?