How do you manage expectations in this world? And I'm seeing this as a general comment, not about you guys specifically, but it feels like AI video is perhaps the most obvious case of amazing Twitter demos or social media demos where it looks fantastic. And then you get on the tool—not necessarily you guys—and then it's work, right? It's a struggle. All creative processes involve struggle. How do you manage this so that people are not disappointed and jump to the conclusion that, oh, AI just doesn't work?
It's all Twitter videos or Twitter or X demos.
Cristóbal Valenzuela34:49 I actually wrote a long post about this very recently, but there are a couple of things. The first one is setting the right expectations for people. I think, as you were saying before, most of people's experience with AI these days, if you look at a macro level, has been with chatbots. And the way you interact with a chatbot is you give the system one prompt, one answer, and you expect one answer back. And it needs to be true, and it needs to be good, and it needs to be like one single thing that I do, right?
Cristóbal Valenzuela35:14 I think if you're completely new to creative AI or using AI to make images or videos, you might come with a very similar expectation, being like, I have this incredibly creative thing in my head, which is a complex scene that I only can visualize internally. I'm going to go into this software, I'm going to type the words that I think describe what I'm seeing, and then once it's out, I'm going to be extremely frustrated because it doesn't match what I had in my head.
Cristóbal Valenzuela35:36 And the conclusion is, therefore, that it doesn't work. And so for me, it's like you're watching a Christopher Nolan movie and you realize he used a camera for that. You go and buy the same camera, and you press the button to record, and you watch the output, and you're like, those two things don't compare.
This camera does not work.
Cristóbal Valenzuela35:57 Right, exactly. It's not me, it's the camera. The camera doesn't work. I make that kind of comparison because I think ultimately this is a creative medium that requires you to experiment, spend time understanding how it works. If the assumption is you're going to come and make a film by pressing a button on a camera, you're not going to understand how cameras work. If your expectation is to come to any AI creative software and type one word and get exactly what you want, you're not fully understanding the extent of how they work.
Cristóbal Valenzuela36:19 And so part of it is just helping people manage their expectations and understand how things work in this new world, that some of the things that you're expecting might not actually happen the way you expect them. It just works differently. And it's just a learning curve. For the first couple of times, it will take you time to adjust until you understand how it works, and you start realizing the potential of it, and you start understanding how you can bring it in.
Cristóbal Valenzuela36:35 And I think managing expectations is probably the most important thing for people who are totally new to the field.
What would you say is the current state of the art in video AI, at Runway but across the industry, precisely in terms of managing expectations? What is currently possible? What is not yet possible? What's truly working, and what's not yet working?
Cristóbal Valenzuela37:16 So I think there's a lot of things that have worked. But I think this field is still very nascent, and there's so many things you can solve for. I think image generation is not fully solved, but it's made a lot of progress. I think most of the tasks that people thought were going to take many years to solve are instruction-based prompting or instruction-based generations. All of the things that are very specific and detailed, I think models are getting really good at.
Cristóbal Valenzuela37:47 On the video side, I would say long consistency of scenes, or being able to cut and have consistent characters within the same generations. Those things are also making a lot of progress. I think real time is getting closer and closer, and I think it will be perhaps the sole focus of many companies over the next couple of months. How do you make sure inference happens on a real-time basis, like you can do with language models these days? You can have a conversation with an assistant. You're going to get to that for video very soon as well.
Cristóbal Valenzuela38:18 I think there are many things that haven't yet been solved, and there are things that are getting better. Like consistency overall is getting extremely good. But I think my belief has always been that, as a field, we solve rendering first. We are able to show and create incredibly consistent videos and images. We haven't yet solved control, which is how do you make sure the model is going to create the thing that you want to create. And control has only started to happen over the last couple of years, months even.
Cristóbal Valenzuela38:25 So there's a lot more focus on control.
For a lot of us people following video AI over the last couple of years, and the broad public, not that long ago we were in the world of the Will Smith spaghetti and then the six fingers and all the things that were sort of easy to poke fun at. What's happened in the last year and a half that all of a sudden we seem to have these completely mind-blowing results? Is there any kind of fundamental breakthrough that happened, any work that you guys did that unlocked this level of quality?
Cristóbal Valenzuela39:28 I think it was just time. I think people know if you believe something to be true, it's just a matter of time until it worked. I think for a long time, all of those cultural moments of the six fingers and Will Smith eating spaghetti, I think for me, those are focusing on a specific moment in time and trying to extrapolate from that towards the future, considering that nothing else will change. And I think we had just a different perspective, being like, yeah, those things are imperfect, but you're not extrapolating well based on what happened before that, and before that, and before that.
Cristóbal Valenzuela40:03 I don't think it was one particular thing. It was more of, I think many in the field have believed some things are going to be true, some things will scale really well, and building the infrastructure to get there is perhaps the thing that takes you the longest. Like, training a model is not trivial. It takes time. It takes time to do the right captioning, the right annotation, the right infrastructure around it, the right testing. But those things are coming. I mean, they will be solved over time.
Yeah, let's talk about that last point in detail, if you will. So let's take Gen-4 and Luma. From an architecture and sort of algorithm standpoint, are they the same thing directionally as Gen-3? Are they more of the same thing? Are they different?
Cristóbal Valenzuela40:43 There's a lot of things that we learn at every model that you build. You learn something around what works and what doesn't. And I think a model is not just like one single idea. It's a combination of different ideas, from how you caption, how you do training, how you test, how you benchmark, how you do different parts of a model. If you swap specific architectures on the encoder or the decoder, there's many parts of model building that, for me, are more like an art than a science that you're going to learn just by shipping one model, then shipping another model, shipping another model.
Cristóbal Valenzuela41:10 Gen-4 has a lot of the things that we've learned over the last three generations of models, and the next generation of models that we're going to be releasing are also going to have a lot of the learnings and the things that worked in Aleph and in Gen-4. And I think that should continue to be the case, where it's less about one single algorithmic innovation that's going to change the entire field and more about how do you make sure that those pieces are set correctly?
Cristóbal Valenzuela41:29 Because I think ultimately it's a complex puzzle that has many different parts, and you just need to know which ones are working and which ones are not.
How long does it take to train a new model in the pre-training world?
Cristóbal Valenzuela41:36 A couple of months.
Yeah, a couple of months.
Cristóbal Valenzuela41:41 It depends on the standards that you have and how big the model is, and yeah, there's a lot of—
The current ones, LLM and Gen-5.
Cristóbal Valenzuela42:05 It takes a couple of months. So the first model that we ever released, I think, was Gen-1 video-wise. It was the first model that was ever out publicly and commercially. It was around 2023. I know that because it was on the front page of The New York Times. Big deal at the time. So 2023 was Gen-1, and now we're in Gen-4, Gen-5 almost. Five years or so? Five years. And so at the beginning it was, like, every 12 months, then I think every eight months.
Cristóbal Valenzuela42:28 Then by now, I think things are getting to a point where you can release new pre-trained baseline models every couple of months. I think partially, again, it has to do with the infrastructure, with the amount of work that you—it's hard to see because when people judge a model, they just judge the output. And I think it's a fair assumption that you're judging what you see, but it's very hard to understand what went into the model itself and how good that infrastructure knowledge of the organization can be used to ship another model and another model and another model.
Cristóbal Valenzuela43:06 And I think that, for me, is the most valuable part of Runway. It's not a model that we put out, because models will completely change every now and then. It's the organizational knowledge and the infrastructure that it takes to ship a model like that. And if you're good at that, you're going to start shipping them much faster than before, which happens to be the case for us.
And to the extent that you can talk about your infrastructure, any kind of detail about how that works, how do you handle all the compute that is needed to train those models? Are you an AWS shop? What's the stack? Anything that you can give us a glimpse about on the infra?
Cristóbal Valenzuela43:52 We started building pretty much, I would say, almost everything from scratch. And so we've spent a lot of time building really good research workflows and tooling for researchers. And those are the things that are hard to measure and see because there is no immediate output or no immediate value yet if you're building and spending time on those. But I think if you make the right bets on the way you manage your data, the way you manage your cluster, the way you do deployments, the way you do research and training jobs, and all of that infrastructure knowledge we build internally, in some cases, I'm pretty sure if we take some of those internal tools and we make them products, they'll be successful products on their own.
Cristóbal Valenzuela44:38 But now, a lot of it has to do with knowing exactly why you're building those kinds of things. And in some cases now we've managed to buy some stuff, like we don't have to build everything from scratch. There's enough knowledge in building those from scratch and knowing why those things work and why others don't work. I think what we can share the most is a lot of what we build is very much custom-built. And I think that's an edge if you can afford to do it.
Cristóbal Valenzuela44:45 And I think we've managed to afford to do it because we just started.
What about the data side? So in AI video, there's this well-publicized debate around, in particular, using YouTube videos and this class-action lawsuit and all the things. Where do you all stand on that? And what data do you use to train the current version of the models?
Cristóbal Valenzuela45:26 Yeah, so we don't disclose what data we use, but we've done some announcements, some partnerships around data. We have one with Lionsgate. We announced another one with Getty Images. And so we have our own internal teams that are collecting data. I think quantity matters a lot, but also quality. So making sure you can curate the right data. Garbage in, garbage out. So if you just put a lot of garbage into the model, you're going to get a lot of bad stuff into it.
Cristóbal Valenzuela45:55 But I think quality then goes back to the question you asked me before. It's like, what's good? In art or in video or in filmmaking or in any artistic endeavor, there's no such thing as a right or wrong answer. There isn't, like there is in a chatbot or in a search engine. And so a lot has to do with just training the right eye to select and curate the data itself. And so data for us, more than quantity, is a lot of the quality component to it.
What about synthetic data? Is that a thing in AI video?
Cristóbal Valenzuela46:16 It is. I think it's becoming more of a thing, I would say. It still has its challenges, mostly to generate diversity of data, but definitely something we're exploring.
And still on the technology front, you are closed source. Obviously, this whole back-and-forth and the theme of open source has been one of the key themes of 2025. Can you imagine that at some point you'll open source some stuff, or where do you stand on this question?
Cristóbal Valenzuela46:54 I do feel that depending on what your goal is, you might just choose whatever is the best outcome for your mission. And I think for us, we know there's a lot to be built, there's a lot of product momentum, that research and product have to work really closely. I don't think open source necessarily gives you that level of control. It has other benefits, it has other things that I think are extremely valuable, and I think there's a lot yet to be built.
Cristóbal Valenzuela47:08 And we've open-sourced a lot of stuff before, but I think for where we are right now, we'll probably continue to build models just internally.