Is that a concept of a world model that's built into the image representation?
Exactly. Exactly. So basically, you want these models to also be like a world model. You want these models to know about the world. There's a good chance that you can actually teach your model about the world just by presenting text to it, but it's just not efficient. And a good shortcut would be to bring multimodality in. And the best way of learning about a modality is learning how to generate that, right? So we got to this point that, okay, we've been having Gemini generate images from Gemini 1.5.
So Gemini was multimodal from day one. Gemini 1.5, Gemini 2—it was not great. It really needed a push. And then we figured out how to push this without introducing any regression to other capabilities that the model has, and bring all of these natively into this model. And that was one side that was super interesting for me. Not sad news, but it's really hard to see positive transfer.
So it turned out to be a really, really good model. But it was really hard to see that, wow, I train on images and then text perplexity goes down. That was hard to see. The fact that you train a native model and it's good across all the capabilities is already impressive, but my hope is that multimodality and better models are the way to really push multimodal training to enable positive transfer across modalities. I've worked with people that were experts on this.
For example, one of the things that I remember at the beginning, they were talking about visual quality. And then I remember, it was like, oh, this model is a great model. I sent it to them, and they were like, no, this is not a good model. I was like, what do you mean? And they started showing me two images that, to my eyes, looked the same, but they were saying, no, this is way better.
I was like, no, they're the same. So they had good taste in grasping the visual quality of images. So working with them was really interesting to kind of understand that, okay, there are dimensions. And by the way, their intuition was the thing that actually made Nano Banana a success in terms of being a good product. But it was like, okay, what if we push this towards something beyond traditional image generation? So instead of a translator that, as you said, is text-to-image, it becomes a thinking machine about images.
For example, you enable interleaved text-image generation, where the model can think in not only text tokens, but also in pixel space, right? So it generates text and then generates an image, and it generates another text, another image, and you can leverage that for different problems. One of them is that if you have some sort of a story, text of the story, image related to that text of a story, like a children's storybook, right?
Another one, which I was actually really excited about, was this incremental generation. Let me just give you an example. So if you take DALL-E or Imagen or a standalone image model, right? If you ask these models to generate an image of a scene with 50 details, they might fail, right? And then someone can say that, okay, I can generate a better model that does up to 55 details. And then you say, okay, what about 60?
And then I say, okay, no, let me just go back and train it and then come back to you to cover this. But at the end of the day, there's a threshold that these models can follow instructions to some extent about how many details they capture from the text. But if you have incremental generation, so if you have text and then an image and text and image, you can get your model to generate these details one by one.
So you never expect your model to generate a perfect image in the first shot, right? So you expect your model to plan about this generation. So it says that, let me start with big objects because later I'm going to have a hard time if I put small objects and the big objects don't fit, right? So let me just do that. And then in the next turn, I go with medium objects and smaller, and this is super smart.
And you're never bottlenecked by the capability of a single-shot image generation because you did planning, and then you tune every step's difficulty to match the capability of your model to generate one shot. So that was also one of the things that Nano Banana and native generation, interleaved generation, kind of brought: a completely new perspective to image generation work, which is a little bit far from just translating text into an image.