MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    AI is Already Building AI | Google DeepMind’s Mostafa Dehghani

    Mostafa Dehghani is the AI researcher at Google DeepMind. We cover why new models are already built heavily with previous-generation models, why long-horizon automation is limited by compounding reliability and weak evals, and why multimodal training can teach models about the physical world more efficiently than text alone.

    04/02/2026

    Hosted by Matt Turck · with Mostafa Dehghani, AI researcher, Google DeepMind

    AI self-improvementAI agentscontinual learningmultimodal AImodel evaluation
    Listen now
    YouTubeApple PodcastsSpotify
    1h 5m · 27 chapters
    Contents

    Transcript

    What “loops” in AI actually mean

    1:17
    Matt Turck1:32

    One of the hottest concepts in AI research right now seems to be the concept of loops. So I thought it'd be a fun place to start. This idea that models are going to improve not by being bigger, but by thinking recursively. What does that mean exactly?

    Mostafa Dehghani2:02

    Definitely one of the top active areas for almost every lab to invest in: looping. And it has applications at different levels. The one that is on the micro level is basically the looping that we use, like in architecture or at inference time for test-time compute and stuff like that. And then at a higher level is basically the loop that we have over the development of these models, which we refer to as self-improvement. If I want to put it very simply, let's talk about self-improvement as this general concept, right?

    Mostafa Dehghani2:40

    If I want to put it very simply, it is really just the continuation of the trend that we've been riding for decades, right? Think about it: in classical machine learning, humans had to sit down and manually engineer the features. And you had to decide what the model actually pays attention to. And deep learning and neural networks came along and said, okay, let's just remove that. Let the model figure out the representation itself. And that was actually a huge deal.

    Mostafa Dehghani3:08

    And we somehow removed a massive human bottleneck and human bias. And then, instead of just hand-designing architectures, we started learning them too. Instead of curating every piece of training signal, we scaled to basically data-driven approaches and let the data speak. And self-improvement and this loop in development is just the next step in the same direction. And the whole idea and the whole point of it is you're removing the human bottleneck and bias from improving these models, right?

    Mostafa Dehghani3:46

    And now you say that, okay, not just humans don't have to handcraft features anymore, but also we don't want the human to sit in the loop every time that the model has to get better. And I think that's basically on the development side. So it's not radically new. It's the same story, just a new chapter of the same story. I think every time that we removed human judgment from this process, we kind of got over a bottleneck. I would say this self-improvement and looping over the development is kind of like doing that at the highest level, which is basically improving these models.

    Mostafa Dehghani4:21

    If you want to go to a more detailed level of looping, we can talk about ways of increasing test-time compute for these models and how we let these models loop over their process within a specific problem to refine it, to think about it. And I think the most familiar form is just like chain of thought and letting the model think with extra tokens. Beyond that, you can think about different ideas that let the model increase the compute for any specific problem.

    Mostafa Dehghani4:50

    Like, what if I have dummy tokens that they can use as read-and-write tape to verify what I've done and go through the solutions or the process that I'm doing over different steps and understand what has been done wrong, what has to be done next, or even network sparsity, which is basically reusing part of the model multiple times. And this sort of looping has also been shown to be super helpful, mostly because you just let the model throw more compute on a difficult problem.

    Self-improvement as the next chapter of machine learning

    5:04
    Matt Turck5:27

    So that's self-improvement at inference time. I think you alluded to earlier, there is also a bigger concept that's maybe, I guess, more science fiction, except it seems to be becoming a reality very quickly, which is this concept of recursive self-improvement, or RSI. That seems to be what a lot of people are talking about. I think ICLR is coming up in a few weeks, and there's a bunch of papers focused on that. So what is that? What is recursive self-improvement as a concept?

    Mostafa Dehghani5:51

    It's actually interesting because you referred to that as something that looked like a bit of a sci-fi situation, where these models are actually improving themselves. And that's true because a few years ago, when you wanted to talk about this, you could just write a perspective paper at a conference and talk about it at a super high level. But if we go and check out what is happening right now, it's happening to a really good extent, and somehow most people don't realize that this is already happening, especially over the past few months.

    Mostafa Dehghani6:32

    In almost every lab, the new generation of models are built heavily using the previous generation of models. I think that's basically the case everywhere. And it's not fully automatic yet, but the direction is super clear. And it's easy to imagine that we're going to get to a situation with full automation. These models are going to improve themselves and keep learning from the world. And again, it has a relation with other concepts, like continual learning and other concepts that we are still not yet at the most advanced point of.

    Mostafa Dehghani7:07

    But if someone comes and says that, oh, I have an idea to get a model to calculate the gradient and update its weights on the fly, it just feels very normal. It's not something that, wow, this is such an amazing idea. I think what is missing right now is long horizon and full automation. And we are moving in that direction super, super fast. The moment that we have this full automation, I would say, we can close the loop of self-improvement, and then the problems become mostly providing compute for these models to actually do what they want to do.

    Are Karpathy’s autoresearch agents an early form of AI self-improvement?

    7:32
    Mostafa Dehghani7:33

    And as I said in my earlier comment, we just got rid of the human bottleneck for improving these models, which I expect to see a huge jump again from such development.

    Matt Turck7:46

    So people may have seen or heard about Karpathy's AutoResearch project a few weeks ago. Is that an example—presumably reasonably narrow to make it work—of a self-recursive loop?

    Mostafa Dehghani8:09

    That is, definitely. And I think that was one of the early examples of seeing these models actually doing something super sensible on the research side. So we've been seeing them doing a lot of good work on improving the engineering part of the development loop. But on the research side, you think about, okay, maybe some sort of gut feeling or intuition is needed, and a researcher with a long time of playing with these models and experience can do this, but not necessarily a model.

    Mostafa Dehghani8:44

    I think we're seeing the sign that maybe that kind of golden part of a successful recipe that mostly comes from the intuition of a good researcher is coming to these development loops through these models. And it's a bit hard to think about: does it mean that we can replace every genius researcher with these models very soon? Maybe. And I don't know how soon, but this is definitely a sign of something that we kind of doubted a few years ago.

    AI building AI: how close are we?

    8:56
    Mostafa Dehghani8:56

    We couldn't believe that. Wow, this is going to happen that early, which is very exciting.

    Matt Turck9:20

    I want to play it back just to make sure that people listening to this understand. I mean, we're talking about AI building AI. And I think a few months ago, if you talked to researchers, people would say, "Oh yeah, we already use AI to build AI," but that really meant that we use AI tools and reasoning models to come up with ideas and thoughts about building models. But here, what we're talking about is AI automatically updating itself, updating its weights in a recursive manner, leading to potentially a dramatic acceleration in progress.

    Matt Turck9:39

    And what you're saying is that this is largely upon us, and a question of longer horizon and basically more compute. Is that fair?

    The biggest bottlenecks: evals, automation, and long horizons

    10:02
    Mostafa Dehghani10:02

    I think so. This is one. And the other one is also, I'm not going to say that soon we're going to have these models fully automated. There are actually many problems that we have to solve, but directionally, I can see how this can happen. It's not something that I would look at as super hard. It's hard, but very possible. Okay.

    Matt Turck10:15

    So what are the roadblocks? You talked about compute. Is evaluation one of them? Because presumably the model needs to understand what is right and what is wrong in terms of the quality of the answer. Is that one?

    Mostafa Dehghani10:44

    A hundred percent. At the end of the day, you can only improve what you can measure, right? And then getting evaluation is just hard. At the end of the day, it becomes almost a philosophical problem, not just a technical one. This is actually a very interesting observation. If you have a team of super-competent people, most of the time they can make massive progress on a problem if there is some concrete eval to hill-climb on. But if there's no eval, it's just really hard to make progress.

    Mostafa Dehghani11:19

    And the fact that we don't have evals that—or even defining evals that can maybe measure how close we are to the point that we can actually get a self-improvement loop—we just don't have that. And it's just making it much harder to measure the progress in that direction. But there are proxies, and there are definitely some evals that we're going from. Maybe we can evaluate every step of the model toward this direction, and maybe we can evaluate up to this many turns of the model, or maybe we can evaluate the model helping itself to improve in either a specific framework, in this specific setup, and this part of machine learning that needs iteration.

    Mostafa Dehghani11:57

    It's also quite interesting because the difficulty of building evals—the infrastructure that you need to reliably run evals that are super complicated—is also super hard. It's quite funny, but sometimes figuring out, okay, how can I create an environment for a model that operates safely within Google, right, and does all the jobs that a research engineer or research scientist can do in a safe setup? Because right now, we're not confident about them doing the right things all the time, and measuring how much they can push and how long they can push a task is very difficult.

    Can formal verification unlock recursive self-improvement?

    12:36
    Mostafa Dehghani12:36

    And connecting all these points into an environment that these models are operating in, and then getting them to run efficiently and bringing diversity to evals, is definitely one of the bottlenecks of making progress in this direction.

    Matt Turck12:52

    A couple of weeks ago, we had a fun conversation with Karina Hong of Axiom Math, and we talked about formal verification. Is that a promising area from your perspective? Is something like formal verification what would enable you to make sure that the improvement loop keeps continuing?

    Mostafa Dehghani13:20

    In my opinion, formal verification is one of the most powerful keys to enable self-improvement, but it's not the key. And if you think about it, for math, code, logic, it's great. You can run a proof; either it checks out or not. If you go to other domains that are a little bit messier, like, for example, you cannot write a formal proof that a doctor's recommendation is good, right? So it's not easy to extend this formal verification to all the domains in the real world.

    Mostafa Dehghani13:58

    But one question that is actually an interesting question, which is very relevant to formal verification, is: how can we look at these methods and formal verification and build that kind of tight and honest feedback loop for the messy parts of the world? I think that that's very inspiring: to build on top of these formal verification methods to extend to domains that are not easy to verify, but you need some sort of clean and tight feedback loop to be able to make progress.

    What is model collapse?

    14:06
    Matt Turck14:11

    So the same problem as reinforcement learning, right? The second you start veering away from math and code, you start getting into very messy territory. Is model collapse one of the issues to think about, or is that orthogonal?

    Mostafa Dehghani14:34

    Model collapse is definitely a risk, right? And I would say model collapse mainly happens when you have a loop that is completely closed, right? And if you don't have any outside signal and just the model, for example, talking to itself or operating in a very restricted environment, there's a good chance that your model collapses. But if you have a strong verifier or some sort of real reward signal that anchors these kinds of signals that are coming from AI-generated data, for example, it can be quite powerful.

    Mostafa Dehghani15:00

    I think the key here is to stay grounded to something real, and then you can most likely avoid things like model collapse. But yeah, I mean, again, it's a risk, but it's not definitely a major problem.

    Matt Turck15:06

    And perhaps to make this accessible to everyone, can you define what model collapse is in the first place?

    Generalization vs specialization in AI

    15:33
    Mostafa Dehghani15:33

    So basically, when you have some sort of data and environment that these models are interacting with, but those environments and data are designed, for example, by another model, this is just an example of that, right? And then you become really, really good at this specific part, and then suddenly you lose generalization to anything beyond that. And this is one of the definitions, or one of the cases, that model collapse would result in.

    Matt Turck15:47

    So you mentioned losing generalization. Is that particularly, in the context of RSI, a worry—that either you have those self-reinforcing loops, but they need to be fairly narrow, or you have more general models, but then you kind of have the loops?

    Mostafa Dehghani16:13

    This is an interesting question. Again, generalization versus specialization. Let me go a few steps back. We had this discussion many, many times. How should we do a trade-off between generalization and specialization when we are developing these models? I think long-term, you want a model that knows everything and knows when to go deep versus wide, right? Imagine you have an agentic actor, right? If you're an agentic coder, if your agent is super strong at every step of operation, like a really, really good programmer, it's amazing.

    Mostafa Dehghani16:52

    It's super specialized. But for many of the problems, like coding problems, you need some sort of planning and understanding what's going on and collecting information and, based on the context, deciding what to do. And then after you define the steps, then your super strong specialization just kicks in. And before that, being a generalist is super useful. Definitely, generalization is one of the things that you need to get to the ultimate side of AGI. But short-term, I would say building a specialist model is probably the fastest way to learn what is actually possible.

    Mostafa Dehghani17:21

    And in many cases, these specialized models are becoming stepping stones toward a generalist model, which is super valuable, right, Matt? So you can imagine that if I'm actually thinking about self-improvement, maybe I need to make sure that, in a very specific area, I can build that. Maybe I focus on coding, and then if it works out, then I go through how to widen that and how to bring more into this specialized setup.

    What is a specialized model today?

    18:04
    Mostafa Dehghani18:04

    One thing that I always say is that people don't care what category their problem falls into, right? And if a human calls something a problem, then AI should be able to solve it. And I think that's fundamentally a generalist need, right? So at the end of the day, you need generalization, and going through this spectrum of super-generalized models and super-specialized models is more about long-term, short-term, and how to take advantage of each side during this process.

    Matt Turck18:14

    What's a specialized model today? Is that a separate model, or is that a broad general model that's trained in a specific way, including, in particular, through RL?

    Mostafa Dehghani18:40

    Okay, so here's the point. We used to have constraints, like compute. And then if we wanted to push a model to be so tall, we would choose specific dimensions. And then we'd say, okay, we want to allocate the compute that we have to that and then make this model really good at this, something that is extremely expert at this. So that was basically the trade-off that we were trying to make, given the compute budget that we had. As we go through this phase of compute becoming more available, cheaper, maybe we're constrained with other stuff, like data and stuff.

    Mostafa Dehghani19:18

    One of the other trade-offs that pops up is especially in post-training, this game of whack-a-mole that sometimes it's really hard to get your model to be good across the board. So you try to kind of make it good at something like multimodality. Somehow, you see some regression on the coding, and you make it good at coding and multimodality, it becomes slightly worse than a model that you had at math and reasoning. So it's hard to kind of find a balance.

    Mostafa Dehghani19:43

    And part of it is because post-training does a little bit of overfitting. At the end of the day, when you post-train a model, you are trying to overfit it to the best local optimum you have. When the recipe becomes, how can I find the best local optimum, it becomes the problem of, okay, there's no local optimum that is good for everything. So you need to kind of choose, right? And then, seeing this, you end up making some decisions along the way and saying that, okay, maybe for me at this stage, because of the need that I have in my organization, with respect to the competition that is going on, I need to choose this specific axis.

    Mostafa Dehghani20:25

    For example, some companies have a very strong focus on coding, which is, okay, I make my job super easy—or not super easy, but much easier than the competitors that want to basically ship a model that is good across the board. I think short-term, it's very, very effective because, first of all, during development, you care less about all the dimensions. So maybe it's just faster to iterate. You kind of free up some space from the mind of your researchers and engineers that, okay, forget about this.

    Could top AI researchers themselves be automated?

    20:57
    Mostafa Dehghani20:57

    Just let's push this to the max. And then the other one is also like, you don't hit the trade-off immediately. And the specialist model is that, like, okay, I'm going to pick this specific axis and then make the model really, really good at this. Sometimes, again, this is a decision based on the place that you are at organizationally, competitors, and stuff like that.

    Matt Turck21:20

    Great. You said something a few minutes ago that I thought was so intriguing, which is this idea that the Karpathys of the world and you of the world could be automated. What happens if the brightest minds in the world get automated and the AI creates itself? At some point, is there just no one who knows how the AI works? Is that an actual possible future?

    Mostafa Dehghani21:42

    This is actually very philosophical. I don't know. Let me give you one quick thing that I thought about a few days ago. I have a daughter, five years old. I've been impressed over the past few years. Very interestingly, I've been proven wrong multiple times about the timeline that I had in mind. For example, sometimes I say, like, oh, this is going to happen in six months. Never happened. Sometimes, like, oh, this is just so hard.

    Mostafa Dehghani22:08

    Like, within the next 10 years, there's absolutely no chance to solve it. And then boom, in two months, three months, someone had a brilliant idea and they solved it. So it's really hard to predict the future. And I was thinking, like, okay, so you're talking about Karpathy and, again, other researchers, but I'm thinking about, okay, what about the next generation? If my daughter at some point comes to me and asks, like, okay, what should I do?

    Mostafa Dehghani22:32

    What do you recommend to study? What major? And what branch of science or research should I kind of dig in and be the expert on? I really don't have a good answer. Almost it doesn't exist. And it's just really hard to predict the future. What I know is there are a few skills that are probably key to be able to make an impact in this world.

    Mostafa Dehghani22:53

    And also be relevant, staying relevant. One of them is being strategic and having all the parameters on your table when you're making a decision, and becoming an absolute expert about a very specific subject most likely is not going to be useful in the near future. I think the brilliance of Karpathy is not that he's a good programmer or a good teacher. Definitely, he's a good teacher, but I'm saying these are not the most impressive parts of it.

    Mostafa Dehghani23:29

    The most impressive part for me is that he has a really good overall view of what is happening. By putting himself in the stream of information, he can make a decision about, okay, what is the next most impactful thing to do? And now, the things that he does to make an impact are very different from the things that he used to do five years ago. And I think he can continue doing that. What are the things that he's going to be doing in five years?

    Mostafa Dehghani23:45

    I don't know. But I know he's smart enough to figure it out and still keep making an impact on the board.

    Matt Turck23:50

    So AI researchers are not researching their way out of a job just yet?

    Mostafa Dehghani23:54

    Hopefully we are smart enough to do that.

    If AI builds AI, does data matter less than compute?

    24:02
    Matt Turck24:09

    All right. Maybe that's more of a macro question as I think about where the value lands in this ecosystem, but the AI just keeps creating itself. Then is data still needed in that equation, or is that all compute?

    Mostafa Dehghani24:33

    The concept of data is a little bit broader than just tokens, right? And if you think about data as whatever the model can get signal from, either it is predicting the next token in raw text, which we use in pre-training, or a super complex environment that the model interacts with and then gets signal from, this is something that basically we can refer to as data, right? And it's not like data, or the value of having good data or working on data, is going to disappear and compute is going to become the only thing.

    Mostafa Dehghani25:11

    At the end of the day, I think the work that we're doing on the data side most likely is going to shift toward building environments or making sure that these models can interact with physical worlds, and then it becomes more of a problem of, okay, how can I provide more grounding for these models? They are good at improving themselves, but as long as I expose them to real-world data and real-world environments. So providing data becomes more about, okay, how can I give access to this specific model to something that we never had?

    Mostafa Dehghani25:45

    For example, something came to my mind, which is, again, a little bit sci-fi, but how can I make smell accessible to these models? Right now, there's no good way, but then data becomes like, okay, information or anything that is, for us, because of all the senses that we have, really easy. Right now I'm sitting here, I know how hard my chair is, what the temperature of this room is. All this sensory information is something that is coming to me, and then this work that I'm saying is based on all this input, right?

    Mostafa Dehghani26:15

    And then providing this for a model that does self-improvement is already a really hard problem. So I would say that the work on the data would shift toward making this sensory information more available to these models in a way that enables them to really improve themselves, given all this information, in a more effective way.

    Post-training vs pre-training: where will progress come from?

    26:22
    Matt Turck26:46

    Yeah, interesting. Yeah, there seems to be a big trend towards agents as a service. We're seeing startups emerge in that field. Okay, super, super interesting. Zooming out from self-improvement for a second, the big theme of the last year has been the acceleration of post-training in addition to pre-training, so the whole reinforcement learning aspect of things. Where do you expect gains to come from in the next few months or year? Is that more post-training? Is that more pre-training?

    Matt Turck26:50

    Is that both? Is that something else?

    Mostafa Dehghani27:15

    The answer to this question really depends on when you actually ask this question. And it's obvious that we're going to be having a bit of a swing back and forth between pre-training and post-training. At the end of the day, I want to say that pre-training is still the foundation, and you can never post-train your way out of a weak base model. But right now, the current return on post-training is really strong. And I started working on post-training myself a few months ago, on Gemini post-training.

    Mostafa Dehghani27:38

    Mostly coding and agentic. I can see how a brilliant small idea can make a model, for example, 10x better in terms of behavior at a fraction of the cost of the pre-training. Again, we can see how post-training is the place to make a lot of impact and improve these models. But on the other hand, I know it's also the case at different companies, but at least at GDM, a lot of exciting recent work is going into the pre-training side, with new recipes, new ideas.

    Why pre-training is not dead

    28:14
    Mostafa Dehghani28:15

    And I would say the work that we're doing on the pre-training is going to unlock a lot of downstream possibilities. Post-training is just a different mode of operation. It's also super interesting for me because I'm, again, a little bit new to this side of the operation. But at the end of the day, I always expect scaling going on between post-training and pre-training.

    Matt Turck28:24

    Your comments on pre-training are sort of against that narrative that appeared a few months ago that pre-training was dead. That's not your take at all, right?

    Mostafa Dehghani28:55

    I think everyone has ideas on the pre-training side. At the end of the day, going for that idea is a function of complexity and the expected gain. And sometimes you feel that, okay, there are low-hanging fruits. And instead of bringing this complex recipe to the pre-training, the one that I have, which is simple, elegant, super scalable, I'm going to push this and then move the effort to the post-training. And then, at some point, the base model becomes the bottleneck. And then you're happy to take the complex recipe and bring it to the pre-training and then keep pushing it.

    Mostafa Dehghani29:13

    I think pre-training is not dead. I would say maybe the old—it's also a little bit difficult to talk about old and new because the time frame is very different. So when I say old, maybe I'm referring to, like, two weeks ago or something. But the way that we used to do pre-training maybe, like, a year ago or two years ago, maybe diminishing returns are obvious. But I can see how new ideas are bringing fresh energy into the pre-training and suddenly just open a door toward something exotic that might actually drastically change. So exciting stuff for Gemini 4 whenever it comes out.

    What is continual learning?

    29:45
    Matt Turck30:06

    You mentioned continual learning earlier, and that's another one of those hot topics that people have been talking about. Can you define continual learning for us so that this conversation is educational for a broad group of people? Maybe compare and contrast that with the self-improvement loop. Those are two different things, but help us understand the difference.

    Mostafa Dehghani30:33

    Definitely, they're related, but they're distinct, right? So self-improvement is about a model getting smarter over time and improving its capability, like the model itself doing it. Continual learning is mostly about a model staying current, right? Think about a doctor that keeps reading new research, and they refresh their knowledge about stuff and they're trying to make sure that their knowledge doesn't go stale. The shared enemy between self-improvement and continual learning is a model with frozen weights over time while the world is just going, right?

    Mostafa Dehghani31:11

    If you have a model that is just frozen and the world is moving, then you neither get self-improvement nor continual learning. But continual learning is mostly focused on making sure that if there's fresh knowledge in the world, the model's knowledge cutoff is not in the past. So it's constantly, for example, overnight, all the news, everything that is happening in the world, everything is just updated. So if today you ask a question from the model, that knowledge, which is super fresh, is already in the weights of the model.

    Mostafa Dehghani31:50

    So it doesn't have to kind of depend on an external source to bring it in. And it's hard. It's really, really hard. And one of the big problems is catastrophic forgetting, where you get your model to learn about new information after you're done training that model, and then suddenly you see regression in the knowledge that you learned already in the main training phase. And it's a very active area of research right now.

    How real is continual learning today?

    31:53
    Matt Turck32:02

    And what's the reality of continual learning as of now? Is that built into existing systems? Not at all? About to?

    Mostafa Dehghani32:17

    There are two sides of it. One side is, I think, the research is not yet to a point that you think, "Oh, this is the recipe. I just need to exploit it and push productionization," right? But basically, every time that you have a new problem that is key, you have this phase of exploration where people try different ideas and jump over from this idea to another idea, which could be so different.

    Mostafa Dehghani32:57

    And then when you're confident about this kind of working to some extent, you go to the exploitation mode and say, "Oh, let me just make it as good as it can be. This is the way to push it, and let's scale it. Let's develop infra for it, make it super fast, productionize it, and see what happens." I think that is not yet there. The other one is also, again, as I said, because we've never had a super-confident recipe for continual learning, building infra for it and investing in something that is fast is hard.

    Mostafa Dehghani33:29

    Given that, I've seen very impressive progress on this within Google DeepMind. It's kind of interesting because it is one of the things that can be heavily theoretical. I've seen people who are doing a lot of theory work get into this problem, and they're having a lot of fun, and they're also making a lot of impact. It's impressive how much progress we made on this, but I don't think that we have yet any idea that everyone says, "Oh, this is it. Let's just do it. Let's push this."

    Mostafa Dehghani’s background and path into AI

    33:43
    Matt Turck33:55

    Great. I'd love to talk about you and your background. Tell us your story in a few minutes. How did you come to do this work, and what was your journey to AI and then your journey to Google DeepMind?

    Mostafa Dehghani34:21

    So I did my PhD at the University of Amsterdam on machine learning, mostly on the language model side, and text and search and retrieval. And then I think what kind of pushed me toward trying really hard to be on the mainstream and be part of this group that are hustling to make really good progress, I did a few internships back in 2016 and 2017. And the funny story is, I did an internship at Google Brain in early 2017, and then it was amazing.

    Mostafa Dehghani34:47

    It was just like, I went to this team, they were working on LSTMs for summarization. Summarization was actually one of the most interesting problems at that time. I was amazed. I was like, this is so good. I just want to keep doing this for the rest of my life. This is it. And then I got a return offer to go back and do another internship at the end of the same year. The recruiter told me that there's this team that just published a paper.

    Mostafa Dehghani35:13

    Maybe you've heard about it, like Transformer, and they're looking for an intern. And I remember I had a chat with Łukasz Kaiser. And then Łukasz was talking to me and was saying, like, oh yeah, we have this idea of building a Turing machine based on Transformer. And he was so excited about this. And then we finished the conversation, and I started sort of sending a message to the recruiter, and I was like, I don't know if I want to go with this team.

    Mostafa Dehghani35:47

    It's just like, they're doing something random. Everyone's doing LSTMs, but why should I go and work with a group of people who are working on this random architecture, like Transformer? It's just like, it's going to die. And then he tried, and he couldn't find any other team for me to join. So I joined this team as an intern, and that changed my life. Being among these super brilliant, super smart people that believed in some vision and direction where almost everyone was excited about something else was very inspiring.

    The story behind Universal Transformers

    36:13
    Mostafa Dehghani36:13

    And then we worked on, again, this Turing machine idea, which turned into the Universal Transformer paper, where recurrence in depth and reusing parameters were coming out of it. And still, this is making a lot of impact after almost 10 years.

    Matt Turck36:26

    Tell us about that quickly. So that was in 2019, I believe. And you were a co-author of that paper. And that was very much that idea that we started with at the beginning of this conversation of loops and recursive stuff.

    Mostafa Dehghani36:50

    So Universal Transformer, we wrote that paper in 2018. And I think it was also rejected one time from one conference. And it was accepted in 2019. I don't remember exactly, but yeah. Yeah, I think it was accepted at ICLR, but it was rejected from NeurIPS or something. The whole intuition was there is something about reusing parameters and a model going through its output another time. And so basically, you generate something and then you kind of pass it into the model again, and then the model has the chance of doing this.

    Mostafa Dehghani37:19

    So we started with, I remember Łukasz had this algorithmic dataset, which I remember he used to call algorithmic tasks. And it was part of this codebase based on TensorFlow, like Tensor2Tensor was the name of the codebase. It's still there. And I remember I can even find my pull request into that for pushing the Universal Transformer code. And we saw that basically there are some problems like copying an input to the output, or, like, doing something algorithmic with a super long input on the output side, which is super easy, but the normal models, like a normal Transformer, were failing awfully at this.

    Mostafa Dehghani38:01

    And we saw that looping could do it perfectly. And then at that point, I remember we had this bAbI dataset from Meta, and it was doing great on that. And then the idea of test-time compute, which basically you train with a fixed amount of compute, but at test time you unleash your model to do more computation, throwing more FLOPs on the input, was coming to our mind. We were super excited about this. And then we ended up actually kind of introducing this adaptive computation mechanism into this, which was again some sort of inspiration from Alex's paper for LSTM.

    Mostafa Dehghani38:41

    And then, like, a very interesting ride, because we were pushing for something that, like, at that time, sounded exciting. And I have a guess, like maybe at that time, the whole field was a bit too focused on using adaptive computation for decreasing the cost on simple problems. But now we know that maybe we can actually use adaptive computation to increase the cost for hard problems. It's actually, like, the other side of the same coin, right? So because at that time we were, like, maybe resource-constrained and everything.

    Mostafa Dehghani39:13

    So we were really thinking about why we were spending so many FLOPs, like, going through all the layers and everything for a dot at the end of the sentence. If that token is, like, do we really need 24 layers? So how can we decrease that? But now we have a different perspective to that, which is, like, how can we increase this for a physics problem that we want to run the inference for maybe for two weeks? So that was really fun to work on that with these brilliant people.

    Mostafa Dehghani39:46

    And I think just recurrence in depth and reusing the parameters, or I've seen later some people actually framing it as negative sparsity, which is a great way of connecting it to mixture of experts, that in mixture of experts, you have FLOP-free parameters, so parameters that are not actually bringing any FLOPs. And in looping, you have parameter-free FLOPs, where you don't have extra parameters for the extra FLOPs that you're throwing on this. So it goes the other direction of sparsity, and it's quite effective.

    Mostafa Dehghani39:55

    We're seeing a lot of excitement in this direction.

    How Vision Transformers changed AI

    39:56
    Matt Turck40:12

    Fascinating. Another fundamentally important contribution to the field that you did was the Vision Transformer paper in 2022. So the paper was called “An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale.” Do you want to walk us through what that was?

    Mostafa Dehghani40:34

    There's also a funny story for that. I got into vision and multimodality with that paper. So I'd never worked on any vision problem. It was mostly because I was sitting next to people who were working on vision. My desk was next to people who were working on vision, and that was the reason that I got interested, because I was just talking to them. I was like, “Oh, this is actually interesting.” And then I remembered that at that time I was working on what we call the PaLM paper with Aakanksha and other folks.

    Mostafa Dehghani41:07

    And I was like, why do we have 400-billion-parameter language models, but the biggest model that we have on the vision side is just maybe 100 million, like a ResNet? Why is there no benefit of scaling? We started looking into this with folks, like, okay, maybe there's something in Transformer that actually makes it scalable. And then maybe we can move away from convolution to try this. And at the end of the day, I don't want to say that that's the only way of scaling.

    Mostafa Dehghani41:37

    Maybe if a group actually spent enough time on convolution, they can also make it scalable and as good, but there was also a benefit of doing that simply because the rest of the machine learning field, which was working on language, was using this architecture. So they were building infrastructure for it, making it faster, and sometimes the hardware is designed based on this architecture, at least for the short term. So we started pushing, and then I remember that we had a bunch of ideas that, okay, what if each pixel is a token?

    Mostafa Dehghani42:02

    And then the cost was going high. The context was getting super long. And then we had a lot of back and forth. And it's also quite funny because we started thinking about this problem from a very complicated point of view. So we were trying to mimic convolutions to be able to get this working. And it ended up like I had a bunch of colleagues also in Zurich, and they started trying the simple idea of, what if we just divide the image into patches of pixels, 16 by 16, and then get each patch as a token and forget about overlapping patches or windows and stuff like that?

    Mostafa Dehghani42:34

    That's it. Chop the image and then feed it to a Transformer and scale. Go with a lot of data, and then let's start with something discriminative to train this model. And it worked. And it was also a little bit of a surprise for us that we were all thinking about something fancy and very, very complicated, maybe in the integration of having convolutions and stuff.

    Mostafa Dehghani42:56

    But something that worked was basically the simple idea of patchify, feed it to a Transformer, scale it up, and then boom, you had a really, really good model for representation learning.

    Matt Turck43:23

    Yeah. And to play it back at the highest level, that basically meant that you could apply the Transformer architecture to images, wherein in the past you had two different families. You had the CNN world and the Transformer world for text. And your breakthrough was to prove that Transformers could scale equally well to images, which basically paved the way to a Gemini 3 today, which is like a natively multimodal model. Is that fair?

    Mostafa Dehghani43:46

    Yeah, that is true. So basically, with that, we kind of took a step toward having videos, like adapting Transformers, and audio, adapting Transformers. So basically, again, even if this is not the only architecture that would be in a multimodal model, it made it really simple to train these models natively because you have a single architecture and can have all the modalities in during training.

    Gemini, multimodality, and Nano Banana

    43:47
    Matt Turck44:09

    Great. So that's a perfect transition into your work into Nano Banana and the future of image AI. So you are part of the Nano Banana team, which must have been so much fun when this came out and went just completely viral. What an incredible product. So since then, there's been a couple of releases: Gemini 3.1 Flash Image, yeah, at the end of February. So a lot of people assume that image generation works as a translator, meaning that the AI reads the text of the prompt and then translates it into picture instructions and then draws it.

    Matt Turck44:33

    But as we were saying, Gemini is natively multimodal, so how does that work? How does a model actually process the text and the pixels at the same time to build the image?

    Mostafa Dehghani45:06

    I think the reason that maybe I got to the generation—okay, by the way, there's also one thing: I'm not an expert in image generation. When I started working on this, I remember I had meetings with people, and they were talking about computer graphics and all the old ideas or intuitions. And I had zero idea of what was going on. I was like, I know how to train a transformer and scale it, and if it helps, I can basically contribute to this.

    Mostafa Dehghani45:32

    But again, it was fun because I worked with a group of super smart, brilliant people with really, really good intuition. And I think the reason that I was excited about this was, again, this is maybe not super relevant to Nano Banana itself, but just to mention, I was excited about the idea of positive transfer across modalities. So when you think about multimodal natively, one part of it is that, oh, I'm adding capability to my model.

    Mostafa Dehghani46:05

    So my model can understand images and understand videos and understand audio, but also text, but also can generate all these modalities. So I have a model that actually does all these together. This is for sure exciting from the product point of view. You have a model that is a great model for generating all these different outputs, and users are finding it very useful and interesting. But the most exciting part for me was, can I see a glimpse of transfer from these modalities?

    Mostafa Dehghani46:38

    For example, if I train a model to become good at generating images, does it also become better at generating text? There are different intuitions of why this should happen. I think there's something very old in the literature on the linguistic side that they call reporting biases. So, for example, you visit your friend's place, and then you see that they have a banana-shaped sofa. When you go home, the chance of talking about that sofa compared to a normal sofa is much higher.

    Mostafa Dehghani47:09

    So you can actually talk to your friends or partner later: oh, I went there, and their sofa was in the shape of a banana, which was really fun. But if it was normal, it's weird if you go somewhere and it's like, oh, by the way, I went to my friend's place and they had a sofa, which was super normal. So this is the language reporting bias. So language doesn't talk about things that are at the middle of the distribution, right? But if you have an image or if you have vision input from anything in the universe, you have that information.

    Mostafa Dehghani47:29

    There's no need for reporting it. It's just there, right? So because of that, picking up a lot of knowledge about the world through language is just not really efficient. I don't want to say that it's impossible, but it's not efficient. To learn about gravity, if you have your model trained on videos, it's much easier to get the model to learn about gravity because it just happens in a video than training your model on all the textbooks to kind of learn about the concept of gravity or what is actually gravity.

    Why multimodality helps build a world model

    47:46
    Matt Turck47:50

    Is that a concept of a world model that's built into the image representation?

    Mostafa Dehghani48:16

    Exactly. Exactly. So basically, you want these models to also be like a world model. You want these models to know about the world. There's a good chance that you can actually teach your model about the world just by presenting text to it, but it's just not efficient. And a good shortcut would be to bring multimodality in. And the best way of learning about a modality is learning how to generate that, right? So we got to this point that, okay, we've been having Gemini generate images from Gemini 1.5.

    Mostafa Dehghani48:44

    So Gemini was multimodal from day one. Gemini 1.5, Gemini 2—it was not great. It really needed a push. And then we figured out how to push this without introducing any regression to other capabilities that the model has, and bring all of these natively into this model. And that was one side that was super interesting for me. Not sad news, but it's really hard to see positive transfer.

    Mostafa Dehghani49:16

    So it turned out to be a really, really good model. But it was really hard to see that, wow, I train on images and then text perplexity goes down. That was hard to see. The fact that you train a native model and it's good across all the capabilities is already impressive, but my hope is that multimodality and better models are the way to really push multimodal training to enable positive transfer across modalities. I've worked with people that were experts on this.

    Mostafa Dehghani49:37

    For example, one of the things that I remember at the beginning, they were talking about visual quality. And then I remember, it was like, oh, this model is a great model. I sent it to them, and they were like, no, this is not a good model. I was like, what do you mean? And they started showing me two images that, to my eyes, looked the same, but they were saying, no, this is way better.

    Mostafa Dehghani50:11

    I was like, no, they're the same. So they had good taste in grasping the visual quality of images. So working with them was really interesting to kind of understand that, okay, there are dimensions. And by the way, their intuition was the thing that actually made Nano Banana a success in terms of being a good product. But it was like, okay, what if we push this towards something beyond traditional image generation? So instead of a translator that, as you said, is text-to-image, it becomes a thinking machine about images.

    Mostafa Dehghani50:40

    For example, you enable interleaved text-image generation, where the model can think in not only text tokens, but also in pixel space, right? So it generates text and then generates an image, and it generates another text, another image, and you can leverage that for different problems. One of them is that if you have some sort of a story, text of the story, image related to that text of a story, like a children's storybook, right?

    Mostafa Dehghani51:13

    Another one, which I was actually really excited about, was this incremental generation. Let me just give you an example. So if you take DALL-E or Imagen or a standalone image model, right? If you ask these models to generate an image of a scene with 50 details, they might fail, right? And then someone can say that, okay, I can generate a better model that does up to 55 details. And then you say, okay, what about 60?

    Mostafa Dehghani51:31

    And then I say, okay, no, let me just go back and train it and then come back to you to cover this. But at the end of the day, there's a threshold that these models can follow instructions to some extent about how many details they capture from the text. But if you have incremental generation, so if you have text and then an image and text and image, you can get your model to generate these details one by one.

    Mostafa Dehghani52:01

    So you never expect your model to generate a perfect image in the first shot, right? So you expect your model to plan about this generation. So it says that, let me start with big objects because later I'm going to have a hard time if I put small objects and the big objects don't fit, right? So let me just do that. And then in the next turn, I go with medium objects and smaller, and this is super smart.

    Mostafa Dehghani52:36

    And you're never bottlenecked by the capability of a single-shot image generation because you did planning, and then you tune every step's difficulty to match the capability of your model to generate one shot. So that was also one of the things that Nano Banana and native generation, interleaved generation, kind of brought: a completely new perspective to image generation work, which is a little bit far from just translating text into an image.

    Why image generation is getting faster and more efficient

    52:44
    Matt Turck53:05

    Fascinating. Does part of this contribute to efficiency? So especially Nano Banana 2, you have the Flash aspect of this. So you're able to create amazing images very fast and, apparently, seemingly very efficiently. So what's behind the scenes? Is that what you described? Is that MoE? How are you able to do that?

    Mostafa Dehghani53:26

    First of all, I was involved in the original Nano Banana, Nano Banana Pro, and then the last version. I again, because I jumped on the post-training and coding agents, and I find it exciting, this one, the team actually shipped it. But if I want to say, super high level, what exactly are the things that make the model faster and more efficient? Part of it is just the size of the model. So Nano Banana was Pro size, and this one is just Flash.

    Mostafa Dehghani53:55

    So definitely the parameter size, configuration of the MoE, and stuff. The other one was people actually spent quite a lot of time figuring out, nailing down a distillation recipe, both on the side of knowledge and other things that basically you kind of need to distill to something like a process that is lighter than the full process. Surprisingly, a lot of infrastructure work for serving. So we have really, really, really brilliant people that are serving engineers, and it's kind of impressive that you sit at your desk and then they come.

    Mostafa Dehghani54:21

    And they say, "Oh, by the way, casually, I made the model 10x faster." And they're just saying it in a very casual voice. Like, wow, this is impressive. We had also a lot of work on optimizing the serving, how to serve these models. And again, because these models are operating differently from just normal language models, they're not necessarily the same as next-token prediction.

    Hot takes

    54:44
    Mostafa Dehghani54:44

    This is definitely something that a good serving engineer can figure out: okay, I can think about a different way of doing that. And we had also a lot of improvement on the efficiency side by their work.

    Matt Turck54:52

    All right. So as we get towards the end of this conversation, I thought it'd be fun to end with a few hot takes, if you're ready for them.

    Mostafa Dehghani54:52

    Yeah, absolutely.

    What the AI field is getting wrong

    54:53
    Matt Turck54:58

    All right. What is one thing the AI field is getting wrong right now?

    Mostafa Dehghani55:31

    Not easy to pinpoint specific things, but again, this is just my personal opinion, and maybe I have colleagues and/or other people sharing this with me. But I think we're underestimating how hard jagged intelligence is to fix. We're underestimating how much it matters. And we talk about almost like people laugh and go, if you have a model that does a very difficult math proof but has a difficult time counting letters in a word. As I said, people just laugh and move on, but I think it's actually pointing at something deep and unresolved about these systems, the way that these systems can represent and process knowledge.

    Why continual learning is underrated

    56:17
    Mostafa Dehghani56:18

    And it's not a bug that you can patch. So definitely, we see that this is happening. People sometimes, or we have these problems that something is off really bad. And then you can, oh, let me just patch by adding something to the system instruction or developer instruction. It's a bit of a structural property of how these models actually learn. So I would say this is probably one of the things that we're not getting super right at this point. Great.

    Matt Turck56:22

    What is one idea in AI research right now that is underrated?

    Mostafa Dehghani56:48

    Something that is underrated, like you mentioned, continual learning. I think this is definitely underrated. As I said, sometimes the problem stays in the exploration mode until we are confident about something, and then it goes to the exploitation mode. I think we are past the time that we really had to push this to exploitation. So maybe foundation models are essentially right now frozen in time when the training ends, right? And then everything is built on top of this frozen model, like in a RAG pipeline and fine-tuning workflows and retrieval systems.

    Mostafa Dehghani57:22

    And all this elaborate infrastructure is all based on this assumption that these models are frozen. And it's a bit too strong of an assumption to make. And I think we are going to get to the point that we need to change these assumptions, and maybe we need to think about it a little bit more actively and start pushing it toward something that we actually push to productionization. And maybe it's a little bit underrated right now, like continual learning.

    Does RAG go away over time?

    57:26
    Matt Turck57:29

    So you think RAG goes away over time?

    Mostafa Dehghani57:51

    It's not going to look as it is today, and it's going to be different. But saying that it's going to go away completely, I'm not sure about that. And one of the reasons that I say that is RAG is not just about bringing fresh information to the model when it wants to solve a problem about the current state of things, but it also has this in-context learning. And there is a difference between in-context learning, the information that you have in the context of the model, compared to the information that you have in the weights of the model.

    Mostafa Dehghani58:19

    Continual learning and RAG are doing different things for bringing this fresh information. Maybe it changes in a way that it doesn't need to trigger RAG for everything, but I'm pretty sure that there's going to be some tail of the distribution that we're going to do RAG still for it. What's the task?

    What people are too confident about in AI

    58:21
    Matt Turck58:24

    All right. Last couple of hot takes. What do you think people are too confident about?

    Mostafa Dehghani58:49

    So people think that pushing the technical side is sufficient, that if we just get a model that is smarter, everything is going to follow. And in my opinion, a version of AI that is really, really brilliant at technical problems, but has a blind spot about everything else, is not going to be able to actually create meaningful progress in the world. And the fact that people are kind of confident about this, that everything else is going to follow, or everything else is just a small list, I think is wrong.

    Mostafa Dehghani59:31

    We have governance, we have regulation, we have social trust, we have, for example, distribution of access and the benefit in the world from this technology. And even the institutional capacity to absorb and adapt this technology is something that maybe we don't pay enough attention to. And these are not really soft problems, if not harder than the technical part. They're really hard. And the pace of technical progress is definitely currently running ahead of the world's capacity to develop this kind of mechanism.

    If he were starting from scratch today

    59:56
    Mostafa Dehghani59:57

    And this gap is getting bigger and bigger. But what I'm saying is basically the field needs to hold both things at once. So maybe that's one of the things. All right.

    Matt Turck1:00:09

    And last one, and I don't know if that's a hot take or maybe just advice for anybody entering the field today: if you were going to start from scratch today, what would you work on?

    Mostafa Dehghani1:00:36

    I don't want to start from scratch. It's hard to start from scratch. I can tell you there are two things that I think would be nice to spend more time on, and there's one thing that I'm very excited about. I'll start from the thing that I'm really excited about. And I would say, short term, it's really exciting to push it. I am actually trying to even be able to contribute to this direction.

    Mostafa Dehghani1:01:05

    And that's full automation of super-long-horizon tasks, things that you have a machine working on for maybe two weeks, one month. The agents today are very impressive, right? And the demos are very remarkable, but there's this compounding reliability problem that doesn't get talked about enough. And, for example, imagine if an agent has to take 100 sequential steps to complete a task, and imagine if each step has 95% success, which is great. Given the models that we have today, 95% is really good.

    Mostafa Dehghani1:01:50

    95 to the power of 100, which is less than 1%. And this math is brutal, right? And this 95% per step, as I said, is very, very optimistic. Long-horizon automation definitely isn't impossible, but it requires a level of per-step reliability and error recovery that current systems maybe don't have. And if we want social trust and people really using it, at the end of the day, people don't experience average performance of these models. They experience the failures.

    Mostafa Dehghani1:02:22

    If you have your model doing a dumb mistake, the damage in trust that it makes is bigger than the benefit of getting 100 things right, like 100 imperfect things right. So this reliability in these long-horizon tasks is something that we definitely need. Besides, as I said, two more philosophical, high-level things: I would definitely work on the grounding problem and how we can build AI systems that are robust and connected to the physical world. As I said, soon the concept of data, how to enable these models to be very good at self-improvement, becomes, how can I ground these models in the real world?

    Mostafa Dehghani1:03:05

    So this is definitely something that would be the bottleneck of self-improvement if we don't actively think about it. We should definitely move away from these statistical patterns in text and pixels. And the other thing that is maybe related is even thinking about a better definition of intelligence itself, right? And it's a little bit philosophical, but it's definitely a practical question. The whole field and us, we are building more and more something that we haven't really defined. We are trying to make these models smarter and more intelligent, but the definition of intelligence is just so hand-wavy and fuzzy that it's hard to actually measure meaningful progress, which is related to your question about evaluations.

    Mostafa Dehghani1:03:53

    It's good. We have proxies, benchmarks, scores, capabilities, and even vibes, which I find super useful. But at the end of the day, we really need a systematic way of maybe defining intelligence. That is hard. And again, making progress based on what we have right now is good, but at some point that becomes a little bit more important to really pinpoint what is the target and what is the goal, and then push toward that with maximum speed.

    Matt Turck1:04:05

    All right, Mostafa, it's been an absolutely fantastic conversation. Thank you so much for spending time with us. Really enjoyed it. Really appreciate it.

    Mostafa Dehghani1:04:11

    Thank you. Yeah, thank you so much for having me. It was fun to chat, and thanks for the invite.

    Matt Turck1:04:31

    Subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.