MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    DeepMind Gemini 3 Lead: What Comes After "Infinite Data"

    Sebastian Borgeaud is the Pre-training Lead at Google DeepMind. We cover why Gemini 3 improves through many integrated changes rather than one breakthrough, how pre-training shifts from an effectively unlimited-data regime to a finite-data one, and why evals must predict both scale-up performance and post-training results while avoiding benchmark contamination.

    12/18/2025

    Hosted by Matt Turck · with Sebastian Borgeaud, Pre-training Lead, Google DeepMind

    Gemini 3AI pre-trainingscaling lawssynthetic datamodel evals
    Listen now
    YouTubeApple PodcastsSpotify
    55 min · 28 chapters
    Contents

    Transcript

    Cold intro: “We’re ahead of schedule” + AI is now a system

    0:00
    Matt Turck0:57

    Sebastian, welcome.

    Oriol’s “secret recipe”: better pre- + post-training

    0:58
    Sebastian Borgeaud0:58

    Thank you. Hi, Emmanuel.

    Matt Turck1:31

    So I was hoping to start this conversation with this tweet from Oriol Vinyals, who's the VP of Research and Deep Learning at Google DeepMind, the Gemini co-lead, who said when Gemini 3 came out that the secret behind the model was remarkably simple: better pre-training and better post-training, which, when you think about the leap that Gemini 3 represented over the prior state of the art, sounds remarkably modest. So I was curious about your perspective. Is it as simple in some ways as that?

    Sebastian Borgeaud1:45

    Yeah, I'm not sure it's a big secret, at least from my perspective. This seems quite normal. I think people sometimes have the expectation that from one Gemini version to another, there's a big thing that changes and that really makes a big difference. In my experience, there's maybe one or two of those things that make a larger difference than other things, but it's really a combination of many, many changes and many, many things from a very large team that actually makes Gemini 3 so much better than the previous generations of Gemini.

    Why AI progress still isn’t slowing down

    2:09
    Sebastian Borgeaud2:09

    And I think this is probably a theme that will recur later, but it's really a large team effort that comes together in a release like Gemini 3.

    Matt Turck2:25

    What does that tell us in terms of where we are in AI progress? What sounds from afar as sort of turning some knobs gives us such a leap. What does that mean in terms of what we can expect going forward?

    Sebastian Borgeaud2:50

    There's two things. The first one is it's still remarkable how much progress we're able to achieve in this way. It's not slowing down. There's so many of these knobs and so many improvements that we find on a day-to-day basis, almost on a day-to-day basis, that make the model better. So that's the first point. The second point is we're not really building a model anymore. I think we're really building a system at this point. People have sometimes this view that we're just training a neural network architecture and that's it, but it's really the entire system around the network as well that we're building collectively.

    Are models actually getting smarter?

    3:04
    Sebastian Borgeaud3:04

    And so that's the second part.

    Matt Turck3:35

    The big question on everybody's mind is: what does that mean in terms of actual progress towards intelligence? And we don't necessarily need to go into the whole AGI thing, because who knows what that means. But is the right way to think about this kind of model progress as an actual path towards intelligence, versus trying to succeed on this benchmark or that other benchmark? What gives you confidence that the core model is getting smarter?

    Sebastian Borgeaud3:56

    The benchmarks definitely keep improving. And if you look at the prompts and how the benchmarks are set up, they are becoming increasingly difficult. And even for me, who has a background in computer science, some of the questions the model answers would take me a significant amount of time to answer. This is just one view. It's the benchmark view. And we evaluate those frequently, et cetera. We're being very careful about holding out the test set, but still, there's often some fear of overfitting to those and just benchmarking.

    Sebastian Borgeaud4:19

    is what people call this. But that's one aspect. I don't think those fears are very well-founded. But the second aspect, and that's the one that really fills me with confidence, is the amount of time people spend using the model to make themselves more productive internally is increasing over time. Every new generation of models, it's pretty clear the model can do new things and help us in our research and our day-to-day engineering work much more so than the previous generation of models.

    Two–three years out: what changes first?

    4:36
    Sebastian Borgeaud4:36

    So that aspect should give us confidence as well that the models are becoming more capable and actually are doing very useful things as well.

    Matt Turck4:56

    I'm always curious, as an AI researcher who's so deep into the very heart of all of this, if you zoom out, are you still surprised by where we are? From your perspective, are we well ahead of where you thought we would be a few years ago? Are we on track? Are we behind, possibly?

    Sebastian Borgeaud5:24

    I think it's easy to say we're on track in hindsight. I think if I'm being honest with myself, I think we're ahead of where I thought we could go. Starting work on LLMs in 2019 or 2020, it's kind of hard to believe the scale of everything we're doing, but also just what the models are capable of doing today. If you kind of looked at scaling laws back then, they were definitely pointing in that direction. And some people really believed those deeply.

    Sebastian Borgeaud5:48

    I'm not sure if I would have bet a lot on that actually materializing and being where we are today. One interesting question that follows from this is, where does that take us if we assume the same kind of progress we've seen in the last five years? I think it's going to be very cool, what's going to happen in the next few years as well.

    Matt Turck6:03

    What do you think on that front? Does that mean AI comes up with novel scientific discovery? AI wins the Nobel Prize? Where do you think we are going in the short term, like the next two to three years?

    Sebastian Borgeaud6:24

    I think, yeah, that's part of it. So on the science side, I think DeepMind historically has done a lot of work, and for sure there's a lot of work in that direction as well. I think we will be able to make some large scientific discoveries in the next few years. That's one side. I think on the other side is in my day-to-day work as well, both research and engineering.

    AI doing AI research: faster, not automated

    6:34
    Matt Turck6:54

    I'm very excited about how we can use those models to make more progress, but also to better understand the systems we're building and develop our own. Yeah, there's this big theme in the industry about automation of AI research and engineering, which, if you extrapolate it, leads into AI 2027 kind of scenarios where there's a discontinuity moment. Just at a very pragmatic level, what does that mean, using AI for your own work today? And what do you think that's going to mean in a couple of years?

    Sebastian Borgeaud7:25

    I think it's not so much about automation, but more about making us go faster and spending more of our time in the research part at a slightly higher level. A lot of the day-to-day work in research on language models is dealing with quite complex and large systems at the infrastructure level. So actually, quite a bit of time is dedicated to running experiments, babysitting experiments, analyzing a lot of data, collecting results. And then the interesting part is forming hypotheses and designing new experiments.

    Frontier labs: same playbook or different bets?

    7:45
    Sebastian Borgeaud7:45

    And so the last two parts, I think, are something where we'll be very much involved. The first part, I think, especially in the next year, with more agentic workflows being enabled more and more, should be able to really accelerate our work there.

    Matt Turck8:20

    Is your sentiment that the various frontier AI labs are effectively all working in the same direction, sort of doing the same thing? One fantastic but, in some way, perplexing thing that we all experience as industry participant-observers is this obvious phenomenon of every week or every other week or every month, there seems to be another fantastic model, and we're completely spoiled. So Gemini 3 just came out at the same time, like two hours ago, literally before we were recording this. Claude 2 came out.

    Matt Turck8:37

    What do you make of that from your perspective, and how do you think that plays out? Is anybody going to break out, or effectively is the industry going to continue with the handful of top labs, plus some newer labs that are appearing?

    Sebastian Borgeaud9:06

    For the first question, there's definitely similarities between what the different labs work on. I think the base technologies are kind of similar. I'd be surprised if we weren't all training transformer-like models, for example, in terms of the architecture side. But then there's definitely specialization, I think, happening on top of that, and different, maybe, tree branches in the tree of research that are being explored and exploited by the different companies. I think historically, for example, DeepMind has—and still, I think, on the vision and multimodal side, we've been actually really, really strong.

    Sebastian Borgeaud9:38

    And that continues to be the case today and shows in both how people use the model, but also in the benchmarks, of course. And then other things like reasoning, et cetera—OpenAI came up with the first model, but we also had a strand of research on that. So there's similarities, but it's not exactly the same, I would say. For the second question, I don't know if I have a good answer. One thing that's clear is, to make progress on a model like Gemini today, you do need a very large team and a lot of resources.

    Sebastian Borgeaud9:58

    Now, that doesn't necessarily mean that what we're doing today is optimal in any form, and some disruptive research could definitely come along and allow a smaller team to actually take over in some form. This is one of the reasons why I actually enjoy being at Google so much as well: Google has this history of doing more explorative research and has a really high breadth of that research, and that continues to be the case, mostly in parallel to Gemini. But we're definitely able to also utilize that and bring some of those advances into Gemini.

    Post-transformers: will a disruption happen?

    10:19
    Matt Turck10:37

    Are there other groups, whether at DeepMind or elsewhere in the industry, that are working in semi-secret or complete secret on post-Transformer architectures, where one day something will come out and we'll all be surprised? Are there groups like that in the industry?

    DeepMind’s advantage: research × engineering × infra

    10:51
    Sebastian Borgeaud10:51

    I believe so. There's groups doing research on the model architecture side for sure within Google and within DeepMind. Whether that research will pan out, it's hard to say, right? It is research. So very few research ideas work out.

    Matt Turck11:22

    And so, in the meantime, the core advantage that one company may have over the other is just the quality of people. In the case of Google, I guess the vertical integration. That tweet from Oriol that I was mentioning got quote-tweeted by Demis Hassabis, and he was saying that the real secret was a combination of research and engineering and infra. So is that the secret sauce at Google, the fact that you guys do the whole stack?

    Sebastian Borgeaud11:46

    It definitely helps. I think it's an important part. Research versus engineering is also interesting. I think over time, that boundary has blurred quite a lot because we're working on these very large systems now. Research really looks like engineering and vice versa. And I think that's a mindset that has really evolved over the last few years at DeepMind, especially where maybe there was a bit more of the traditional research mindset before. And now with Gemini, it's really more about research engineering.

    Sebastian Borgeaud12:01

    The infrastructure part is also very important. We are building these super complex systems. So having infrastructure that's reliable, that works, that's scalable, is key in terms of not slowing the research engineering down.

    What a Gemini 3 pre-training lead actually does

    12:26
    Matt Turck12:26

    And Gemini 3 was trained on TPUs, right? Not on NVIDIA chips. So it's truly integrated. Okay, so I'd love to do a deep dive on Gemini 3, but before we do that, let's talk about you a little bit. So you are the pre-training lead on Gemini 3. What does that mean? And then let's go into your background and your story.

    Sebastian Borgeaud12:54

    I am one of the Gemini pre-training leads. So what this entails, it's a mix of different things. So part of my job is actual research, so trying to make the models better. But these days, it's less running experiments myself, but helping design experiments and then reviewing results with people on the team. So that's the first part. The second part, which is quite fun, is more of the coordination and integration. So it's a fairly large team at this point. It's a bit hard to quantify exactly, but maybe 150, 200 people that work day-to-day on the pre-training side between data, model, infrastructure, evals.

    Sebastian Borgeaud13:29

    And so coordinating the work of all of these people into something that we can build together is actually quite complicated and takes quite a bit of time, especially time to do well. To me, this is super important because actually being able to get progress out of everyone is really what makes us make the most progress, rather than enabling maybe one or two or a small group of 10 people to run ahead of everyone else. That might work for a short period of time, but over longer periods of time, what's really been successful for us is being able to integrate the work from many, many people.

    From Europe to Cambridge to DeepMind

    13:59
    Matt Turck13:59

    So in terms of your personal background, I'm always curious: where did you grow up? What kind of kid and teenager were you? I'm always trying to reverse engineer those top AI researchers. Where do they come from, and how did you become who you are?

    Sebastian Borgeaud14:21

    I grew up a bit all over the place in Europe. I moved around quite a bit. So I was actually born in the Netherlands, and I moved when I was seven to Switzerland. So my dad is from Switzerland and my mom is from Germany. So I did most of my school and the beginning of my high school in Switzerland, mostly in French and also in German in parts. And then at age 15, I think, I moved to Italy, where I finished high school until I was around 19.

    Sebastian Borgeaud14:53

    And at that point, I was going to go to ETH Zurich to do my studies, but I think just by random events, one morning I just looked up the top universities in some kind of ranking, and I saw Cambridge was at the top. So I thought, I'll just apply. Why not? And, yeah, a few months later I got the acceptance letter. So I decided to move to Cambridge, where I did my undergrad and master's in the Computer Lab.

    Matt Turck15:01

    And when you were growing up, were you just a super math-strong kind of kid, computer science kind of kid?

    Sebastian Borgeaud15:32

    My dad has a technical background. So I remember, when I was 10 or 11, starting to program a bit with him and learning. And I kind of always liked that. And then I always had, like, easiness in math and science at school. I remember never having to really study for math exams, but always doing quite well. That definitely changed at university. But that was my high school experience.

    Matt Turck15:36

    And what was your path from school into where you are today?

    Sebastian Borgeaud15:58

    Yeah, so that's, again, a bit of a lucky moment, I would say. One of the lecturers we had in my master's was someone who was also a researcher at DeepMind. And I just remember at the end of the last lecture, I was packing my stuff, and I was like, oh, I'll just ask him for a referral. What's the risk, right? He might just say no, but whatever. And so I actually took the courage and I went up to him and asked if he would give me a referral.

    Sebastian Borgeaud16:18

    He was like, sure, send me your CV and I'll see what I can do. And that's kind of how I got my interview at DeepMind. This was in 2018. And so I joined DeepMind at the time, just DeepMind, not Google DeepMind, as a research engineer after university.

    Matt Turck16:25

    And what did you do at first, and how did that evolve to being one of the pre-training leads on Gemini 3?

    Sebastian Borgeaud16:54

    Yeah, so at the beginning, having joined DeepMind and DeepMind being known for RL, the first project I managed to work on, or decided to work on, was something on the RL side. So specifically, we were training some unsupervised network to learn keypoints on Atari environments and try to get the agent to play Atari. So I did this for about six months, maybe. It wasn't enough, or in the sense I didn't like the synthetic aspect of this. I always wanted to work more on real-world data and have more of a real-world effect.

    Sebastian Borgeaud17:30

    I think in general, I like to build things and build things that work. I don't really like the academic pure research part. And so that kind of drove me to start working on representation, so creating these, or training these neural networks that have good representations to do different tasks. And one funny anecdote here is something I tell a lot of the people on my team, but the first effort I joined on this was called Representation Learning from Real-World Data. And at the time, we had to add this "from real-world data" to the name of the project because people would assume otherwise it would be synthetic environments or synthetic data.

    Why he left RL for real-world data

    18:06
    Sebastian Borgeaud18:07

    And that definitely has shifted completely since then. So yeah, that was kind of my first project on that side, and specifically LLMs and transformers. We were looking at architectures like Transformers and models like BERT and XLNet that were learning these representations and trying to improve those representations and do research on that side.

    Matt Turck18:11

    Great. And then you worked on RETRO, right? Do you want to talk about that?

    Sebastian Borgeaud18:34

    Yeah. So after that, we started working on scaling up LLMs and LLMs in general. So we started this work first on Gopher, which is, I think, the first DeepMind LLM paper that was published. So already at that point, it was a team of maybe 10, 12 people. So already at that point, it was pretty clear you couldn't just do that research on your own. And this is really where I started doing pre-training and pre-training at scale, and developed my research taste, but also what I enjoy about this.

    Sebastian Borgeaud19:09

    So we trained the first dense Transformer model. I think it was 280 billion parameters, I think 300 billion tokens at that time, and trained that. We definitely would not do things like we were doing them back in the day, but it was great and a very fun learning experience. After that, there were kind of two projects that emerged. The first one was Chinchilla and the second one RETRO. So in Chinchilla, we were reexamining how you should scale the model size and how you should scale the data, especially from a training compute-optimal perspective.

    Sebastian Borgeaud19:48

    So the question is, you have a fixed amount of training compute: how do you train the best possible model? Should you increase your model size or should you increase your data size? And there was some previous work in this domain from OpenAI specifically that we reexamined, and we actually found that you want to scale the data side much more quickly than what was thought before, rather than scaling the model side. Funnily enough, this is still really relevant in our day-to-day work today, especially because it has a lot of implications on the serving cost and how expensive it is to use the models once they're trained.

    Matt Turck20:01

    So that was one side.

    “Research taste”: integrate or slow everyone down

    20:28
    Sebastian Borgeaud20:28

    The other line of work was more on RETRO, and this is more on the architectural innovation side of things. So here we were looking at how you can improve models by giving them the ability to retrieve from a large corpus of text. So rather than having the model learn and store all the knowledge in its parameters, you give the model the ability to look up specific things during training, but also during inference.

    Matt Turck20:40

    You used the word research taste, which I think is super interesting. What does that mean? How would you define that? And how important is that for a researcher?

    Sebastian Borgeaud21:05

    Yeah, it's very important these days. And it's quite hard to quantify. But the first thing that matters, maybe, is that your research is not standalone. This is what I was mentioning before, but your research has to play well with everyone else's research and has to integrate, right? So let's say I have some improvement on the model, but it makes the model 5% harder to use for everyone else. This is probably not a good trade-off, right? Because you're going to slow down everyone else and their research, which would then cumulatively slow down the overall research progress.

    Sebastian Borgeaud21:34

    That's the first thing. The second thing is being allergic to complexity, but complexity is quite subjective in terms of what people are familiar with. But still, we have a certain budget of complexity we can use and a certain amount of almost research risk we can accumulate before things go bad. And so being aware of that and managing that is very important. So oftentimes we don't necessarily want to use the best-performing version of a research idea, but we'd rather trade off some of the performance for a slightly lower-complexity version because we think that will allow us to make more and more progress in the future.

    Sebastian Borgeaud21:53

    So these are kind of the main two things, I think, around research taste.

    Matt Turck22:05

    That's fascinating. And then presumably a part of it has to do with having an intuitive sense for what may work and not work, given there's only so much compute you can use. Is that fair?

    Sebastian Borgeaud22:23

    Yeah, definitely. That's also an important part. I think some people have that much more than others, and a lot of experience really helps. But for sure, we are bottlenecked on the research side by compute. If we had a lot more compute, I think we'd make a lot more progress a lot quicker. And so you have to guess, to some extent, which part of the research tree you want to explore, and then within that, what are the right experiments.

    Fixes vs moonshots: how they balance the pipeline

    23:00
    Sebastian Borgeaud23:00

    But then also knowing most research ideas fail, right? And so you need to figure out at what point have I done enough in this direction to know to move on to something else, or should I keep pushing? And then the other interesting thing is, especially in deep learning, a negative result doesn't mean something doesn't work. It means you haven't made it work yet, often. And so being aware of that as well is quite tricky.

    Matt Turck23:17

    Since we're on this topic of research and how to organize a research team to be successful, let's double-click on some of this. So you mentioned trade-offs. Presumably, one kind of trade-off is short-term versus long-term. How does that work? How do you all think about that?

    Sebastian Borgeaud23:38

    This is part of what I spend a lot of time thinking about as well. There's always critical-path things to be done, or like this part of the model needs improving, or we know this part of the model is suboptimal. So we invest quite a lot in just fixing those immediate things. There's a few reasons for that. The first one is we know this will make the model better. So it's a fairly safe bet, but also we know that things that don't look quite good or quite perfect often tend to have issues later, either when you scale up or when the model just becomes more and more powerful.

    Sebastian Borgeaud24:18

    And so actually being very diligent about tackling those and fixing those is really important. So that's kind of the first part. The second part is slightly more exploratory research. So ideas that could land in the next version of Gemini or the version after that, that have maybe a bit bigger effect on the model performance, but aren't quite validated. How we balance these is, I don't think I have a very clear answer. It's also a bit periodical. So when we're doing a scale-up, for example, there's often slightly more exploratory research because there's nothing right now that needs to be fixed in parallel.

    Sebastian Borgeaud24:36

    But just before we are ready to scale up a new architecture or a new model, it's very much like, let's de-risk the last pieces. It's very execution-focused.

    Research vs product pressure (and org structure)

    24:37
    Matt Turck25:06

    How does that work, a little bit in the same vein, the tension between research and product? So, as we were discussing earlier, you all are in this constant race with other labs. And so, is there presumably some pressure, like, "Oh no, we need to have a better score or win IMO," or whatever it is? So, a very pragmatic, immediate product goal versus stuff that we know is going to improve the model over time. How does that work?

    Matt Turck25:12

    I guess it's just a variation of the same theme.

    Sebastian Borgeaud25:37

    This is why I like Google as well. There's actually very little of that, I think, because all of the leadership has a research background. They're very much aware that, yes, to some extent you can force and accelerate specific benchmarks or certain goals, but in the end, the progress and making the research work is really what matters. So I personally, at least on a day-to-day basis, never really feel that pressure.

    Matt Turck25:50

    How is a team at DeepMind organized? Your pre-training team is several hundred people, if I heard correctly. Is there then a post-training team? Is there an alignment team? How does everyone work together?

    Sebastian Borgeaud26:14

    At a super high level, we have a pre-training team and a post-training team. On the pre-training side, we have people working on the model, on the data, the infrastructure, evals as well. Very important. I think people often underestimate the importance of evals research, and it's actually quite hard to do this well. And then, yes, there's a post-training team, and of course there's a large team working on infrastructure.

    Gemini 3 under the hood: MoE in plain English

    26:24
    Matt Turck26:46

    Let's switch tacks a little bit. And as promised, let's go fairly deep into Gemini 3, if you will. So, Gemini 3 under the hood: the architecture, Deep Think, pre-training data scaling, all those good things. So, starting at a high level on the architecture, was there a big architectural decision that explains the difference? And then how would you describe that architecture?

    Sebastian Borgeaud27:16

    At a high level, I don't think the architecture has changed that much compared to the previous one. It's more of what I was saying before, where a few different things come together to give a large improvement. At a high level, though, it's a mixture-of-experts architecture, Transformer-based. So from that perspective, if you squint enough, you will recognize a lot of the original Transformer paper pieces in that.

    Matt Turck27:21

    Can you describe, to make this educational for people, what an MoE architecture is?

    Sebastian Borgeaud27:47

    At a high level, the Transformer kind of has two blocks. So there's an attention block, which is responsible for mixing the information across time, so across different tokens. And then there's the feed-forward block, which is more about giving the memory, but also the compute power for the model to make these inferences. And those operate on a single token at a time, so they operate in parallel. So in the original Transformer architecture, this is just a single hidden-layer neural network.

    Sebastian Borgeaud28:21

    So it's a dense computation where the input gets linearly transformed into a hidden dimension, you apply some activation function, and that one gets linearly transformed again into the output of the dense block. So that's the original paper. And then there's a lot of work before Transformers as well on mixture of experts. And here the idea is you kind of decouple the amount of compute you use with how large the parameter count is. And so you dynamically route, effectively, to which expert you want the computational power to be used on, rather than having that coupled.

    Native multimodality: the hidden costs

    28:30
    Matt Turck28:43

    Gemini is natively multimodal. In practical terms, what does that actually mean for the model to think about text, images, or videos?

    Sebastian Borgeaud29:01

    Yeah, what this means is that there's no specific model trained to handle images and a different model trained to handle audio, a different model trained to handle text. It's the same model, the same neural network that processes all these different modalities together.

    Matt Turck29:13

    Presumably there is a cost aspect to this. Does being natively multimodal mean you're more expensive from a token perspective?

    Sebastian Borgeaud29:43

    Yeah, this is a really good question. There's kind of two costs to this. I would say that the benefits largely outweigh the costs here, and this is why we train these models. But the first cost is maybe less obvious to people, but it's this complexity cost and this research bit I was talking about, because you're doing a lot more things, and especially different modalities interact in some ways. This can interact with different parts of the research and has a complexity cost. So we have to spend time thinking about these things.

    Scaling laws aren’t dead (but scale isn’t everything)

    30:03
    Sebastian Borgeaud30:03

    The second cost is, yes, images are often larger in terms of input size than pure text. And so the actual computational cost, if you do it naively, is higher. But of course, then there's interesting research to be done on how you make these things efficient.

    Matt Turck30:42

    All right, let's talk about pre-training, since it's the area that you cover in particular. So, starting with the high-level question, we mentioned, of course, the term scaling laws towards the beginning of this conversation. And we talked about Chinchilla a few minutes ago as well. In 2025, there was this much-discussed theme of the death of scaling laws, particularly for pre-training. Is Gemini 3 the answer that shows that all of this is not true and that, indeed, the scaling laws are continuing?

    Sebastian Borgeaud31:09

    Yeah, the discussions there, to me, were always slightly strange because my experience didn't match those. I think what we've seen is scale is a very important aspect in pre-training specifically, and how we make models better. I think what's been the case, though, is that people overvalued that aspect. So it is a very important aspect, but it's not the only aspect. So scale will help to make your model better. And what's nice about scale, it does so fairly predictably.

    Sebastian Borgeaud31:37

    And that's kind of what the scaling laws tell us, is as you scale the model, how much better will the model actually be? But this is only one part. The other parts are architecture and data innovation. These also play a really important part in the performance of pre-training, and probably even more so than pure scale these days. But scaling is still an important factor as well.

    Matt Turck32:01

    Right. And we're talking about pre-training specifically, right? Because this year we seem to have scaled RL in post-training and scaled test-time compute and all the things. But for pre-training, you're seeing not only scaling laws not slowing down, but you're seeing some acceleration. Do I understand this correctly, due to data and different architectures?

    Sebastian Borgeaud32:23

    I think the way to put this is these all compound. So scale is one axis, but the model and data also will make the actual performance better. And yes, sometimes the innovation part outweighs the benefits of scaling more. And sometimes just raw scaling is the right answer to make the model better. So that's on the pre-training side. And yes, on the RL and RL scaling side, I think we're seeing a lot of the same things we're seeing in pre-training, or we saw in pre-training.

    Sebastian Borgeaud32:43

    What's interesting here is because we have the experience of pre-training, a lot of the lessons apply, and we can reapply some of that knowledge to RL scaling as well.

    Matt Turck32:57

    Speaking of data, so what is the pre-training data mix on Gemini 3? I think you guys had a model card out for a bit that talked about some of this. So what went into it?

    Synthetic data: powerful, dangerous

    33:07
    Sebastian Borgeaud33:07

    Yeah, it's a mix of different things. So the data is multimodal from the ground up, and there's many different sources that go into this.

    Matt Turck33:36

    Another classic question in this whole discussion is: are we about to run out of data? So there's always: do we not have enough compute? And the other question is: do we not have enough data? Clearly, there's been a rise in the usage of synthetic data this year. In your day-to-day work, or perhaps in general, where do you think synthetic data helps, and where does it not help?

    Sebastian Borgeaud33:56

    Yeah, so synthetic data is interesting. You have to be very careful in how you use it because it's quite easy to use it in the wrong way. And what's often the case as well with synthetic data is you use a strong model to generate the synthetic data, and then you run smaller-scale ablations to validate the effect of the synthetic data. But one of the really interesting questions is: can you actually generate synthetic data to make a model that you want to train in the future, which will actually be better than the model that generated the synthetic data in the first place?

    Sebastian Borgeaud34:26

    Can you actually make that one better as well? And so we spend a lot of time thinking about this and doing research in this direction. The other part of your question, “Are we running out of data?” I don't think so. So there's more. We are definitely working on that as well. But more than that, I think what might be happening instead is kind of a shift in paradigm where before we were kind of scaling in the data-unlimited regime, where data would scale as much as you would like.

    Reasoning traces: what he can’t say (and why)

    35:00
    Sebastian Borgeaud35:00

    And we're kind of shifting more to a data-limited regime, which actually changes a lot of the research and how we think about problems. But one good analogy of this is before LLMs, a lot of people were working on ImageNet and other benchmarks, and there was a very data-limited regime as well. So a lot of techniques from that time start to become interesting as well.

    Matt Turck35:29

    And perhaps that's one of those, and I don't know to which extent you can talk about it, if not talk about it in general, but there is this concept throughout the industry of training models based on reasoning traces. So basically forcing the model to show its work, how it got to a certain outcome, and then taking that to train the next model. Is that something that you do or that you think is interesting or a future direction? What is your perspective?

    Sebastian Borgeaud35:32

    Yeah, unfortunately, I can't comment on the specifics.

    Matt Turck35:39

    This is how I know I'm asking the right questions. But maybe in general, is that something that people in the industry do?

    Sebastian Borgeaud35:48

    I believe so. And this also falls into the previous question around synthetic data you were asking, and our approach to that is similar.

    Matt Turck36:13

    And perhaps with that, taking this into a futuristic conversation, another big question and theme seems to be indeed how can models learn from less data, which I think is what you were alluding to, talking about a data-limited regime. Again, whether at DeepMind or in general, are you seeing interesting approaches?

    Sebastian Borgeaud36:35

    To use the famous analogy, a model can learn like a child. Just to maybe clarify what I said earlier, in the data-limited regime, I didn't necessarily mean with less data, but rather with a finite amount of data. So the paradigm shift is more from, like, we have infinite data to we have a finite amount of data. The second point is, in some sense, model architecture research is exactly what you mentioned. So when you make an improvement on the model architecture side, what it typically means is you get better results, a better result if you use the same amount of data to train the model.

    Sebastian Borgeaud37:00

    But equivalently, you could get the same result as the previous model by training on less data. So that's kind of the first aspect of that. But it is true, in terms of the volume of data needed today, we're still orders of magnitude higher than what a human has available to. Of course, there's the whole evolution process as well, which I find these high-level discussions quite hard to understand or follow because you have to make so many assumptions to convert that amount of data into what is today's pre-training data.

    Long context + attention: what’s next

    37:18
    Sebastian Borgeaud37:18

    But at least at first order, it does seem like we're using a lot more data than humans do.

    Matt Turck37:25

    What other directions in overall pre-training progress are you excited about throughout the industry?

    Sebastian Borgeaud37:41

    Gemini 3, I think, had a really good leap in the long-context capabilities of the model. And I think that's really enabling the ability of models and agents today to do this work where you have maybe a codebase and you do a lot of work on it. So your context length really grows. I think there's going to be a lot more innovation on that side in the next year or so to make long context more efficient, but also just to extend the context length of models themselves.

    Sebastian Borgeaud38:14

    So that's on the capabilities front, I think, is something where pre-training specifically has a lot to offer, and it is very interesting. Relatedly, I think for us, at least on the attention side, we've made some really interesting discoveries recently that I think will shape a lot of the research we do in the next few months. And I'm personally very excited about that. Again, I think I want to emphasize the point that I made towards the beginning, but the way things work is it's really a culmination of many different things.

    Retrieval vs RAG vs long context

    38:40
    Sebastian Borgeaud38:44

    So there's a lot of small, medium-sized things that we can already see coming up where I think, wait, we fixed this issue, we fixed this bug. This is interesting research that shows promising things. And all of these things coupled, I think, will drive a lot of the progress. It's interesting thinking about RETRO that we talked about a bit earlier.

    Matt Turck39:21

    You're the co-author of RETRO, which was about efficiency and smaller models doing more. And now you are in the world of Gemini 3, which is like massive amounts of data and training in very long context windows. Do you think that this paradigm of having, again, larger models, large context windows, obviates the need for kind of RAG and search, and that everything gets folded into the model? I mean, obviously there's a corporate data part, but in general, there's some interesting questions here.

    Sebastian Borgeaud39:42

    So first of all, I think RETRO was really about retrieving information rather than storing it, not necessarily about making models smaller. So it's about how we can use the model to do more reasoning already in a pre-training sense of reasoning, rather than just store the knowledge. So this is still very much the aspect today. The interesting part is the iteration cycle, maybe, of pre-training used to be a lot slower than that of post-training until fairly recently.

    Sebastian Borgeaud40:07

    And so making these large changes on the pre-training side is quite costly in terms of risk and how long it takes. And then you have approaches like RAG or search, which you can do during post-training and iterate much more quickly on, which give very strong performance as well. I think deep down, I do believe that the long-term answer is to learn this in a differentiable end-to-end way, which means probably during pre-training, or whatever that looks like in the future, learn to retrieve as part of the training and learn how to do search as part of the larger training.

    Sebastian Borgeaud40:36

    And I think that that's kind of where RL scaling maybe starts that process, but I think there's a lot more to do on the architecture side. But this is something that we'll see in the next few years and not immediately, I would say. The one thing I want to highlight is people often talk about model architecture, and that's definitely one part of what makes pre-training better, but there's other parts as well, like infra and data and evals specifically, that don't always get the same mention.

    Sebastian Borgeaud41:15

    Evals specifically are extremely hard, and it's even harder in pre-training, I would say, because it kind of has these two gaps you need to close. So on the one side, the models we train regularly are much smaller and less powerful than when we scale up. So that means the eval has to be predictive of what the performance will be, has to still work for the large model, and point in the right direction. So it has to be a good proxy on that side.

    Sebastian Borgeaud41:37

    And then there's a second gap as well, which is when we evaluate pre-training models, there's a post-training gap as well. So the way the models get used is they don't just get used after pre-training. There's more training happening after. And so the evals we use in pre-training, or pre-trained models, have to be good proxies of what happens after as well. And so making progress on evals is really important and quite hard, and has also driven a lot of the progress we have in terms of being able to measure what an actual improvement is on the model or on the data side.

    The real boss fight: evals (and contamination)

    41:49
    Matt Turck41:54

    And evals at DeepMind, that's all internally built? Like, you have your own set of evals?

    Sebastian Borgeaud42:17

    Yes, to a large extent, and more and more so, because what we found is that external benchmarks, you can use them for a little while, but very quickly they become contaminated. So they start to be replicated on different forums or different parts of the web. And then if we end up training on those, it's really hard, basically, to detect leaked evals.

    Alignment: pre-training vs post-training

    42:28
    Matt Turck42:41

    So the only way you really have to protect against cheating yourself and thinking you're doing better than you are is by actually creating held-out evals and not really keeping them. In the same vein, is alignment a part of what you all think a lot about at the pre-training level, or is that more of a post-training kind of conversation, or both?

    Sebastian Borgeaud42:52

    It's a majority of post-training, I would say, but there's definitely some parts of it which are relevant to pre-training. I can't go into too many details here, but some parts are relevant to pre-training, and then we do think about that as well.

    Matt Turck43:07

    And at a very simplistic level, I always wonder, again, in the context of Gemini or otherwise, if the corpus is the internet, there's a lot of terrible things on the internet. Is alignment 101 that there's stuff that you just do not include in the model?

    Sebastian Borgeaud43:24

    This is an interesting question, and I don't think I have a definitive answer, but you don't want the model to do these terrible things. So at a fundamental level, you do need the model to know about those things. So you have to train a bit, at least on those, so that it knows what those things are and knows to stay away from those. Otherwise, when a user would mention something terrible, the model wouldn't even know what it's talking about, and it might not be able to say, "This is something terrible."

    Deep Think + agents + “vibe coding”

    43:32
    Matt Turck43:45

    Let's talk about Deep Think, like the thinking model that was released a few days after Gemini 3. So first of all, is that a different model, or is that part of the same model? How should one think about it?

    Sebastian Borgeaud43:48

    I'm not allowed to—I can't comment too much on specific things.

    Matt Turck43:49

    I'm sorry.

    Sebastian Borgeaud43:50

    Okay, fine.

    Matt Turck43:58

    When the model thinks and you wait for 10 seconds or 20 seconds, or whatever time, what happens behind the scenes?

    Sebastian Borgeaud44:30

    Yeah. I think this has been covered quite a bit in some of your previous podcasts as well. It's about generating thoughts. And so, rather than just doing compute in the depth or on the model side, you also do compute and allow the model to think more on the sequence-length side of things. So the model actually starts to form hypotheses, test hypotheses, invoke some tools to validate the hypothesis, do search calls, et cetera, and then, at the end, be able to review the thought process to provide a definitive answer to the user.

    Matt Turck44:40

    The industry has normalized around that paradigm of chain of thought.

    Sebastian Borgeaud44:41

    That's fair, yeah.

    Matt Turck44:48

    Can you talk a little bit about the agentic part of this and Google Antigravity? What do you find interesting about it?

    Sebastian Borgeaud45:06

    Yeah, this is, I guess, what I was mentioning before around my own work especially. I think that's interesting. A lot of the work we do on a day-to-day basis is more execution-based, babysitting experiments, et cetera. And I think this is where I at least see the most impact from those. Bringing it back to the topics of pre-training, I think that the perception and vision side is very important for this because now you're asking models to interact with computer screens.

    Sebastian Borgeaud45:25

    So being able to do screen understanding really, really well is critical. And so that's an important part on the pre-training side, at least.

    Matt Turck45:42

    And in Antigravity, there's a whole vibe-coding aspect, truly vibes, in that you don't even really see what happens when you ask. Is vibes—same question—is that a pre-training thing? Is that just a post-training thing? How do you build vibes into a model?

    Sebastian Borgeaud46:10

    Yeah, this is interesting. I think you can probably ask five different researchers and you'll get five different answers. There's also this notion of a large-model feel. GPT-5 historically had some of this, presumably, where larger models maybe feel differently. I wouldn't put it in these terms specifically, but I think vibes comes down to this, and actually pre-training probably plays a larger role today in some of that, and how the model feels in general, than post-training. I think this is, yeah, this is in general for vibe coding specifically.

    Sebastian Borgeaud46:23

    I think that's maybe more of an RL scaling and post-training thing, where you can actually get quite a lot of data and train the model to do that really well.

    Continual learning: updating models over time

    46:34
    Matt Turck46:59

    So zooming out a little bit, maybe for the last part of this conversation, I'm curious about where things are going in general. There was a key theme discussed at NeurIPS this year around continual learning, and I'm curious about your perspective, especially from a pre-training perspective, right? Because we are in this paradigm where every few months or years, we—and by we I mean you—train a very large new base model. First of all, what is continual learning? And two, how does that impact pre-training if continual learning becomes a thing?

    Sebastian Borgeaud47:30

    Yeah, I guess continual learning is about updating the model with new knowledge as new knowledge is discovered, right? Let's say a new scientific breakthrough is made tomorrow. The base model we trained yesterday wouldn't actually know about it in its pre-training. First, I think a lot of progress has been made on this front in the last few years. I think this is mostly around pre-training, around search: use search tools and then make search calls. Then they would have access to that new information.

    Sebastian Borgeaud47:59

    In some sense, this is also what RETRO, that we talked about, was doing by retrieving data and then trying to externalize the knowledge corpus with the reasoning part. So that's the first part. I think the second part, on the pre-training side specifically, is what I was mentioning about long context as well. And one way of doing this is if you can keep expanding the context window, the model keeps getting more and more information in that context. And so you kind of have this continual learning aspect as part of that.

    Sebastian Borgeaud48:18

    But then, of course, there's more of a paradigm shift. Maybe this is what people discuss: can you change the training algorithm such that you can continuously train them on a stream of data coming from the world, basically?

    Matt Turck48:27

    Beyond continual learning, what do you think is hot slash interesting or intriguing in current research today?

    Sebastian Borgeaud48:53

    Yeah, there's a lot of, again, there's a lot of small things right now that accumulate. So that's kind of the first thought that comes to my mind. And that historically has really driven progress. So I wouldn't just bet against that continuing to drive progress. The things I mentioned before around the long-context architecture and long-context research is one aspect, I think, on the attention mechanism as well on the pre-training side. And then this paradigm shift from infinite data to the limited-data or finite-data regime is something as well, I think, where a lot of things will change and there's a lot of interesting research.

    Sebastian Borgeaud49:25

    That's kind of on the pre-training alone side. The other side, which is quite interesting today, is these models—the amount of people using these models is growing quite rapidly. And so more and more, what we have to think about on the pre-training side as well is how expensive is the model to use, to serve, and have really deployed at a large scale? And what things on the pre-training side specifically can we do to make this model have better quality and maybe be cheaper to serve and consume fewer resources during inference?

    Advice for researchers + founders

    49:35
    Matt Turck49:56

    For any student or PhD student listening to this, if they want to become you in a few years, what problems do you think they should think about or focus on that's not like a year or two out, but more interesting sort of a few years out?

    Sebastian Borgeaud50:07

    One thing that's becoming increasingly important is being able to do research, but being aware of the system side of things. So we are building these fairly complicated systems now. So being able to understand how the stack works all the way down from TPUs to research is kind of a superpower, because then you're able to kind of find these gaps in between different layers that other people weren't necessarily able to see, but also to reason through the implication of your research idea all the way down to the TPU stack.

    Sebastian Borgeaud50:56

    And people that can do that well, I think, have a lot of impact in general. So in terms of specialization, it's really, really thinking about this research, engineering, and systems aspect of model research, and not just the pure model architecture research. That's one. I think personally, I still have a lot of interest in this retrieval research as well that we started with RETRO. And I think it wasn't quite ripe until now, but things are changing. And I just think it's not unreasonable to think in the next few years, something like that might actually become viable for a leading model. And why was it not ripe, and why may that change?

    Sebastian Borgeaud51:28

    I think that's around the complexity side of things I was mentioning, and also the fact that, for the capabilities it brings, you can iterate much more quickly in post-training. So what I was saying with search and post-training data, you can give very similar capabilities to the model in a much simpler way.

    Matt Turck51:48

    And as post-training grows and RL scaling grows as well, do you think there are areas of AI right now that are overinvested in, where there's a disconnect between what makes sense and where the industry is actually going and investing dollars in?

    Sebastian Borgeaud52:03

    I think it's got a lot better. I think maybe two years ago, what I was seeing is people were still trying to very much create specialized models to solve tasks that were maybe within half a year or a year of reach of generalist models. And I think people have caught up to that much more and now kind of believe that for generalist tasks, or tasks which don't require extremely specialized models, trying to use a generalist model—and maybe not the current version, but the next version—might be able to do that.

    Sebastian Borgeaud52:38

    And then, so what that means is research in terms of how you use models and the harness, et cetera, is becoming increasingly important. And also how you make models and these harnesses more robust to making errors and recover from such errors.

    Matt Turck53:07

    Yeah. In that vein, do you have any advice or recommendation for startups? So, seen from the perspective of a founder or the VCs who love them, there is this feeling that the base models are becoming ever so powerful and then trained on multiple datasets. So it used to be the model is able to converse, but now it's able to do financial work and cap tables and that kind of thing, which seems to shrink the area of possibility for startups. Do you have thoughts on that?

    Sebastian Borgeaud53:30

    Yeah, I think so. Maybe look at what models were able to do a year or a year and a half ago, and then look at what models are able to do today and try to extrapolate that. I think the models—the areas where the models are improving—I think will continue to improve. And then there's maybe some areas where there's not been that much progress, and that might be more interesting areas to do research. I don't really have a specific example in mind right now, but that would be the general advice.

    “No end in sight” for progress + closing

    53:35
    Matt Turck53:40

    What are you excited about for the next year or two in terms of your personal journey?

    Sebastian Borgeaud54:00

    What I like very much about my day-to-day is working with many people and being able to learn from a lot of researchers. And that's what drives me to a large extent. Every day I come to work and I talk to really, really brilliant people, and they teach me things that I didn't know before. And so I really like that part of my job. As I was saying multiple times at this point, there are just so many different things that will compound and different things where there's headroom to improve.

    Sebastian Borgeaud54:29

    I'm really, really curious because right now I don't really see an end in sight for that kind of line of work to continue giving us progress. So actually being able to see this through and see how far this can take us is really interesting. But at least for the next year or so, I don't see this slowing down in any way.

    Matt Turck54:35

    Great. Well, that feels like a wonderful place to leave it. Sebastian, thank you so much for being on the pod. Really appreciate it. That was fantastic. Thank you.

    Sebastian Borgeaud54:36

    Thank you, Matt.

    Matt Turck54:57

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.