All right, let's switch to the model part of the discussion. In December, you guys released Mistral 3, which was a big release, still with the MoE architecture, which is at the core of what you guys have been doing. You mentioned efficiency earlier in the conversation. Maybe walk us through the general thinking and approach in a highly competitive world of AI models, both in terms of closed source, but also very much open source and all the Chinese labs. What is it that you guys are trying to do, and how do you position?
Yeah, so we've released Mistral Large 3, which is an MoE. MoEs are really nice systems to train because of the lower amount of FLOPS, which makes us able to push performance a lot more during training. They are not necessarily the best format for on-prem deployment because, as of today, if you want to get the best efficiency out of a mixture-of-experts model, you require a lot of volume because you're looking at deployments across dozens of GPUs, usually. And to justify that amount of GPUs, you need to have the right throughput.
We are training large MoEs to get the best performance with the most efficiency during training. We're also continuing to train dense models at other scales because, depending on the environments in which our clients want to deploy, this might be the more cost-efficient solution. I think both architectures are still valuable on edge as well. Sometimes you just don't have the RAM capacity to deploy something like a sparse mixture of experts, and so going dense is helpful there as well. But yeah, definitely for training, mixture of experts and their lower FLOPS are very interesting.
What is the ultimate goal of the model effort? I mean, clearly you guys are a frontier AI lab, but are you trying to create the best models and solve AGI, or are you trying to be the best open-source model compared to the Chinese labs or whatever open source eventually comes out of the U.S.? What is it that you're trying to do?
We're trying to get the best models that we can, and the models that are most useful for the use cases that we cover in enterprise. And so, typically, with the rise of agentic behavior, one thing that's very important is how you deal with various contexts, how you deal with various documents being added to the input. And so having the capabilities to do architecture iterations, really trying new things in terms of model training, is critical. So we're pushing the boundaries of what the current models can do with the compute capacity that we have, but we're also trying to focus on the things that are most annoying in our deployments today.
And so one of the considerations that has been solved with a few harness tricks is the context of those agentic systems. So it's visible typically in vibe coding, but it's definitely applicable to a lot of other use cases where, through all of the tool calls, you'll have to consolidate and summarize the context to be able to fit everything and have the model focus on the right parts. To me, this is just an artifact of the current architectures. We're trying to fit things in linear context windows where, essentially, the questions that we're asking aren't really necessarily all linear.
And so we rely today on the file system for this. And I think that was the big change and realization through vibe coding, is that agents are good enough at manipulating file systems that they can use this as a replacement for their context window, basically. They can select parts of what they want to read. They can select parts of the tool results. And this minimizes the context-length requirements. This is the state today. I think we can do much better.
And I think there are a lot of improvements to be done on those types of questions.
Do your agents run on sandboxes?
It depends on the types of agents, but the answer would be yes. If it's coding agents, usually we have sandboxes that will let the agent iterate and run. I think the depth of the isolation will depend on the use case. Typically, if the file system is just representing textual context and you're not expecting the agent to do much action on it, then you don't really need a full sandbox. You just need some representation of that context as a file system, and it can be any sort of abstraction.
But if you are, I don't know, typically running asynchronous code development, then yes, you need a sandbox.
Great. What is the current constraint that you guys are facing to make Mistral 4, when it eventually comes out, do much better than Mistral 3? Is that a question of more compute, or is that a question of data? And in particular, are you guys doing anything around synthetic data that you can talk about?
Definitely compute, and the current deployment that we have will help, as it's going to be giving us a lot more Grace Blackwell capacity than we had in the past. And so that's something that we're very excited about. And when you add compute, you also have to add data. And so we've been hard at work making sure that our data mixtures are as high quality as ever and growing in size. But as you mentioned, one of the ways to do this is through synthetic data.
In terms of where we use synthetic data the most, I think a lot of the interesting work that's happening is for the post-training part, where we can build environments that look similar to an enterprise and then try to synthetically create queries that are hard and that will require multiple hops. And so all of this work, in addition to the coding work, the reasoning work, is really what makes the final model able to perform in the various environments that we work in.