All right, Dan, what is your perspective on all of this? I really appreciated Tim's post because I think one thing that I really appreciated is that there's some AGI talk that, if you just trace the exponential, at some point you get the thing that will eat up the universe or whatever, which I always found a little bit odd to think that way. I appreciate the thing in terms of the actual physical constraints because, like Tim said, these are physical systems with physical inputs and actually doing physical computation.
I think my perspective was that if you look at where the systems are today and you look at the models that we've trained, we are just so far from even using the last generation of hardware as efficiently as possible, not to mention all the new hardware that's being built out. On the technical side, I'd say there are two major points I wanted to make in my post. One, if you look at the models that are the really great ones that we know today.
In my blog post, I mostly talked about open-source models because they talk a little bit more about how they train and the resources behind it. We don't have public figures behind how much OpenAI and Anthropic are using. But if you look at the DeepSeek model, for instance, this is one of the best open-source models we have out there today. It was trained at the end of 2024 on last-generation, kind of nerfed GPUs, H800s instead of H100s. The H800 is nerfed in all sorts of ways from NVIDIA to get around the export restrictions at the time.
They were trained with about 2,000 H800s, according to the report, for about a month. When you compute how long that took, when you see how much compute was actually available on the chip, you get something like a 20% effective chip utilization or something like that. The term of art is called MFU, model FLOP utilization. Basically, that's a 20% utilization number. Meanwhile, earlier in the 2020s, we were seeing lots of training runs on older hardware with different model architectures that were easily achieving 50–60% MFU.
If you just take that number and then say, hey, maybe there's a way to get it out there. Since then, my good friend Tri Dao has released a whole new set of kernels on how to train these models better. And you say, okay, there's a 3x there just from that one piece. Then the other thing to realize is that that is a model that is being used today in early 2026 as the base for some of the best open-source or open-source-adjacent models out there.
It would have started training the base model at least a year and a half ago. So let's call it mid-2024. Since then, we've started building out completely new clusters with the current generation of hardware. On NVIDIA, these are the Blackwells. There are companies like Poolside that are building out tens of thousands of B200, GB200 chips. There are other folks like Reflection who are building out tens of thousands of B200 chips. This is comparing: we have a new generation of hardware where, even if you take the exact same precision as you had before, exact same everything, 2 to 3x faster compute, 10x larger clusters, plus maybe 3x lurking in terms of just pure optimization.
That's 3 times 3 times maybe another 10. That's another 90x of compute available. And that's not even looking at future buildouts. That is literally clusters that you can point to today that people have started training on that you might hope that at the end of that you'll get much better models. The point I really wanted to make was: if you just look at it from those basic inputs, you can look around, you can squint a little bit, you can see up to two orders of magnitude more compute available compared to the models that we are indexing on today.
Now we can argue about, are there going to be diminishing returns in terms of scaling up? Do we expect the scaling curves to hold and all that? But you can just look around and see it. That's 100x more compute. I think from the physical, just a pure compute perspective, there's a lot more available, a lot more that we're not doing. This is not even to mention a bunch of the great points that you mentioned, Tim. These are all 8-bit training runs.
We've just started writing the papers about how to do a 4-bit training run properly. There's new things like, on the GB200, you have 72 of these connected really quickly. I don't think we've even seen the first pre-trained model come out of that yet. GPT-5, I think, was the first time that you saw in one of OpenAI's reports, hey, this was trained on H100, H200, and GB200, which to me suggests that was actually pre-trained on one of the really old clusters. Maybe some fine-tuning was done on the new GB200s.
You make the point that not only is the hardware underutilized, but you also say that the models themselves are a lagging indicator. The models that we see today that we can play with today have been pre-trained on clusters that were built out a year or two ago because you need enough time to get the cluster running, you need enough time to do the large pre-training run, and then you need enough time to really post-train it, fine-tune it, do all the RLHF and all that stuff.
So the models that we have a snapshot of today that, at the beginning of a conversation, say maybe it is AGI, maybe it isn't, are already trained on clusters that are a year and a half old. We've built out much larger clusters since then. You can expect that they're going to use them for pre-training. The models that we see today, that we index on quality today, are actually trained on pretty old hardware, and we've got new generations of hardware, more software choices we can make, not to mention architectural choices.
Tim, you were mentioning this thing about, you need to move data and then compute on data. We've actually seen the transformer change in architecture a little bit slowly for researchers, a little bit slowly for my taste, but you've seen the fundamental way we do the computation change. Five x or two x there, now you're talking 100x, 150x more compute. There's a lot more compute out there to train better, higher-quality models. If I understand this whole discussion correctly, all of this is about pre-training, right?