Yeah, maybe it becomes a self-fulfilling prophecy, right? If you have enough smart people that decide it's the thing, then maybe it does happen. All right, so maybe to close the conversation, let's talk about your work and how you do your work. So you published a book this past year, in 2025, on how to build an LLM from scratch. I believe that you are writing the sequel currently, How to Build Reasoning Models from Scratch. Is that the title?
And you produce an incredible amount of work. So people should find you on your website, on your Substack newsletter. How do you absorb all of this knowledge? To which extent are LLMs part of your workflow? I'm just curious how you work these days.
Good question. So I must say, I don't have a magic approach or anything. I think the thing I have maybe is I get very excited about things. And then when I'm excited about something, it goes very easily and very fast. I don't know. If you notice, maybe I write only about certain topics. I don't cover image models at the moment, for example, because I am just very excited about LLMs. And then, I don't know, I just can't help it.
I get very excited, read all the things about it, write about it, and that's mainly it. I almost go by intuition, basically. And I'm kind of lucky in that sense that with my blog, what I find interesting, other people also right now find interesting. So I think there's a lucky coincidence that I honestly write only about things I find interesting. So I'm not trying to force myself, oh, I have to cover XYZ because it's something that should be covered.
It's more like, oh, how does this, let's say, recursive language model work? Let's just read the paper and then I write about it. More like getting excited about things. Yeah, the book writing is also a bit different because my blog is more research-paper-focused, where when I get excited about something, I read about it and put it in there. For the book, I'm similarly excited, but that's more like a coding book where it's like the fundamentals.
Because I think that's, to be honest, the best way to understand something: to see it actually working. It's not any hand-wavy figures. I mean, there are a lot of figures. Like right now for chapter six, I just finished it the other day, 21 figures. They take the most work. Maybe one day LLMs can help me with that. But figures help to explain the code and everything. But code basically doesn't lie. It either works or it doesn't work, and I think that's a very useful way to learn also.
And it's just also a lot of fun for me. When I write code and I see it working, it's very satisfying, and you have something that actually works. So I should say I'm not building, in that book, any production-level systems. It's called Build an LLM from Scratch, or build, I mean, large language model from scratch. But it's not like an LLM that you would use in production. Well, if you make a few tweaks, you could use it in production.
But the goal is code readability and teaching LLMs, basically, because I think to see actually, how do I format my training data? How does it get processed? What is the loss function? What gets updated? I think this explains so much more than if I say, oh, so it does next-token prediction, and then hand-wavy, hand-wavy here and there. You can actually literally see how it does that and what feeds in. And the same with RLVR. We had hopefully not too bad an explanation in this podcast at the beginning of GRPO.
But if you actually see the diagram, I have the numbers in there, but the numbers could be wrong, though, if you have a figure and you draw arrows and everything. But then if you implement that in code and you get exactly the same results and the model trains and you get 50% accuracy, it's a nice thing where you can say, oh, it's actually working. It's not just made-up numbers. It actually works. And also that's how I learn about LLM architecture.
So I have this blog, The Big LLM Architecture Comparison, with now, I think, 13,000 words because I keep extending it. I read the paper, I look at the architecture of an LLM. Do I really understand it? So I draw the architecture. But then do I really understand it? And often, unless it's like a 1 trillion parameter model, I often code the model. And so I have still my GPT-2 architecture, and they are relatively similar. And so if there's a new architecture, I take the most similar one and make a few changes to it.
But then the beautiful thing here is also someone already implemented that in Hugging Face, the Transformers architecture. So I have a reference model I can run. I run my model and I can see, do I get the exact same logits, the same numbers, if I have the same prompt? And so with that, you can actually self-check yourself. Did I implement everything correctly? Are the results correct? And I think that's just a lot of fun. It's a lot of work, but it's a lot of fun.
And it doesn't lie. It gives you the correct answer. Perfect. On some prompts, someone else extended that and I found a bug. And now I have a better understanding of how they implemented the YaRN scaling. That is something you would never understand by just reading the paper. You have to really, I don't know, look at the code and toy around with that. And so, yeah, that's basically how I try to work. I try to combine reading and coding, and LLMs I also use, but I try to use it.