MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Jeremy Howard on Building 5,000 AI Products with 14 People (Answer AI Deep-Dive)

    Jeremy Howard is the Co-founder at Answer AI. We cover why Qwen and DeepSeek offer production flexibility commercial models cannot match, why autonomous coding agents like Devin produce slower and less maintainable software than shared human-AI workflows, and how a 12-to-14-person team aims to build thousands of AI products without scaling headcount.

    05/15/2025

    Hosted by Matt Turck · with Jeremy Howard, Co-founder, Answer AI

    Open-source AIAI agentsAnswer AIHuman-AI collaborationAI product development
    Listen now
    YouTubeApple PodcastsSpotify
    55 min · 17 chapters
    Contents

    Transcript

    Highlights and takeaways from ICLR Singapore

    1:39
    Matt Turck1:37

    Jeremy, welcome. Thanks for doing this.

    Jeremy Howard1:39

    Thank you, Matt. Nice to see you again.

    Matt Turck1:49

    You're just coming back from ICLR, which was in Singapore this year, a couple of weeks ago as we're recording this. Any highlights for you? What caught your attention?

    Jeremy Howard2:21

    It was actually really interesting that I was there with most of my team, and we all turned up on day one. And for the first time at a conference, I realized I didn't find anything interesting. And then I went and sat at a local cafe and sent a message to my colleagues and said, "Hey, do you want to catch up? I'm here." People started dropping by, and I had extremely interesting conversations with them. I quickly realized, actually, we've gone so far off the academic normal research track now that no one cares about anything we care about, on the whole, and vice versa.

    Current state of open-source AI

    2:39
    Matt Turck2:55

    We're going to go into how you're building a very different kind of research lab and a very different take. But as we do this, one of the key things that you all seem to be advocating for in pretty much your entire career is open source. And I was curious, before we get into the specifics of Answer AI, of maybe your high-level take on open source today in the AI world. If you listen to different people, some people say, well, last year was the year where open-source AI really caught up with commercial closed-source AI.

    Matt Turck3:14

    If you listen to other people, they say, well, actually, the gap probably widened, and you don't see that much open-source actual deployment. What do you make of it from your perspective?

    Thoughts on Microsoft Phi and open source moves

    3:45
    Jeremy Howard3:45

    Well, just judging by the actions that we've chosen, we're very well funded, so on the whole our choices at this stage are not very cost-driven. Two of the three models that we rely on in our day-to-day production stuff are open source, being Qwen and DeepSeek. I guess we've decided we like them very much. They give us a level of flexibility to tune things to be exactly the way we want, which we can't get from commercial models.

    Matt Turck3:51

    But for the rest of the world—you guys are deep experts—but for the rest of the industry, where do you think we are?

    Jeremy Howard4:05

    I think that's all that matters, though, right? It's always the early adopters that tell you where everybody's going to be in two years' time. So it's also interesting that all the best open source is coming out of China now. So they seem to have a much more kind of pro-social approach to this, whereas, oddly, the U.S. companies that used to lead the way have drawn up the drawbridges, and that's going to cause China to keep moving faster, because when you're in that more collaborative and competitive environment, you just go way ahead, as we've seen in every other area of open-source software over the last 30 years.

    Jeremy Howard4:41

    Yeah, I think it's going to keep going in that direction. I will say the one slight outlier is Google, who've really come out in front. Way out in front.

    Matt Turck4:43

    Gemma 3, or?

    Jeremy Howard4:47

    Gemini 2.5 Pro is just a fantastic commercial—

    Matt Turck4:52

    As a commercial product, yeah. But for open source, Gemma, you found that interesting?

    Jeremy Howard4:56

    Yeah, but they are doing more on the open-source stuff. Gemma's excellent.

    Matt Turck4:58

    And what about Microsoft Phi?

    Jeremy Howard5:13

    It's interesting. We're not using it. We've spoken to the team, the Phi team. We've kind of tried to use it. In the end, it's kind of a bit hamstrung by weird decisions that they made. I hope they keep trying. I think they're in the right direction. For a long time they had terrible licenses, and finally they've got good licenses, but they've also made a lot of kind of bad choices around how they evaluate and portray the models.

    Responding to OpenAI’s open source announcements

    5:41
    Jeremy Howard5:41

    Yeah, they're still not quite at the point that they're that practical, but it feels like they're on a really excellent trajectory. So I hope they keep going the way they're going, but try to really open things up a lot more and really try to challenge the best.

    Matt Turck5:49

    And the fact that OpenAI announced intentions to open-source at least one more model, what do you make of that?

    Jeremy Howard6:14

    I don't really take any account of vaporware announcements. We'll see if and when it happens. Vaporware announcements can only ever be made as a way to try and stay relevant. So I'd rather look at actual things they're doing. So, for example, o3 is an interesting model. I hardly ever use it. One of our team who's been building Triton kernels has found that for some of the things he's doing, it's the only one that's really able to help. Also, their Deep Research tool is still the best.

    The real impact of the Deepseek ‘moment’

    6:29
    Jeremy Howard6:29

    So, it's nice. We've got lots of good products coming out from the open-source world and closed-source world. And this multi-party environment, I think it's great for everybody, actually.

    Matt Turck6:43

    A few months into that big DeepSeek moment that some call the Sputnik moment for AI and all the things, does that still, from an AI community perspective, feel as important today as it did then? Or was that overblown in the press?

    Jeremy Howard7:01

    Well, I never understood it at the time. We'd been using DeepSeek for ages, and they just—I don't know how these things suddenly pop into the public perception. We'd been asking all of our vendors, "Add DeepSeek and Qwen models," and they're all like, "Oh no, we wouldn't do that. They're Chinese," blah blah blah. And then, for whatever reason, one particular one gets picked up by the public at large, and suddenly everybody's like, "Oh, look at us, now we've got one on our serving infrastructure."

    Jeremy Howard7:33

    It's like the very same people that just two weeks earlier were telling me they'd never put a Chinese model on their serving infrastructure. So I found it all really weird. DeepSeek's been great for a long time. They've kept on improving. They'll keep on improving. For me, there was no technology DeepSeek moment. R1 was just another expected journey along their path. And it was interesting to me to see how, again, just some things can break into the public awareness, I guess.

    Matt Turck7:51

    On Twitter/X a few weeks ago, I seem to remember you seemed to indicate that you thought that it could be problematic for OpenAI, I don't know, from a cost perspective or a business model perspective.

    Jeremy Howard8:16

    Yeah, well, they've got too much money. It's like what happened to Google, like, five, 10 years ago. They encouraged their engineers to use as much compute as possible so that they could advertise how much compute they used for these results. Definitely OpenAI has been guilty of that kind of thing: "Look how much money we're raising, look how much compute we're using." The GPT-4.5 debacle, which was just always coming. They kept on kind of exponentially increasing the amount they were spending on their models, whilst the return, the kind of utility of those models, was only increasing logarithmically.

    Jeremy Howard8:52

    And you kind of very quickly hit this point where it's like, "Oh, we made this great model, but basically it's never worth the time and cost to use it." I think that's a really foreign concept for a lot of people at OpenAI, to be like, "Oh, I see, those users actually care about resources. We can't just give them slightly better things and say, now it costs a lot more and you have to wait a lot longer." And it was so unsuccessful, I think they're shutting down that product, or they've shut down that product, if I understand correctly.

    Progress and promise in test-time compute

    9:02
    Matt Turck9:22

    And in terms of what those big labs produce, you just mentioned o3 and Deep Research a minute ago. So, the whole test-time compute approach and all the things, it seems from your perspective as deep experts, like a very promising avenue. Do you see that continuing to improve model performance?

    Jeremy Howard9:44

    Yeah. It's weird that people took so long to care about that. It's another of these things that lesser-known papers have been kind of indicating for a long time, is adding a few tokens to give some breathing room or thinking room or whatever is important. I guess the weird thing is people seem to treat it as being like, oh, it's like part of some kind of exponential curve. And actually, it's like, no, that was a really obvious thing that lots of people knew had to be done, was take advantage of inference-time compute, not just train-time compute.

    Jeremy Howard10:18

    It's not some infinite money supply. It's a thing where you get most of the juice out of it in the first year or two. So we're still in that phase at the moment, and we'll start to hit the fall-off point pretty soon, just like we did for training. But the nice thing is now people are starting to understand on the training side, for instance, that, okay, COGS, cost of goods sold, actually does matter. And people are starting to invest in efficiency again, which is nice to me because that's kind of always been my thing, is I care a lot about efficiency because I care about maximizing the number of people that can use the technology to make their lives better.

    Where we really stand on AGI and ASI

    10:53
    Jeremy Howard10:53

    And when it's expensive, that reduces the number of people a lot. So it's nice to see both people starting to care about costs and companies like DeepSeek showing we really do have lots and lots of opportunity to bring costs down. And when you bring costs down, you often bring performance up as well, like speed up, because you're doing more with less.

    Matt Turck11:09

    You guys made the very clear decision to not be an AGI lab. So if test-time compute is going to start producing decelerating returns at some point, what do you make of the whole AGI, ASI kind of thing? Is that something that you think about a lot?

    Jeremy Howard11:32

    I don't really have a strong opinion about it. I don't think we have any more evidence that ASI might be close now than we did 15 years ago, 15 years before that. At any of those times, you could say, like, oh, ASI might be close. Because it can use natural language, our brains think we're dealing with a different kind of thing. But all that's changed is the user interface to that thing. It's doing the same things it was always doing before, but now we have natural language input, natural language output.

    Jeremy Howard11:55

    It can kind of ingest and calculate with natural language, text, audio. These are really useful things, but I think the impact on our brain of thinking, like, oh, this is closer to superintelligence, is maybe our brains are getting a little bit tricked.

    Matt Turck12:20

    So the emotional aspect is disproportionate, but the underlying raw performance of the models for generative AI have improved, right?

    Jeremy Howard12:32

    Yeah, I mean, they've kept improving. They've been doing stuff with neural nets for a bit over 30 years. They've largely kept improving throughout that time. Obviously, now they're getting a lot more people and investment put into them.

    Matt Turck12:46

    But you dispute, or at least you question, the sort of recent exponential aspect of this. You feel it's mostly sort of a question of perception and emotion, or do you think there's something—

    Jeremy Howard13:23

    No, I mean, there's definitely positive feedback loops with the additional money and people coming into it. So obviously things go harder and faster, and getting to the point that we can use things like synthetic data effectively also has some nice positive feedback loops. So there's definitely some exponentials at play. Like basic economics says, that's the exponential that is at the start of a sigmoid curve, and they're impossible to tell the difference for a while. So we're starting to see that with training already.

    Jeremy Howard13:50

    The amount of people and money being put into AI, it's going to keep going up, but again, the low-hanging fruit's been done now. So yeah, I think we'll continue to see the tailing off of growth. But regardless, none of that means ASI. It really is a continuation of trends we've been seeing for a while.

    Matt Turck14:05

    You mentioned a few minutes ago test-time compute being something that everybody knew about, and you were surprised how long it took to actually get rolled out into a product. Is there a next thing that everybody knows about, and it's surprising that it hasn't been rolled out in product?

    Jeremy Howard14:34

    Not of that size. I think something that we've heard from lots of folks, most notably, I guess, Yann LeCun, is that the autoregressive approach to inference seems dumb. We're starting to see some models with a diffusion feel appearing, and obviously Yann LeCun's got his JEPA-based approaches. But that side feels more complicated. The test-time compute side always seemed pretty straightforward. If in 10 years' time nobody's got JEPA or diffusion or whatever to work in NLP, I wouldn't be like, oh, that must be they didn't try hard enough. Maybe it doesn't work.

    Jeremy’s journey from philosophy to AI

    15:05
    Jeremy Howard15:05

    But I think it probably will work. And I think that'll probably be a significant jump in performance when you can sketch out the entirety of the solution first and then fill in the blocks in gradually increasing levels of specificity, versus only being able to think by emitting tokens.

    Matt Turck15:34

    So boring, but fascinating. Okay, great. You mentioned having done deep learning for 30-plus years. One thing I find fascinating about your story and your resume is that at university you studied philosophy and somehow became a world-class AI scientist. Were you the kind of kid that was just, like, super great at math the whole time, and then you just chose philosophy on the side because it was kind of a different pursuit?

    Jeremy Howard15:55

    The answer to that's kind of complicated. I did come top of my school in math, and my school was kind of the top school for math. So at one level, it's like, okay, there's some data point there that I was good at math. Another data point, though, is, like, I dropped out of first-year math at university because I didn't have any of the background that anybody else had. Like, I didn't have any parents that were academics or teachers, so everybody else who was doing that level of math, everyone had a parent that was helping them.

    Jeremy Howard16:25

    And at high school, I never worked hard at math. I never felt like I liked math very much. I never did particularly well at it. It wasn't until the final year of school where I studied the rubric and figured out exactly how to get the highest possible score, and I did. Yeah, but what I actually got good at was, like, spreadsheets and databases. And so every holiday, pretty much every holiday during the last three years of school, I went and did work experience, always basically working with personal computers because nobody else, none of the older people, knew how to do it because they had grown up on mainframes.

    Jeremy Howard16:46

    So I had some expertise, which was just kind of self-developed coming out of my interest in computer games, really. And so then—

    Matt Turck16:50

    It seems that story of computer games comes up so often. It's sort of so amazing.

    Jeremy Howard17:11

    Yeah. And so then by the time I was in university, yeah, I was able to get a job at McKinsey & Company, and I was working 80 to 100 hours a week, which left no time to actually study philosophy. So I didn't do well at philosophy either. So yeah, it was all self-taught, and it was very easy to be ahead of everybody at that point because nobody else studied it. I felt pretty hopeless for a while because it felt like everybody else knew what they wanted to do at university, but university didn't have a thing for the things I cared about.

    Jeremy Howard17:43

    So I thought the things I cared about must be stupid and pointless. I still wanted to study computer games or something. I wanted to study these fun things that I thought were interesting, and I thought they were going to be really important. And I thought, well, if they were, then all these older, smarter people would be saying that too. So obviously they're not. And so I wanted to figure out where I've got it wrong.

    Matt Turck17:52

    You were at McKinsey and then you were a successful entrepreneur. You started a company called, what was it, Fastmail?

    Jeremy Howard17:57

    Fastmail and Optimal Decisions. I started two companies at the same time.

    Matt Turck18:17

    And then when you and I first met, I think you had just joined Kaggle, but you joined Kaggle as president because you were the top competitor on Kaggle, if I remember correctly, so on one of those world-class data science competitions. So how does a McKinsey person that then becomes an entrepreneur and works on—

    Jeremy Howard18:26

    Well, I was the most shocked person. Yeah. So at that time, Kaggle was just one guy, Anthony. That was his—

    Matt Turck18:26

    Anthony Goldbloom.

    Jeremy Howard18:55

    Little hobby thing. And he didn't have any investment, he didn't have a board, he didn't have anything, but I thought it was super cool. And after I sold my second company, I was so hopeless because I'm just a manager. I've spent all my time telling this group what to do, very insular. So I thought, I should do this. I should try a Kaggle competition because it's going to teach me to actually get a lot better at this. So that's why I joined a Kaggle competition.

    Jeremy Howard19:21

    And I was trying to think, the first one, yeah, the first one I joined, I won it. Like, flabbergasted. All the other people in the competition were professors and PhDs. And yeah, the way I thought of myself was, like, business guy, philosophy major, couldn't do university math and had to drop it. And I thought, like, that's amazing. These professors and PhDs apparently, clearly they haven't learned as much about how to actually solve predictive modeling problems as I did just by spending 20 years solving predictive modeling problems, figuring stuff out as I went.

    Jeremy Howard19:59

    I took it very seriously. I was like, okay, much to my surprise, I'm one of the best in the world at this. So I should do something with that. I had no idea that this self-taught thing that I've spent all this time doing had worked out so well, but it did. And much to my surprise, the academic approach was way less successful. So I thought, oh, that's—yeah, I should make that my life's work now then, making the most of this skill that I seem to have.

    Becoming a Kaggle champion and starting Fast.ai

    20:07
    Matt Turck20:16

    AI, right? I mean, that in some ways sort of feels like a continuum of the same mission of democratization and teaching. Is that fair?

    Jeremy Howard20:53

    So I was the second person at Kaggle and was the chief scientist and president of the American company, not just the one that ended up becoming Kaggle. Kaggle, we created this kind of motto: making data science a sport. And what we wanted to do was to highlight the world's best data scientists, make them rich and make them successful and give them accolades and make people want to be like them. And it worked really well. We worked hard on the PR side.

    Jeremy Howard21:15

    We got top competitors into the press and onto TV. Fast.ai was kind of an extension of that to be like, okay, we want everybody to be able to use these technologies and improve their lives with them and improve other people's lives with them.

    Matt Turck21:36

    Fast.ai being online courses to teach deep learning that have been immensely popular. Is that a fair description?

    Jeremy Howard21:44

    That's one of the four things we did. It's the tip of the iceberg, you see, but it's the kind of smaller part of the work.

    Matt Turck21:44

    Fast.ai?

    Jeremy Howard22:12

    Yeah. So the cycle basically went: we were trying to figure out how to get AI into a lot more people's hands. The cycle was to teach best practices to as many people as we could in as usable and compelling and useful a way as possible. And then, in that process, identify all the places that people couldn't achieve the things that they wanted to achieve because it's too expensive or too slow or it just didn't work or whatever. Then we'd spend months doing research to try to see if we could find ways to get over those problems or find other papers that had already gotten over those problems.

    Jeremy Howard22:44

    Then we'd spend months implementing those things in software. And then a year later, we would do another course with all of those improvements, and the cycle continues. So we did that cycle five or six times. Yeah, it's kind of like, I don't know if you remember the restaurant El Bulli. Ferran Adrià—it was the best restaurant in the world for a while. And they spent six months of the year researching chemical engineering and how foods react, and six months a year serving the things they created.

    Answer.ai mission and unique vision

    23:04
    Jeremy Howard23:04

    So the courses were the things I guess most people interacted with, but that was really us showing the results of all the work that we had done that year and finding out how successful it had been.

    Matt Turck23:23

    Let's talk about Answer.AI. fast.ai is now a part of Answer.AI, but at the beginning, Answer.AI was a separate effort. So Answer.AI is a new kind of AI research lab. And you mentioned that it was partly inspired by Edison's lab. Do you want to talk about the general philosophy?

    Jeremy Howard23:41

    Sure. So it's definitely not a research lab. This might sound minor, but it's an R&D lab. But they feel and look extremely different, as was Edison's Menlo Park lab, which was an R&D lab.

    Matt Turck23:42

    Menlo Park, New Jersey.

    Jeremy Howard24:06

    Yep. The way I like to think of it is, what would Michael Faraday have said just after he figured out how to harness electricity? If you're like, hey, Michael Faraday, I'm a VC. I'm thinking of funding this. Tell me what it's for. He'd have been like, oh, it's energy. It's great. Cool. What's it for? What products will you make? What's the market? He would have had no idea. He didn't know about record players, air conditioners, microwaves, whatever.

    Jeremy Howard24:37

    I think the discussion probably would have been like, oh, well, previously we've seen steam and we use that to kind of make trains and better mills and better weaving machines. So the VC might be like, oh, great, so you're going to improve the unit economics of trains, mills, and weaving machines. Yeah, but I think it might be even more than that. So this is where Edison's team was very successful. He basically created a team of tinkerers to just mess around with electricity and make stuff and see what happened.

    Jeremy Howard25:13

    So there was a lot of development, but there was also a lot of research. So it's like, okay, well, finally we've got a good light bulb. It's not going to be any good unless everybody's got electricity coming to their homes. So now we're going to have to do research to figure out electricity distribution. But in the teams doing the development, they weren't all just developers. And the teams doing research, they weren't all just researchers. They were all a mixture of boilermakers and physicists and engineers, and they all just worked together to figure stuff out in this loop of like, okay, developing this requires this research, which creates this new opportunity to develop this, which also needs this research.

    Jeremy Howard25:33

    And that ended up creating General Electric.

    Matt Turck25:34

    GE.

    Jeremy Howard25:53

    Which was, I don't know, was that like the world's first true global industrial conglomerate? Putting aside the various East India companies, they were really doing something else. They had 4,000 or 5,000 various different products at its heyday, General Electric, all with the common theme that they used electricity. I really like that picture. There are many bad things I could say about Thomas Edison, but there's no question that GE made a lot of products that were valuable to society, and people were prepared to pay more to buy them than it cost to make them.

    Jeremy Howard26:15

    So they became profitable and successful. So we want to do that in AI. So it's very different to a research lab, which is like, hey, let's try and get lots of citations on a paper in a field that's so highly recognizable to its peers that they recognize it as being something they want to appear at their next conference, which really encourages things to be as familiar as possible and evaluated on things that are as similar to other things as possible, and really not be too bold.

    Jeremy Howard27:08

    And to have people who fiddle around with researching stuff, and once they get to a point it kind of looks like it could work, they then throw that away and go and do something else. So yeah, we're the opposite. We're trying to build things that lots of people use, and if there are places where there's constraints we need to get past, then we do the research to get past those constraints, and it's all about results. Yeah.

    Matt Turck27:23

    You mentioned a lot of people making things that a lot of people use. fast.ai is now a part of Answer.AI. I think you guys are structured as a PBC, as a public benefit corporation. So that's part of it, right? Making AI broadly accessible?

    Jeremy Howard27:45

    Yeah. Our charter is to basically create as much societal benefit from AI as possible. That's our job as a company. And my job as a CEO is to help make a company that does that.

    Matt Turck27:48

    A lot of it is open source as well, in the same vein?

    Answer.ai’s business model and early monetization

    28:15
    Jeremy Howard28:15

    Yeah, quite a bit. But we need to make money to be successful at that mission. It's a 20-year-plus mission, but making money would not be sufficient on its own. So we're trying to find this mix of building profitable products that people want to pay us to use. And if there's things that we can give away for free, which there have been, then we are happy to do that as well.

    Matt Turck28:26

    But there's a business model coming up. So you co-founded the company with Eric Ries of Lean Startup fame. You guys are VC-backed. You raised $10 million from Decibel and others.

    Jeremy Howard28:29

    And then another $8 million from angels.

    Matt Turck28:46

    And then another $8 million from angels. And then, so, is there already a business model? Is that something that's coming up? In other words, are you charging for some of the products, or not yet? To rephrase the question, is a business model active today?

    Jeremy Howard29:09

    Slightly. We've done some little test runs. We've made a small amount of money. The main thing, though, we've been working on is more the platform for us to rapidly experiment with AI applications. And actually, that's the thing we've made money from as well, is we've let 1,000 people use that platform as a kind of pre-release beta test. It's an interesting position to be in. It's kind of a bit AWS-like. The platform is built for us.

    Jeremy Howard29:16

    Obviously, it's very useful for other people as well.

    Matt Turck29:16

    Right.

    How a small team at Answer.ai ships so fast

    29:33
    Jeremy Howard29:33

    So that question of how do we think about it or invest in it, we're kind of like, okay, we're primarily going to think about it and invest in it as an internal tool, and we're going to focus on us as customers. But that's quite a good way to create tools that other people like anyway, or at least people who are a bit like you.

    Matt Turck29:49

    Fantastic. So actually, that's a perfect segue. Let's get into that. So from afar, one thing that's truly remarkable is that you guys have an incredible shipping velocity, and you're ultimately a team of 12, last I read. Maybe you're a bit more now.

    Jeremy Howard29:55

    Yeah, I think we're 14 now. We're kind of aiming to be around 12. There were a couple of people we found we couldn't say no to.

    Why Devin AI agent isn't that great

    30:25
    Matt Turck30:27

    So part of it is the team, and we can talk about this in a minute, but part of it is what I think you just described, which is that you built this whole system of dialogue engineering which has superpowered your productivity. So let's get into that. And perhaps as a sort of introduction to this, we're talking about AI-assisted or AI-driven code development and product development. So you guys tested Devin, the autonomous agent, quite thoroughly, and the results were, I think, less than impressive.

    Matt Turck30:43

    So maybe talk about that part, and then we'll talk about how you guys go about it.

    Jeremy Howard31:07

    Sure. I guess we have a number of basic theses underpinning our strategy, and one of them is that humans and computers work better together than separately. And then there are various things that kind of come out of that. And one key one, for example, is that the human and the computer should be able to see all of each other's work and be operating in the same direct environment. Devin is the opposite of that, but we don't assume our theses are correct.

    Jeremy Howard31:37

    We like to explore people using other approaches. Devin's approach is to hand off things to a computer and have it go do it. So, yeah, we were keen to see how well that could work, because certainly in a close human-computer interaction, if there are parts that a computer can just do on its own, you want them to kind of go off and do that and then come back, give the human those bits to work with. So it's always useful for us to know what exactly can a computer do.

    Jeremy Howard31:50

    So Devin's an agent, a bunch of tools, tool calls, tied together with an OpenAI model, if I understand correctly. Yeah, it actually ended up definitely supporting our thesis, which is that there were so many places we wished we could have got involved and could have worked together and been like, "I know how to debug that," or, like, "It's going on wild tangents." It ended up taking a lot longer to do things than we would do them on our own, and ended up with much less high-quality software. And it ended up with kind of vibe-coded stuff that we couldn't really build on top of anyway, which, in a lot of situations, has no value.

    Jeremy Howard32:53

    There are certainly situations where scripts, spreadsheets, VBA, Access databases are useful, but nobody builds them as the components of their technology strategy. They're the things you use for little tools, small, well-defined things. And that's where these kinds of AI agent things seem to be okay sometimes. But Devin, I think, I can't remember the details, but of the 24 things we tried to get her to do—

    Matt Turck33:06

    Yeah, I think I read that you had three or something like that. And I think you guys were saying that part of the issue was not just that it failed, but it failed unpredictably. It was not always failing at the same thing at the same time.

    Jeremy Howard33:09

    Yeah, absolutely.

    The future of autonomous agents in AI development

    33:10
    Matt Turck33:26

    And you think that's a world where that's a temporary issue, a little bit related to the earlier part of the conversation about AGI, ASI, or whatever? Is there a world where, I don't know, the underlying models improve so much that this ends up actually working?

    Jeremy Howard33:45

    I don't really make predictions. I've just always liked to make things based on kind of where we are now, what works now, and what's, in a kind of a very direct extrapolation, where we might be in like a year. That's always worked pretty well so far in my career. That kind of helps understand the details of the technology well. So these things like DeepSeek R1 or whatever, they don't come out of the blue.

    Jeremy Howard34:14

    They don't seem like wild jumps or whatever. They're all part of the same path. Well, obviously, AI models will be able to do larger and larger pieces, and more and more stuff, make less and less mistakes, because they're getting better. Does that mean we'll get to a point where humans have no useful input to provide at all? That's the same as saying, well, we get ASI. And if that happens, then the world's a very different place.

    Jeremy Howard34:34

    There's no point in me doing anything. I have no particular reason to believe that will happen. And for every point until we get there, if we ever get there, by definition there's going to be humans interacting with AI. And so that's what I care about: how do we do that in an optimal way? And I feel the optimal way is to have the humans and the AI in the same environments, sharing that, where they can both then bring the bits that they're best at.

    Dialogue Engineering and Solve It

    34:43
    Matt Turck34:52

    So let's talk about now dialogue engineering, Solve It. So dialogue engineering is the process, Solve It is the platform.

    Jeremy Howard35:26

    This tool we've built for rapidly creating proof of concepts or testing AI applications is called Solve It, and it uses an approach called dialogue engineering. And the basic idea is that, unlike prompt engineering, where you're just creating a single sentence or paragraph or whatever, that's actually part of a whole back-and-forth dialogue. All of the previous steps get sent to the AI model as well, not just the prompt, and they all greatly influence how it responds. And how it responds influences you as to what you then add to the dialogue.

    Jeremy Howard36:04

    We've basically built a system to allow you and the AI to together construct a dialogue, and that dialogue is always editable. So you can delete parts of it or reorder it or turn it into a hierarchy, run code in it, run code on it, ask the AI to run tools over it or to respond to parts of it or whatever. So it's kind of like, imagine bringing together Cursor and Claude Code and ChatGPT and Jupyter Notebooks smushed together. It's kind of got all of that functionality, but then when you do that, you end up with something way more than the sum of its parts.

    Jeremy Howard36:30

    So Eric Ries is using it—he's using it like six hours a day for helping him write his book. And again, it doesn't write any of his book for him because he's the human. He's a good writer, it's his ideas. But when you've got a really long book like he has, identifying places that themes were brought up but not properly rounded out, or examples were mentioned but not properly sourced, or segues were missing or whatever, all these issues, it's like this amazing editor.

    Jeremy Howard36:58

    I mean, he has a human editor as well. And the human editor also works in Solve It.

    Matt Turck37:02

    I hadn't realized. I thought that was just for code, but that's for text as well.

    Jeremy Howard37:28

    No, it can't be just for code. 'Cause if we're going to make 5,000 products like GE did that use AI, they're going to be as broad as GE's products were that used electricity. So yeah, I'm using it to help me manage my business. I'm using it to help me do system administration of our servers. Yeah, I use it for everything. I kind of live in it. I'm either writing stuff to make it better, or I'm using it to use AI.

    Jeremy Howard37:58

    So yeah, definitely, I can tell we're more advanced users of AI than anybody else because people can keep complaining about all the problems with AI. And I'm always like, I don't have any of those problems because we just use this really different approach where the human is considered a vitally important part of the process.

    Matt Turck38:11

    To unpack that, so some of the problems being, for example, hallucination. But if you're in constant dialogue with the AI, you can recalibrate what would be a hallucination. So that—

    Jeremy Howard38:30

    Yeah. And we've got grounding and retrieval built in, and it can operate over your documents. So for example, Eric's got, as you can imagine, hundreds and hundreds of documents and interviews and stuff developed for the book, and they're all there readily available through a kind of agentic search-type mechanism. And then they can be brought into the dialogue, and the AI knows about the chapter he's writing and knows about the to-do list he's working on for the chapter and knows about what things he's resolved so far.

    Jeremy Howard38:58

    And we've also got it connected up to a kind of LLM-powered text editor, which is collaborative. So it's kind of like a Google Docs, I guess, but for the AI. Yeah, all these things, they all kind of come together, and increasingly I'm writing the next version of Solve It in Solve It.

    Matt Turck38:59

    Mm-hmm.

    Jeremy Howard39:20

    So these are the kind of positive feedback loops that we were really hoping for when we started the company, this kind of R&D cycle I described. It's interesting. It does mean not that much yet is ending up in the outside world. We have our own nice loop going on and we're kind of happily doing that. But we're starting to look more now like, okay, let's try to start bringing more of this out into the world as well.

    Jeremy Howard39:39

    It's a challenge to find a way to do that that doesn't slow down our path too much, but also does get these societal benefits that we were hoping for.

    Matt Turck39:55

    The plan is to eventually commercialize or at least make Solve It broadly available. You mentioned 1,000 beta testers, or is that going to be the secret sauce for Answer AI that's going to enable you to create those 5,000 or X many thousands?

    Jeremy Howard40:14

    Yeah, I think it'll be like AWS. I think we will continue to make it more and more available to people. The way we did those first 1,000 was we barely even mentioned Solve It, the product. We just talked about creating a—which we did do—we created a course called How to Solve It with Code based on Pólya's classic math book, How to Solve It. And we said, we'll open up registrations, but if it gets to 1,000 people, we'll close it off.

    Jeremy Howard40:44

    And that was hit within 24 hours. So there was a lot of interest in understanding how to solve problems with code. And we took people through the process in the platform because there's no other platform that's designed for this kind of iterative approach like Solve It is. Now we've had hundreds of people of those come back to us and say basically, this has actually changed my life. I can now do things I could never do before. I got this job, I started this startup, I solved this long-running academic problem.

    Jeremy Howard41:13

    So it's been a bit mind-blowing, but it does feel like it's not something we can just say, here's the product, go use it. It kind of needs the training as well, the process. So maybe we're thinking at the moment, the next one will be like, maybe we'll cap it at 10,000 rather than 1,000 and go through a similar process. But we want to make it, each time we do it, the amount of training required is halved. We want to get better and better at that.

    Jeremy Howard41:21

    By making the product better and better. So we're going to try to turn the 10-week course into a five-week course.

    Matt Turck41:28

    And when would the next course come out, and how do people apply? Like anybody listening to this, how does one—

    Jeremy Howard41:30

    I think like two or three months.

    Matt Turck41:36

    Two or three months. Okay. And then how does one get on the waiting list, for anybody that's—

    Jeremy Howard41:40

    Answer AI website. We'll announce it in all of those places.

    Matt Turck41:58

    All right. Solve It and dialogue engineering. And I mentioned that you guys have had this incredible product velocity, which I guess is a result of this.

    Jeremy Howard42:18

    And actually, I mentioned that there's another one that people can use right now. It's free, which is called Shell Sage. But it's a really good example of the general approach. So it's written by one of our team, Nick Cooper. He was previously one of the LLM leads at Stability AI. And basically, it uses tmux as the environment which you and the AI are in. So tmux, as I'm sure you know, is something that runs kind of over the top of your terminal and gives you a persistent history of everything that you typed and everything that was sent back to you.

    Jeremy Howard42:51

    And with Shell Sage, you can at any point kind of invite an AI into that environment and ask it questions. And it has access to all of that history, plus all of the aliases and your .bashrc and information about your operating system. And if you start using it, you'll really get this feeling of what it's like to work with an AI that knows what you're doing. It's tiny. It's got a prompt and maybe 100 lines of code.

    Jeremy Howard43:02

    I don't know. It's really small, very simple idea, but it kind of gives you a sense of the power of this basic thesis.

    Matt Turck43:04

    And that's available?

    Jeremy Howard43:25

    And we use it all the time now. Yeah, you can just pip install it. We use it all the time now to, like, if I'm working on a server and there's some weird things in the log, I can just say, "What's this weird thing in the log?" And it knows exactly what I mean. And it'll be like, "Oh, that's a TCP disconnect caused by packet fragmentation. You could run this command." And I'll run the command and I'll do something else, and it'll be like, "Oh, it didn't seem to work."

    How Answer.ai decides which projects to build

    43:54
    Jeremy Howard43:54

    And it knows exactly what happened because it just saw it. And it'll be like, "Oh, I see. Actually, you're running on this version of Linux, that version, you have this issue." It's just so nice to have this dialogue. You can't really dialogue-engineer it because the dialogue is this tmux scrollback buffer, but at least the dialogue has you and the AI in it sharing each other's history. So you get 50% of the way there.

    Matt Turck44:24

    So let's talk about some of the stuff that has come out of Answer AI over the last few months. So again, it's pretty remarkable how much you guys have built with a team of, I guess now, almost 14 or 14. So, just reading through my notes, it goes anywhere from next-generation encoder models like ModernBERT to retrieval and search tools like ColBERT, long-context and memory optimization called Compress, developer tools like FastHTML and MonsterUI. So do you want to maybe pick one or two and talk about what you guys—

    Jeremy Howard44:30

    I think FastHTML, because I think that's an interesting example.

    Matt Turck44:31

    Great.

    Jeremy Howard44:35

    It's ColBERT, by the way, not Colbert, just so you know.

    Matt Turck44:42

    Oh, interesting. Especially, yeah, especially of all people, as a Frenchman, I should pick the right pronunciation. Okay.

    Jeremy Howard45:05

    Most of our retrieval stuff is done by a guy called Ben Clavié. So you can probably guess where he comes from. So FastHTML is a brand-new kind of web application development system. So it's not just another library or framework, but it's a different kind of library or framework. And so that seems really weird. Like, why would an AI R&D lab do that? But when you think carefully about our strategy, it makes perfect sense, which is, if we're going to have 12 to 14 people create 5,000 to 10,000 extremely commercially successful products, we're going to have to be extremely efficient at every level.

    Jeremy Howard45:45

    Nobody currently thinks that the web development stack we have is efficient. It's definitely not. So, yeah, we have to make everything work great. And I'm very much somebody who likes to take advantage of the foundations, the basics of a field, and really take advantage of those as far as we can. I'm a big fan of HTML and HTTP. And then there's a very small JavaScript library that sits on top of those two protocols called HTMX, which basically you can think of as a polyfill for the stuff that browsers should have always been able to do, but can't.

    Jeremy Howard46:26

    So, like, there's a number of HTTP verbs, not just GET and POST, but in a normal browser you can only GET and POST. The result of calling an HTTP verb normally changes the whole page. With HTMX, you can just have it change a part of a page. With normal HTML, the only ways you can call an HTTP verb is by clicking on a link or a button. With HTMX, any interaction causes an HTTP verb to be called. So we really grabbed onto this because we're like, this is taking the basic fundamentals of HTML and HTTP and letting them breathe, getting them to take full advantage of their potential.

    Jeremy Howard47:05

    And so then we built a system on top of that, which also takes advantage of some core capabilities built into the Python programming language. It basically turns out, you might know, like Lisp is built on this idea of S-expressions. You might know that there was originally an alternative representation of these called M-expressions, which, for whatever reason, didn't take off. But actually Python is an implementation of M-expressions, and the language is so flexible that you can basically have those M-expressions do anything you like.

    Jeremy Howard47:45

    And it turns out that M-expressions can fully map to the entire HTML syntax. And therefore you can use Python as HTML syntax. But when you do, you suddenly get this benefit that you don't need any templates or special template languages or whatever. You can create functions that return Python, which are actually M-expression representations of HTML. They can also contain HTMX that can fully integrate with HTTP. You get this sudden blossoming. It's extraordinary. So we created this thing called FastHTML, which brings all these ideas together and lets you create arbitrarily rich, sophisticated web applications in a single Python file.

    Jeremy Howard48:23

    You can use more Python files, right? But how would you do that? Well, it's just—this is the nice thing—it's just Python. So you can create exactly the same kinds of modules and packages and stuff that you can in anything else in Python. So this is one of the things that's dramatically increased our production velocity, is that we've written our own way of writing web apps. Absolutely wouldn't exist without HTMX. That's really the genius behind it, to be honest, as is Python's data model.

    Jeremy Howard48:57

    So we're kind of gluing stuff together that other smarter people already built. And then we're like, "Oh, this actually lets us create a whole new kind of web component, which is kind of server-rendered, but as rich as you want it to be, and entirely written in Python." That's what MonsterUI is. MonsterUI is the first framework of that kind, and it's heavily inspired by shadcn/ui, but we're trying to bring those ideas into the FastHTML world. There's actually a Python library that already started creating a Python version of shadcn/ui called FrankenUI.

    Jeremy Howard49:22

    So we're actually taking advantage of that. So again, we're bringing in ideas from other people. But yeah, it's just been great. And so then we've also written our own deployment platform, which is kind of very soft-launched at this stage. It's got a dozen users or something.

    Matt Turck49:24

    What is that called?

    Jeremy Howard49:45

    sh. So it's nice. All the bits we're using, they're bits we fully understand because we made them, and they're designed around our philosophy and our thesis. And each time we do more of this stuff, we just get faster and faster. It feels like we're giving ourselves superpowers every day now. It's really nice.

    Future of Answer.ai: staying small while scaling impact

    49:47
    Matt Turck50:17

    Is that a guiding principle between Solve It and what you just described, that, of all the universe of projects that you could be picking, one of the guiding criteria is to build stuff that enables you to go faster? Is that the key principle? Is that only one of the principles? I guess the broader question is that in a universe where, if you're going to build 5,000 applications or more, you could be doing anything, how do you choose?

    Jeremy Howard50:43

    It's not far away from it. It wouldn't be quite how I'd describe it, but it's not bad. So we have this idea we call the substrate, which is to say, okay, in a normal company like GE, you want to expand, you want to create 5,000 products. You do that by hiring more people, creating departments, setting up offices in different geographies. So that's a very recognizable thing, that you build up a company. Then you have your multinational, which is a bit of a different kind of company to perhaps your kind of industrial conglomerate.

    Jeremy Howard51:13

    But so they kind of innovated a bit on that idea of, like, what is it that's doing this work and producing these products? So we're not going to have that if we've got 12 to 14 people. We're not going to have offices and departments, and we don't have any roles. We don't have any professional managers. So the substrate for us will be kind of the AI and the automation and the connectivity to the people iterating with that.

    Jeremy Howard51:48

    So things that contribute to the substrate are the things that we particularly prioritize, but not in an abstract way. So, for example, it's all about solving problems as they kind of come up. It's very much a lean startup approach. So, for example, not having professional managers is a bit unusual. And we naturally face this problem of people sometimes wondering, like, hmm, what should I do today? What am I meant to be working on? What are other people working on that I can contribute to?

    Jeremy Howard52:19

    So we're building a system to answer these questions. So it so happens it lives inside Discord. We're using Discord as another kind of shared space-like thing. It's an environment in which all our communications are there: voice, text, attachments, pictures, people. And we share it with an AI model. We can ask it questions, and it can interact with us in that environment. It also has access to all of GitHub. And there's a great example of something that helps us work each day better.

    Jeremy Howard52:45

    And so when somebody tries to ask that AI a question or get help doing something and it isn't able to, then it's a very immediate thing of, like, oh, I need this to work. It doesn't work. Okay, we need to fix it. And again, it's like, is that something that one day might be a product we sell? Maybe. We really like it. It's particularly helpful. Eric Ries is helpful for lots of things, but one thing he's particularly helpful for is he's seen every AI startup under the sun.

    Jeremy Howard53:18

    They all come to him for advice. And so the feedback we keep getting from him is like, I've never seen anything like this before, about pretty much everything we do. So that's also nice, to just kind of see that we are actually genuinely innovating here. But we're always keeping an eye out. We want to take advantage of stuff that other people create as well.

    Matt Turck53:40

    Amazing. Well, you've been very generous with your time. I really appreciate it. Maybe to zoom out, the next, I don't know, five years of Answer AI—you mentioned it was a 20-year project. In five years, are you still 14? Are you a bit bigger? How do you—what does success look like?

    Jeremy Howard54:05

    Yeah, hopefully we'll still be 12 to 14. Yeah, the basic idea is, like, why should we need lots of people? That would feel like we're failing at our thesis if we need to scale by adding headcount. We can always be wrong, and if so, then we'll adjust. But our strategy has been working out so far, which is we don't seem to need to scale through headcount. So yeah, in five years' time, hopefully we might have 50 products that people are using, similar-sized company.

    Jeremy Howard54:33

    There may be some mix of commercializing some of the bits of the substrate, but hopefully more and more providing things that use AI rather than things to help people make stuff with AI. We want to create more like the record players and the dishwashers rather than the transformers.

    Matt Turck54:39

    Fantastic. Well, it's been an absolutely fascinating conversation, Jeremy. Thank you so much for doing this.

    Jeremy Howard54:40

    Thanks, Matt.

    Matt Turck54:41

    Really appreciate it.

    Jeremy Howard54:43

    Great to see you.

    Matt Turck55:03

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.