MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Diffblue: AI For Code Testing with CEO Mathew Lodge

    Mathew Lodge is the CEO at Diffblue. We cover why LLM coding assistants trade generality for errors, how Diffblue analyzes entire codebases to iteratively generate boundary-testing unit tests, and why it generated 3,100 tests for a bank in eight hours rather than a developer's estimated year of work.

    10/25/2023

    Hosted by Matt Turck · with Mathew Lodge, CEO, Diffblue

    AI code testingreinforcement learningunit testingJavadeveloper productivity
    Listen now
    YouTubeApple PodcastsSpotify
    46 min · 1 chapters
    Contents

    Transcript

    Full episode

    0:00
    Matt Turck1:23

    Hi, Mathew. Welcome. Thanks for doing this. So today we are going to talk about generative AI for code. You are the CEO of Diffblue, a very interesting startup based in the UK, originally a spinoff from Oxford University. And amongst other accolades, you were recently named to the CB Insights AI 100 list, the most promising artificial intelligence startups of 2023. And of course, you're listed on the annual MAD Landscape that we produce here every year at FirstMark. So I'd love to take it from the top and maybe have you paint a broad picture, if you will, of the generative AI for code space, which is something that people have been very excited about.

    Matt Turck1:57

    And there seems to be a lot of players doing different things. So it'd be really interesting to sort of understand who does what and the different approaches, and then go into what Diffblue does specifically.

    Mathew Lodge2:24

    Yeah, sounds great. So thanks for having me on the podcast. So if we think about generative AI, everyone's heard of ChatGPT and Copilot, for the most part, which are large language model-based approaches. And essentially today, you see two different technological approaches. One is based on language models, and the other one is based on reinforcement learning. Tabnine was probably one of the earliest startups in this space. So they built language models, they used some of the earlier open-source OpenAI models and GPT, back when OpenAI used to open-source things.

    Mathew Lodge3:02

    And they've taken those and essentially evolved them over time, and they have their own proprietary model at this point. So, similar in architecture in a sense to Copilot, those tools essentially have a prompt that's based on what's in the IDE, the code that's behind the insertion point, so what's come before that point and what comes after. And what you're asking the model to do is predict what comes in between. And so it uses information from maybe you've got a comment about what you want this code to do, or it's some half-written function or something like that.

    Mathew Lodge3:42

    And it's going to try and figure out what it needs to do to fill in. And so these models fundamentally are trained on lots and lots of code examples. And so they learn statistical text patterns for code. And they've been trained in lots of different languages, which is how they're able to generalize across languages. And so these tools essentially make—they're like autocomplete on steroids. If you think about autocomplete, it's obviously been around for decades. But what they've done is essentially make that much more powerful.

    Mathew Lodge4:07

    And that's where a lot of the appeal comes from: an approach that developers are already familiar with. They already know about autocomplete. It's just better autocomplete. It can do more. And so it's very easy for developers to pick these tools up and start using them. And the code is by and large not correct. If you look at some of Microsoft's own statistics for Copilot about acceptance rates by developers, they've done some work looking at exactly how many suggestions from Copilot are taken by developers, how many are rejected.

    Mathew Lodge4:53

    Then you can see the acceptance rate varies depending on the seniority of the programmer. So more senior programmers are more discerning; they reject more completions. Junior developers tend to take more. But in general, it hovers around, depending on the language and the context, somewhere around the 50% mark. And a lot of it depends on the kind of code that you're writing. So essentially, what is happening here is that you get fantastic generality from this. And the generality comes with errors.

    Mathew Lodge5:20

    I mean, essentially, that's the trade-off you're making. The large language model is not going to be completely correct, but it might get you started. And a developer can sort of take it to completion, make it work, turn it into something useful. The approach with reinforcement learning is quite different in that there's no pre-training in reinforcement learning. So the idea with reinforcement learning is that you want to conduct a search of some kind to go find what you're looking for, or you're trying to find a solution as opposed to make a prediction.

    Mathew Lodge5:45

    So it's been used perhaps most famously for game playing. So if you think about things like Google AlphaGo, it's based on reinforcement learning, where you're searching for the best Go move.

    Matt Turck5:46

    Yep.

    Mathew Lodge6:01

    Now, in the case of Go, the reason you have to do this at all in the first place is because the possible number of moves in Go is so large. So that search strategy is only useful when your search space is too big for you to look at all of the alternatives. That's also why you don't see AI for playing chess, because chess's search space is small enough, you can just look at all the possible moves, just try them all and pick the best one.

    Mathew Lodge6:34

    And you can't do that for Go. So essentially, what reinforcement learning does is that you have some kind of prediction for what a good solution is. So the idea of reinforcement learning fundamentally is you can't look at all the solutions. So what you do is you spend a lot more time looking in places where solutions are most likely to be found. So you need some way of figuring out what that most likely space is. In the case of Go, AlphaGo looks at a move tree and makes some predictions about who's going to win the game.

    Mathew Lodge7:04

    And then it just tries different moves. And this is the key of reinforcement learning. The learning happens when you try the different alternatives. And in the case of Go, you play the move, you look at your prediction about who's going to win, and then you might follow a path of moves and keep following that and see: is your prediction for winning getting better or is it getting worse? And if it's getting worse, you might want to backtrack, try some other moves, and sort of further up in the move tree, if you think about it that way.

    Mathew Lodge7:36

    And through that process, at the end of that process, you've got the best move that you can find. It's not guaranteed to be the best move, but as we saw with AlphaGo, it's good enough that it can beat Go masters in some competitions. Essentially, you can use the same for coding. This is what we do at Diffblue, where we're looking for tests. So we're writing unit tests. So a unit test isolates a unit.

    Matt Turck7:41

    We'll talk about in a second what that means, but let's hold that thought for now.

    Mathew Lodge7:58

    Okay. So essentially, we're searching for the best test. And so in our case, we analyze the program, and that gives us an idea about where the best solutions might be found, so where the tests might be, what they might look like. We write that first test, we run it against the code, we see how it performs. What kind of coverage it gets, does it test boundary conditions, how much branch coverage do we have, things like that.

    Mathew Lodge8:31

    That enables us to improve our prediction about what a good test would be. We make that prediction, we build the test, we run it, we try it, and we iterate through that loop. And at the end of that process, we have found the best test that we can for a particular block of code. Google DeepMind has also applied this approach to things like code optimization, where essentially they're searching for more efficient implementations of a particular algorithm. So they shaved about 7% off the sort algorithm in the C++ library, which doesn't sound like a lot.

    Mathew Lodge8:41

    But considering how often that code is called, that's actually quite a big win.

    Matt Turck9:17

    Yep. Okay. And so, are you 100% focused on reinforcement learning? Or is that a combination of reinforcement learning and LLMs? Or is that—people, I'm sure, have heard of GPT having a sort of reinforcement learning from human feedback, or RLHF, being a part of what they do overall. So is that a hybrid approach, or is that solely reinforcement learning?

    Mathew Lodge9:43

    Yeah, so we use an ensemble approach. So we've got different aspects of the models. And we've obviously been playing with large language models to see if we can improve on that process. So some of this is like, how do you make predictions? And you could use large language models to do some of that work. You can have them make some predictions for you about what a good test might look like, and so on. So we do use a number of different techniques.

    Mathew Lodge9:59

    It's not purely all about reinforcement learning. But the big difference is that in large language models, essentially, the reinforcement learning is used to tune the model and have it prefer some responses over others. And so part of what you're doing is correcting for the randomness that's in a large language model and teaching it to prefer some outcomes over others, which is fundamentally what we're doing with reinforcement learning. Reinforcement learning is first and foremost how you guide the algorithm, how you assess the effectiveness of what we're doing, how we make predictions. That is ensemble, but the emphasis is different, right?

    Mathew Lodge10:35

    It's primarily about reinforcement learning as opposed to using reinforcement learning to correct what a large language model has produced.

    Matt Turck11:07

    Okay, great. So I interrupted you a minute ago that you need tests before we go further into how things work. Let's backtrack a bit and talk about what that is, to make it interesting to a broad audience that may not be familiar with the intricacies of this. So, what is a unit, I guess, and what are unit tests, and why did you all pick that area compared to, like, other areas of code generation?

    Mathew Lodge11:30

    Yes. So, the idea with unit tests is you isolate one module of the software, and you test everything inside of that in isolation. So, unit tests have been around for about 25 years as a concept. And you can think of it, if we make a very simple analogy about a production line for a car, right? When you're making a car, you check that the brakes work when you put them on the car. You don't wait until the test drive at the end of the production line to find out if the brakes work.

    Mathew Lodge11:49

    And so, the same idea in software is called shift left. Now, that's a crude analogy, but it works. It gives you the right idea. You're essentially trying to make it so that you find defects as soon as possible. And so the idea with unit tests is that a developer can run them, and they run quickly, and they run them at the time they're writing the code, so that failures in unit tests can very quickly be fixed by the developer, because they're in the headspace of that code.

    Mathew Lodge12:24

    They know exactly what they're doing. They've probably been working on this for a while. Everything is there for them to be able to make a change to the code and fix an issue that they find. And that's one fewer defect that then shows up in QA, because fixing bugs that are found in QA takes much, much longer.

    Matt Turck12:53

    Typically, just to make this super basic for a broad group of people, developers would not test, and they would just write code, and then this would go into a separate group of people who'd actually test. And the idea is that the developers should be testing in a way that they were not originally. Is that part of the idea, or is that too much of a caricature?

    Mathew Lodge13:20

    Yeah, it's a simplification, but yeah, it's pretty much what happens. I mean, so when a developer is working on something, they will test it themselves. The question is how they might do that. They might just test if that piece works. They might build the entire application. And maybe if they're making a change to how the interface works, they go in and make sure that when they do the things in the interface, the right things happen. And so part of the challenge with that is that, one, you need to build the entire application.

    Mathew Lodge13:43

    And to run it, you might need dependencies like a database or something like that. So that can end up being quite slow. So what developers try to do is they don't really want to have to deal with that, so they try and simplify things for themselves and get enough confidence. So it's not like developers are reckless, but those changes that they're putting in, plus the changes that other people are putting in, and then sometimes the software doesn't build, and it's the works-on-my-machine problem, right?

    Mathew Lodge13:59

    Okay.

    Matt Turck14:07

    And so, historically, before Diffblue, over the last 25 years, how did people do it?

    Mathew Lodge14:24

    Historically, what they did is they spent a lot of time in quality assurance. So, there was a separate QA team whose job it was to look at the original, what's the software supposed to do, the original specification, and take that idea and check that the software does what it's supposed to.

    Matt Turck14:29

    What kind of tools do they use to do that?

    Mathew Lodge14:54

    It depends on what kind of software you're building. But let's say if you're building a web application or a mobile application, a lot of it is clicking on buttons and making sure the right thing happens. So working through the interface is typically how they would do that. And in the sort of bad old days, that was all done manually. So people literally sitting in front of screens, clicking on buttons. And obviously, everyone wants to automate that. And there's a vast number of products out there to automate UI testing these days.

    Mathew Lodge15:21

    And web servers, web applications, and browsers have made that easier to automate. And so you can sort of automate the manual process out of this. The fundamental issue, though, is still that all you know when the software doesn't work is that it doesn't work. You don't know why. So the next thing that has to happen is the triage of, like, okay, it does the wrong thing. Why is it doing the wrong thing? And sort of working backwards from that problem to what's the root cause.

    Mathew Lodge15:53

    That by itself might take days or weeks of work, depending on how complex things are. And especially in complex systems where you've got databases, you've got caches, you've got all these other things happening all at the same time, you get into these problems that are hard to reproduce because they only happen when you've got a certain combination of things all happening at the same time. And that makes fixing those bugs very time-consuming. So just root-causing it can be very difficult. And then the developer has probably moved on to working on something else.

    Mathew Lodge16:18

    They have to stop what they're doing, go back to that code that they were working on maybe a month ago, try and figure out what's going on, try and figure out what the problem is, the fix for it, and then they have to send it back to QA to make sure, did we really fix that? So that process is just much longer, much more time-consuming for everyone. But that's what people did.

    Matt Turck16:26

    Yeah. All right. And so now with Diffblue and reinforcement learning, then how does that work?

    Mathew Lodge16:50

    So, the challenge for a lot of organizations, and this is sort of how Diffblue got founded, to your question earlier on about that, is that unit testing is tedious, error-prone stuff to do. It can be quite difficult code to write. And the idea with—because you have to test a single unit—you have to figure out how to isolate just that unit. And so, for a database-driven application, where you're expecting to have a database on the backend, you don't want to have to do that anymore.

    Mathew Lodge17:25

    So, you're going to have to figure out how to essentially emulate the database. And so, you're looking at trying to synthesize test data or synthesize dummy data, pretend to get the same thing back from the database that you would get, and that's called mocking. So, that process itself is very tedious. You have to get exactly the right configuration, you have to use a mocking subsystem, you have to learn how to use that. That just by itself can get very complicated and very slow.

    Mathew Lodge17:57

    And this is why developers don't like writing these tests. And fundamentally, they're not paid to write tests. They're paid to deliver the functionality, the new things in the application, the fixes to the application, the updates. And tests are just there as part of the process to help them. But what it means, especially in larger organizations, is that you've got very large applications that run big chunks of the organization. They might have been around for 15, 20 years.

    Mathew Lodge18:22

    People that first wrote them are long gone. There's no one person who understands the entire application and how it works. And so that process is now very, very difficult to do, to go back and sort of get unit tests when you didn't have them before, or it's not been the norm to write unit tests. It's very, very difficult. It's like a mountain to climb.

    Matt Turck18:32

    So, there's a problem with new tests of new code, but of course, there's the whole legacy thing. Okay. All right. Fascinating.

    Mathew Lodge18:47

    Yeah. So, legacy is a big challenge. And it depends on what you mean by legacy as well. In a lot of cases, the entire reason that unit tests are useful for an application is because it's very important to the organization.

    Matt Turck18:57

    Okay, that's a really helpful thing. So, how does Diffblue solve that problem then? What are the actual sort of high-level mechanics of how it works?

    Mathew Lodge19:19

    Yes. So, our product writes code, this test code, fully autonomously. So, very different from LLMs. And in order to do that, we use reinforcement learning to help us find those tests. So the basic steps are: the first thing we do is analyze the program. So we see the entire source code of the application. So again, different to LLMs where you've got limited context, you've got before the insertion point, after the insertion point; we've got the entire program.

    Mathew Lodge19:56

    And so we do analysis on that, and that enables us to come up with our first guess, essentially, of what a good test would be. So, we take it method by method. So, all the functions of the program, enumerate all of those, understand exactly the context of each method and how you're supposed to call it and what you're supposed to feed into it, all of those things.

    Matt Turck20:00

    So, it's like a mapping exercise of some sort?

    Mathew Lodge20:21

    Yeah. So, you essentially compute what's called a call graph, like all of the functions of the program, how they call each other, what the signatures are, so what data do they need, because you're going to have to create that data to pass it in to test them. And we also look at the structure of the program. We do a little bit of what's called data flow analysis. So we look at how data moves through the program to a certain extent.

    Mathew Lodge20:49

    And that gives us a way for us to produce that first test and essentially gives us our first set of predictions as to what useful tests might be. And so then we write the first test. So we essentially compile it, run it against the method under test. So we call the code that we're trying to test with our guess at what a good test might look like, with the data that we think might work. And we see how it does. So we evaluate the effectiveness of that.

    Mathew Lodge20:55

    So we look at how much coverage did we get? Did we hit branches?

    Matt Turck20:55

    Right.

    Mathew Lodge21:20

    Are these interesting values that would be on the boundaries of changing the behavior of the code? So if there's something in there that says, if a value is greater than 100, then do something different. And so you want to maybe feed in 1 and 100 and 101 and 99 and try and sort of exercise that boundary. We want to know, does it throw an exception? And then what does it do? So, what in the state of the program changes?

    Mathew Lodge21:48

    Because we're going to need to know that. That's our information that we need to write assertions for exactly what that method does. And then we just keep going, and we iterate through that process, and we repeat that again for every single method in the program. So, we have customers that have million-line Java applications. They might get 50,000 unit tests for a million-line Java application. Wow.

    Matt Turck22:06

    Yeah. That's incredible. It just sounds like incredibly complex systems, like interdependency and all those things. When you run those jobs, is that super compute-intensive and it takes a while? Or what's the right way to think about it?

    Mathew Lodge22:31

    Well, ironically, it's far less compute-intensive than large language models because our stuff will just run on a regular CPU. And so, for one of the banks, they had an application where we wrote like 3,100 tests in about eight hours. Now, they estimated that that's like a year's work for a single developer, 3,000 tests.

    Matt Turck22:31

    Wow.

    Mathew Lodge22:39

    And we can do that in eight hours. Some of it depends on what the code is, but that sort of gives you the sense of how much time you're saving. Okay.

    Matt Turck23:01

    And currently you support Java, if I read correctly, because there's that dimension as well. Like, the tests are very language-related. Presumably you chose Java because that's the most widely used for legacy kind of application code. Is that right?

    Mathew Lodge23:10

    Yeah, that's right. And that is why we chose it. So it's been a top enterprise programming language for the last 25 years, something like that.

    Matt Turck23:11

    Yeah.

    Mathew Lodge23:12

    So it's vast amounts.

    Matt Turck23:18

    Do you have others on the roadmap, or it doesn't even matter because the Java universe is so large?

    Mathew Lodge23:42

    No, we will do others in the future. So we just added our second language, which is called Kotlin. So Kotlin also uses the Java Virtual Machine. You get a lot of applications that are hybrid Kotlin and Java today. And so it was a good second language for us to do. Some of our customers also have applications that they wanted us to support. And we'll absolutely do more languages in the future. It's the usual thing with a small company: you've got to do one thing incredibly well when you start out.

    Matt Turck24:08

    Yeah. And since we're talking sort of about roadmap, it feels like, again, a very complex world and a vast universe of customers you can sign and all the things. Is part of the idea over time to expand into other things like, I don't know, code migration? If you spend a lot of time in the gnarly world of legacy software for large institutions, is part of the idea, like, one, you test, but second, it's like, okay, well, this is how you should be writing it or porting it to a different language or whatever it is.

    Mathew Lodge24:44

    Yeah, there's a lot you can do. If you look at what some of the very, very large companies do, like Meta, if you're a software developer at Meta and there's a bug in your code, when you open your IDE, you don't just have your code there; you also have a suggestion for a fix. And so you can apply the technology, apply the approach that we use, to lots of other coding problems where you're trying to find better code, essentially.

    Mathew Lodge25:03

    So there's that kind of thing you can do in the future. And there's all kinds of improvements to existing code, as well as writing brand-new code.

    Matt Turck25:24

    Yeah. Does that mean that the COBOL situation is going to be fixed in the short term? COBOL being that very old language that still runs in production globally at a somewhat massive scale, especially in financial institutions, as far as I know.

    Mathew Lodge25:43

    Yes. So, it's funny you mention that. We have a number of projects underway where they're doing COBOL-to-Java migrations. And so part of that is the testing part is very, very important for them to be sure that the Java code is doing what the COBOL code did. And also, it gives them the ability to then evolve that and refactor that Java code and maintain it over the long term, rather than just sort of taking the application, doing a one-time conversion. They want to update the program, right?

    Mathew Lodge26:18

    So that's a big part of it. Interestingly, the motivation for that is the cost of things like Micro Focus and the software that runs on the mainframe. That's the big driver because PE companies bought all of those software products and raised the price. Yep.

    Matt Turck26:43

    So while we're talking about code in financial institutions, I read somewhere a nice quote from Citibank, and I saw that you guys were exhibiting at various financial services trade show conferences. So is that the core target for you all, or how do you think about go-to-market?

    Mathew Lodge27:09

    Yeah, so we got started in financial services because Goldman Sachs was an early backer. And so a lot of it was word of mouth and working—those guys all talk to each other, people move around. That's how we got started. And so we grew out of that. But more recently, we've got customers like Cisco and AstraZeneca. And everyone has the same problem. It's not limited to financial services.

    Matt Turck27:26

    But you're focused on the enterprise, so large codebases. If I'm a smaller startup, I may have less of that problem, or I'm just too small to be an interesting customer?

    Mathew Lodge27:29

    No, no. You may still have the problem.

    Matt Turck27:30

    Okay.

    Mathew Lodge27:58

    Yeah, yeah. Yes, yes. You may still have the problem. And also, what you find is we sell—our go-to-market strategy is very much sort of land and expand. So we might start with a single team. That's exactly how we started at Citi, with a single team, and expanded from there. They've got a problem, but if things go well, other teams learn about it, they find more uses for it, then we can expand.

    Matt Turck28:08

    Okay. And you sell to the CIO, or you sell to the developer, or you sell to the VP of Engineering? Top-down, bottom-up, the middle?

    Mathew Lodge28:30

    Yeah, it's a little bit of both. So it's usually somebody like a VP of Engineering or Senior Director of Engineering who's the buyer, the economic buyer. Now, we have a free Community Edition. You can just try our product. Very important because most people have never seen anything like what we do before, so they don't know what to expect, so they want to just try it. And then if they like it, then they come back to us at some point later and say, "We'd like to buy a Teams Edition, get started."

    Matt Turck29:05

    And how do people perceive the fact that it's fully autonomous, versus, again, GitHub Copilot, which, as the name indicates, has more of that copilot kind of interaction mode where the human may feel like they're more in control? Like, this is autonomous AI. Are people offended by it, spooked out by it? How do they react?

    Mathew Lodge29:25

    Well, we offer both models as well. So, we offer an interactive model where you can just point to your code and say, "Give me tests," and we will just give you tests then and there. And so that's much more like a developer is used to. But they might be doing that in the context of, previously, we wrote 20,000 tests. And so now they're in test maintenance mode, like they've written some new code, they need tests for it, or they're maintaining the code, and so the test job's obsolete, they need new ones to replace the old ones, and they can do that interactively.

    Mathew Lodge30:09

    So we've made the product so the most important thing is it works the way developers are comfortable with working. You can go the full-on, fully autonomous route where the developers don't use the tool interactively at all. They just get the tests, they're there in the baseline of the code, they can run them. When the tests are obsolete, they can just delete them because we can just regenerate new ones. For some developers, that's a bridge too far.

    Mathew Lodge30:16

    So it's really what people are comfortable with.

    Matt Turck30:30

    Yeah. And I saw, when I was prepping for this, looking at the demo that you have on the website, it's still very much transparent, right? It shows you exactly the line where the problem might be.

    Mathew Lodge30:47

    Yes. Yeah, that's right. And we designed the product to fit in with the way people do testing. We use JUnit and TestNG, the two sort of open-source frameworks for testing in Java. We generate code that uses those. It's what they're used to.

    Matt Turck31:28

    And still in the general vein of how people react to it and humans, obviously, the unavoidable question around all things AI and generative AI in particular is AI and jobs. And I was alluding to a nice quote from an MD at Citibank, Jonathan Lofthouse, one of your customers, MD and Global Head of Markets Technology at Citibank, sort of making the point that this is not taking anyone's job, right? How do you position the product and experience from that perspective?

    Mathew Lodge31:55

    Yeah, it's about developer productivity, which is very important to these organizations because software is really critical to how they compete these days. They are large software organizations. And at Citi, what they wanted to do was reduce cycle times. So, how quickly can we ship? Can we increase the frequency at which we ship? And because that will make us more competitive. And also, how do we improve the productivity of our developers so they spend more time on, frankly, the more important parts of their job, which is implementing what we want in software, as opposed to having to write all of the tests and doing the drudgery?

    Mathew Lodge32:41

    I mean, the analogy I like to make is it's kind of like the shift 30, 35 years ago from assembly language to compiled languages, right? So, what you did as a developer really changed. People used to write code in assembly language back when computers were far less powerful, they had far less memory. And so the original versions of BASIC were written by Bill Gates and Paul Allen and all those guys in assembler. And then higher-level languages came along because they offered this massive productivity improvement.

    Mathew Lodge33:13

    So what you did as a programmer changed. Now, instead of writing individual machine-level instructions, you were thinking about algorithms and data structures and operating at that level. And that's where that productivity gain came from. You got this shift in what you spent your time on to higher-value things. And it's essentially that all over again. And I think that we talked a little bit about large language models and that approach and how it's an extension of what developers do today, but it's also lower impact for that reason.

    Mathew Lodge33:54

    So, a code suggestion, you could really still only go as fast as a human can go, because a human has to review that, decide whether it's useful or not. If it is useful, make it fit into the code and finish the work. You're still basically going at the speed of a human, whereas with fully autonomous code generation, you're much, much faster. Now, it's more of a change to how people work, but the benefit is much higher. And so I think that is really the promise, or it should be the promise, of what machine learning and artificial intelligence should be able to do.

    Matt Turck34:36

    Great. Well said. Very interesting. You have a— just looking around on Twitter and some press and everything, you have some really interesting spicy takes. I guess this is the term of the day: spicy takes. And one that caught my attention recently was, "Prompt engineering is not a thing." Do you want to expand on that and explain what that is?

    Mathew Lodge35:04

    Yeah. This came out of some reflection about exactly how large language models work and do what they do so well, and some experiences using very early versions of large language models for code completion, where what we found is that very, very small changes in what you asked the model to do and the code you asked it to write could result in huge differences in what you got, the output that you got. And that's essentially what people are talking about with prompt engineering.

    Mathew Lodge35:30

    Because what Copilot is doing is basically doing a lot of behind-the-scenes work on the prompt to try and get you the best code for a particular situation. But fundamentally, these models are so big and so complicated that nobody knows what they're going to do. You don't have any predictive ability. And so the idea that there's something called prompt engineering—engineering in the sense of you're applying an approach because it's going to give you the output that you desire—is a complete fallacy.

    Mathew Lodge36:00

    At best, it's trial and error. It's prompt trial and error. And there are rules of thumb about how you construct a good prompt and so on, which is based on the way that the backend model works. But it's not engineering. It's not engineering in the traditional sense.

    Matt Turck36:08

    Okay, so I should revise my LinkedIn profile to no longer say engineer.

    Mathew Lodge36:09

    Yeah, right.

    Matt Turck36:20

    And another one I liked, which maybe you alluded to, was, I think there was a tweet: "One problem with large language models is their largeness."

    Mathew Lodge36:51

    Yes. Well, one of the reasons we've been successful in highly regulated industries is because we made our software so that they could run it locally. They could run it in their own development environment, inside their own infrastructure. They don't have to send their code anywhere. They don't have to worry about any of the privacy or IP implications of any of that. And it's incredibly difficult to run large language models locally, and it's incredibly expensive. And so we built a version of our product where we use the large language model to generate all the tests.

    Mathew Lodge37:21

    So we took our reinforcement learning engine out and put in a large language model, and we did a lot of work on prompts and so on. And our product will write a test in about a second and a half, and we were getting 40 seconds to write a test with a couple of hundred-billion-parameter large language model on the backend. Now, we could have paid more and probably sped that up, but those very practical considerations make a difference, make a big difference.

    Mathew Lodge38:01

    And so practically, whether people are going to be able to use the technology, does it solve some of their privacy concerns? The largeness of large language models works against you in those situations. It's one of the reasons I think that we're going to see more smaller models. And you've got startups now emerging, which are essentially trying to distill down from a large model into something much smaller that will run without a GPU. And they're trying to automate that process for exactly that reason.

    Matt Turck38:15

    Yeah. And I also read somewhere, like, you're a big fan of open source as well—like smaller models, but also open source. Same reasoning?

    Mathew Lodge38:41

    Yeah. Yes, same reasoning. But also what we've seen happen with some of the image generation stuff, diffusion algorithms. So the most successful one of those, based on the Stability engine, has seen very, very rapid improvement because there's a whole bunch of people tinkering there, and they don't have a lot of money. They don't have 1,000 GPUs to run it on. And so the combination of all those people working together has led to some very rapid advances.

    Mathew Lodge39:03

    And very quickly, Stability, the open-source project Stability, got to be better than DALL-E 2, and the spend on DALL-E 2 is immense. So I'm a big fan of open source. Yep.

    Matt Turck39:24

    Yep. And so you're saying open source is going to catch up to— that's sort of the question of the day abundantly on Twitter every other minute, or I should say X, every other minute. But you think the ultimate win or victory of open source is—the writing is on the wall?

    Mathew Lodge39:48

    I think so. I think because it allows that level of innovation I just talked about. And I think also it gives people the ability to specialize these models to a particular task. The genius of ChatGPT is that it is so general, right? It's a text generation engine that will do anything. It'll write you a poem. It'll write you a story. It can pretend to be Shakespeare. And that really just catches everyone's imagination, and it suddenly becomes a household thing.

    Mathew Lodge40:07

    And it's brilliant. But in a commercial context, you don't need it to be all those things all at once. You need it to do one thing incredibly well, which is your problem, whatever your problem is.

    Matt Turck40:22

    Okay. All right, so maybe to close, I'd love to go in a pretty different direction. You are based in the UK.

    Mathew Lodge40:22

    Yeah.

    Matt Turck40:51

    You're a spinoff from Oxford. And for anybody who watches this online on YouTube as opposed to listening to the podcast, your Zoom background is a famous bridge at Oxford University. I'm personally a huge fan of this sort of Europe-to-US or Europe-to-the-rest-of-the-world kind of startups. And it seems that the UK in particular has produced a bunch of really interesting companies. Not to be the VC plugging their companies, but we're investors in Synthesia.

    Matt Turck41:04

    You mentioned Stable Diffusion. Stability AI is originally from the UK.

    Mathew Lodge41:04

    Yeah.

    Matt Turck41:22

    You guys obviously are from the UK. What's, I guess, what's the reason for that? Why is the UK AI scene vibrant? What are the parts of it that you're excited about? What parts need more work? And what was your overall take?

    Mathew Lodge41:43

    Yeah, there's just a huge amount of talent over here. I mean, obviously, I'm sitting here in Oxford, but it's not just Oxford and Cambridge; it's all over the country. There's a lot of research. And as you said, there's a strong history of really good technology coming out of the UK.

    Matt Turck41:48

    Oh yeah, I forgot DeepMind in my little list, which—

    Mathew Lodge42:13

    Yeah, I was gonna say, I have to say, Demis Hassabis and all of those guys, right? UCL and Oxford together. So that's where DeepMind came from. I mean, what's interesting about DeepMind is one of the reasons that Demis said they went over to Google is because they were trying to develop LLMs and they needed access to capital, right? And as I said, I think that is a really good marriage between the VC world and what's coming out, and the talent that's here. There's not necessarily the money in research institutions to take this to places where it can do the most good.

    Mathew Lodge42:36

    And so that's the other important part. And we've been very fortunate to have investors who see the potential.

    Matt Turck42:49

    What do you wish there was in the UK or Europe that would help build the next OpenAI or next Google out of Europe?

    Mathew Lodge43:13

    Well, I don't think Brexit has done us any favors. I don't know if that's even a controversial thing to say anymore in the UK, but I think it hurt us. We had a Brexit exodus of talent, where people just said, "You know what? I'm just going back to France," or, "I'm going back to wherever." And I think especially the kind of resources you have to apply to this, the war for talent, all those other things, better relations with the rest of Europe and being able to work together more effectively is a good thing.

    Mathew Lodge43:31

    Yeah.

    Matt Turck43:42

    Okay, great. All right, well, it's been wonderful. What's happening next at Diffblue in the next few months that people should be on the lookout for?

    Mathew Lodge44:00

    So we have an exciting partnership in the works that I can't talk about now, but we'll be announcing probably in December. And we're deepening a relationship with one of our customers, and we'll have more to say about that in the next few months as well.

    Matt Turck44:12

    Okay, very exciting. Where do people find you online? I mentioned your Twitter. What's your handle? I guess we'll add that in the show notes, but where do people find you?

    Mathew Lodge44:26

    And learn more. Yeah, you can find me on Twitter, Mathew Lodge, Mathew with one T, but it's fairly straightforward. MathewLodge.com, that will give you all the links to all my social media accounts. I'm also on Instagram. I post a lot of photography on Instagram. That's my hobby.

    Matt Turck44:28

    Oh, nice.

    Mathew Lodge44:38

    I took that picture that's behind me, so. Oh, very good. So you can find me there. LinkedIn, I'm linkedin.com/in/mathewlodge. So that's me.

    Matt Turck44:57

    Oh, very chic. Okay. First time hearing your LinkedIn handle. All right. Okay. Well, it's been wonderful. Thank you so much. Really appreciate it. And congratulations on everything you have built so far, and look forward to seeing the announcements in the next few weeks and months.

    Mathew Lodge45:02

    All right. Thank you very much. I appreciate it. And thanks for having me on. Thank you.

    Matt Turck45:03

    Bye.

    Mathew Lodge45:32

    Thanks for joining us for The MAD Podcast. We're back here every Wednesday with new conversations with leaders in the machine learning, AI, and data space. And if you like this show, you can also find a video recording of not only this episode, but many, many more over on the Data Driven NYC YouTube channel. Thanks again, and catch you next week. Thanks for listening.