Welcome, Lukas. So to set it up, you are the CEO and co-founder of Weights & Biases, a.k.a. W&B. Weights & Biases is the developer-first MLOps platform. The company makes it easy to track your experiments, manage and version your data, and collaborate with your team so you can focus on building the best models. The company was started in 2018, and you have raised about $200 million in venture capital, if I read correctly. So you and I last sat down in September 2015, which is kind of a crazy thought.
I guess we were both 16 at the time. And the company that you were talking about at the time was called CrowdFlower, which I believe was later rebranded into Figure Eight. So I'd love to start with this, maybe the journey with that company, lessons learned there, and then what led you to start Weights & Biases as your next big project?
Yeah, totally. So CrowdFlower—you're right, rebranded as Figure Eight—was a data-labeling platform primarily for machine learning applications. And I started that, I think it was 2007, so very early on, and it was very unclear at the time that data labeling would be a large market. And so the journey was really several years, I think maybe before we actually raised any kind of significant funding. So we were really just trying to sell data labeling for quite a while.
And we kind of had an early run in sort of the search space. So we worked with Google and Yahoo and eBay and a lot of the e-commerce companies, and got a kind of a head start there, and then kind of had a lull for a while when there just wasn't a lot of stuff happening, it felt like, in data labeling. And then sort of reignited around self-driving cars, and kind of automotive became this huge application where we were sitting with good tools for that.
So it was a long, long journey. And I got a chance to sell into many different parts of organizations, right? Because different kinds of companies, AI will roll up into different places. I mean, almost anywhere. It depends on the company. Could be like the CMO, could be head of sales, could be, now it's typically in the tech org, like CTO or CIO, but really talking to every part of a business. And data labeling, I think, really wants to be a top-down sale.
But my background is actually building and deploying machine learning models. And so me and my co-founders are kind of looking at the space and thinking, there's not a lot of great tools for the ML developers themselves, and we really wanted to make stuff for that audience. So we started Weights & Biases really with more of an audience in mind, or a customer in mind, than an application. I knew that I really loved working with people working on machine learning.
I always felt a little bad when I would get pulled away from that into more executive conversations. It was really fun to start Weights & Biases and just go right back into the weeds of helping with the details of developer issues, getting ML deployed. Although now, it's funny, as we've grown, now we do talk to a lot of executives, and a lot of our customers, back when we started Weights & Biases, have been promoted several times. So now they themselves have become executives.
Great, great. Actually, to make this broadly interesting to a larger group of people, what is an ML developer? I think there's this concept that there's a lot of people working on machine learning, but there's actually different flavors of people actually building the models, people deploying the models. What's an ML developer, and who's your target audience specifically?
Yeah, I mean, we say ML practitioner in our internal mission statement and stuff because we want to keep it a little bit vague. Like, I've watched—I think I was originally hired as, like, a research scientist and then became data scientist and now might be ML engineer. And so I think these titles change a lot, and in different industries titles mean different things. So I don't mean ML developer like the title. I mean it more by the function. So we want to help the people that are literally just trying to get these models into production and do something.
And there's different flavors, right? There's typically MLOps teams that are more operational, more like, how do we make this a reliable, repeatable process? And we love them. We also love the more kind of research scientist-oriented person who's trying to get the last bit of accuracy out of their model. And now we see a whole new kind of archetype where it's like a software developer that's just really interested in machine learning and wants to get models working inside their company, maybe using a third-party API or something like that.
So we're not too precious about which version of that you are. I think we need to serve all of those flavors of ML practitioners well.
Mm-hmm. Very interesting. All right, so I'd love to spend a good amount of time talking about the product and the platform. So you guys have shipped quite a lot of features and capabilities in a sort of broad, horizontal way. And you have LLM models, you have a platform, you have monitoring, you have Prompts, which I believe is the most recent addition. So maybe walk us through what different parts do, maybe starting historically, like the first one you released, and then how the rest followed from there.
Totally. So we started with a product called Experiments, which at the time was kind of like TensorBoard in the cloud. That was kind of how we thought about it. Like, we'd build a model and then kind of want to share the performance with somebody else, and it was kind of hard to just, like, screenshot your TensorBoard instance and send it over, and then you always want to do slightly more complicated analysis than TensorBoard, and that was tough.
And at the time, PyTorch—I don't think PyTorch ever really had a great TensorBoard integration—so we wanted to integrate with more stuff and just make it friendlier. And so that's, we call it Weights & Biases Experiments. And I think we're well known for that. I don't know that we ever knew that that would be a great business, but it's actually—I've realized—it's a very important thing, right?
I mean, I think without tracking all the things that go into building models, you can't have reproducibility and you can't have explainability. You can't really say that you're doing, I think, ethical AI, no matter how you really define that, if you don't know how the model itself was built.
So tracking what goes into the model, is that data? Is that—what is it? Is that the model parts, the data part?
Totally. I mean, I think there's more things than you'd think, right? So when a model gets trained, the data is super important, right? The training data that it's learning from. Often there's upstream preprocessing steps that are happening to the data before it's fed into the model. There's things called hyperparameters, which might be, a famous one is learning rate, sort of how fast does the model react to new data.
But there's typically thousands of these hyperparameters that people are always trying to tune. There's sort of the architecture of the model, and that's changed a lot over the years. So there's broad categories of model, but even if you think about neural nets, there's wildly different ways you can set that up. But then there's actually a lot of other things that go into it, like what GPUs did you train your model on?
Like, what infrastructure were you using? What libraries were installed on your system when the model trained? And what we kind of realized was that people wanted to keep track of all this stuff, but they didn't want to explicitly write it down. And so one of our core tenets is this kind of one-line integration where you can just get as much value as possible by default and in the background.
It just tracks whatever you do and then just produces: this is the GPU you used, this is the dataset you used, and you don't have to do anything.
Exactly. Yeah, I think that's really important. And then, of course, it's important to make it extensible and flexible. But then the other thing is actually the performance metrics, which are all really complicated, right? People always think, okay, you're trying to optimize accuracy, but no one's ever really doing that, right? I mean, the self-driving car case is really interesting, where it's like, okay, maybe your accuracy at detecting roads goes down, but your accuracy at detecting bicyclists goes up.
What do you do with that? And typically, you actually have thousands of different complicated measures, and you're trying to figure out, okay, is this something I should deploy? And so we've actually added a lot of stuff to support this kind of general workflow. We spend quite a lot of time building something called Artifacts, which keeps track of both the data flowing into the models and the models that come out of your training runs.
And that's a tricky product to get right because the size of these files is so big typically, right? A lot of our customers are training on terabytes of data, and it might be slightly changing every day. They don't want to copy it over, but they do want to have a good record of exactly what the model was trained on. And we looked at all the tools out there, and none of them kind of worked well for that use case, we thought at the time.
So we've built that whole system. And then other highlights: our model registry has been really popular lately. So that's been really exciting to see, and that's kind of simple, right? It's just tracking, okay, what's going into production? When does it go into production? How is it evaluated? It lets you do kind of CI/CD on the model itself.
So it's a centralized place where, regardless of where you are in the organization or what office or what have you, you can just go check out what models. If people say, "Well, how many machine learning models do we actually have in production?" that's where you can find out the answer.
Exactly. And in production is always this complicated thing in a real organization, right? There's often candidate models or ones getting A/B tested. And sometimes you'll have some models in production that flow into other models in production. And so actually supporting real-world workflows ends up being kind of a tricky thing that we're really passionate about. So that's something we're really proud of that's been really popular lately.
And then you mentioned prompts, which is kind of our latest thing that we launched, which is all around essentially using LLMs, which was fun to make because I feel so proud of this. I think all of the major LLMs out there, almost all, were trained using Weights & Biases. So we had seen all these come out, but then we were shocked by how much adoption they've been getting lately.
And people, I think, want to compare asking GPT-4 to training your own model. And we, of course, want to support that.
Yeah, that's really amazing, by the way, that all those LLMs were trained using your products. How exciting, how amazing. Okay. And obviously, LLMs and generative AI, although that's been around for, depending on how you count, a couple of years, this is obviously the big sea change of the last few months. How was that transition for you guys, given that you had built this whole platform pre-explosion of LLMs?
Do you view that as just like a plug-and-play? We have a broad platform, here's a set of new capabilities, or does it change more than that?
It changes more than that. I mean, I think broadly the workflow has a lot of similarities, right? I think if you've used LLMs, I'm sure a lot of your audience has, and you're trying to get them into production, if you're an engineer, there's some things that surprise you, right? These things are very non-deterministic. It's very hard to tell if they're working or not, right? And your workflow is very experiment-oriented, I think, where, with coding, it's so satisfying, right?
Because you kind of write down these steps, and then you debug it, and then it's out, and it's this deterministic thing that does what you think it does. Whereas with LLMs, it's much more like this exploratory process where you try it in different ways and see what happens. And so I think that exploratory workflow is something that we understand well and are really passionate about supporting. I think when you have that exploratory workflow, you really want an easy way to track everything that happens and see if weird things are happening in production.
So those things are the same. But then there actually are some big, big differences. Even the people working with a lot of these APIs are less huge math nerds than some of the people that have been training models for a long time. So you would not believe how much effort we've put into making every possible graph available. We actually have essentially an arbitrary compute engine that you can run before you build a graph because there's been so many requests for different kinds of crazy graphs.
And it's just like our audience loves graphs if you're building models. But I think in the LLM world, the anecdote is something people really pay attention to. So it's like, you're building these prompts and you want to just really see, okay, how is this thing responding? And okay, can you find me some weird examples, and what's happening there? So we spent a lot of time on making text display well and having nice tables of text and different ways to kind of search it.
And embeddings are a really interesting thing where you can kind of find a notion of similarity. But I mean, I think the products end up looking pretty different, even though they support a workflow with a lot of similarities.
Great. What about the monitoring part?
Yeah, I mean, so again, I think the monitoring of LLMs is very nascent, right? We're having this conversation in August 2023, and I feel like anything I say is going to be out of date very quickly, right? But I think probably the observation that is true of models, even worse for using LLMs, is that these systems don't do a great job of notifying you when something's breaking.
So people talk a lot about data drift in the industry, and that's this idea that language changes over time and you want to know that it's changing and sort of have your model notice that and update it. But I guess what I see actually in the real world, doing this for like 20 years, is just incredibly heartbreakingly stupid mistakes that cause models to go totally haywire in production and not tell you.
And I think actually monitoring for that is a pretty unsolved problem. I think people tend to notice these things from downstream user behavior metrics and things like that, and then get really frustrated. And then you sort of fix the one weird thing that you did that caused the model to go haywire. But I think that monitoring these things is tough. And essentially, our point of view is we want to give people ways to build their own custom alerts and their own analyses of what's going on.
Are you seeing a surge in actual use cases of those things? I mean, given that it's all pretty recent. Obviously, there's tons of interest. But just in your almost daily usage on the W&B platform, are you seeing people actually deploying those products in the enterprise?
Well, those are some different questions. So we have seen steady growth in usage since the company started, which feels great. So we're really happy about that. And it's not even like there's one moment. It's just been more and more enthusiasm for ML, and then LLMs, and more and more people use our platform every day, which is awesome. I think that LLMs in particular—we talk to a lot of people, and we don't see a ton of people getting them into production yet.
And I think it's funny, VCs are always surprised when we tell them that. I don't know, I'm bullish on LLMs, but I think we should just acknowledge most enterprises haven't gotten them into production because it's actually hard. And I think this stuff has been obviously working for less than six months. And so I think it's a time where we should be a little bit patient in the short term.
I think we see pockets of applications that work so well that they are getting deployed. And I would say one is search, which just works phenomenally better with LLMs, to the point where people are deploying custom internal search stuff inside their organizations really effectively. And I think chat has just gotten so much better that people are actually really using it for various chat applications. But then there's tons of things that people are trying.
However, I think most of the things are not in production yet. And production is a funny thing too, where I think people sort of gradually productionize these things. I think we see a lot of stuff running asynchronously in the backend.
And so it's useful for something, but that's not exactly production in the same way that serving a classifier for your users live is.
Yeah, actually, as a little bit of a tangent on this, what are you seeing in terms of the most sophisticated organizations that you work with? Like, rough estimate, how many models do they actually have in production, with the caveat that you just described, just to give us a sense for the state of maturity of this industry?
Totally. Well, I'll tell you one funny thing. We have a model registry where we can count how many models people have in production. But the most common question that I get when I go into an organization and talk about the model registry is, what should be the model here? Because in the real world, you typically have a model that's relying on classifications from upstream. So we have organizations that you might think of as having one real application, but inside our model registry, they have hundreds of models that they're actually tracking separately and deploying separately.
And so when I see these crazy studies where they ask CIOs, okay, how many models do you have in production, I have a feeling they have no idea. Or I wonder if you ask the same CIO on two different days how different the answer would really be, because I don't think it's like sort of saying, how many applications do you have in production? I think you probably get a lot of different answers inside an organization.
I mean, that said, we certainly have customers that will have hundreds of models, but I don't know if that really correlates with customers' level of sophistication. OpenAI has been a longtime customer. I mean, I consider them extraordinarily sophisticated, and they have a pretty small number of models in production. So, yeah, yeah, yeah, yeah.
Interesting. And if this whole world of LLM and generative AI is nascent and emerging, what is sort of the bread and butter in terms of use cases at Weights & Biases, in terms of where you see your customers use the platform? What do they do in terms of actual machine learning?
Well, I would say the majority of users on our platform still are building their own models. I think that it is kind of amazing when we look at—I think one thing that surprises people when they look at the company is how broad our customer base is, right? So a lot of people come in, they're like, "Hey, Lukas, why don't you verticalize your messaging?"
Like, you sell into so many different places, but it's like, well, no one vertical is more than like 15% of our usage or revenue or anything like that. So it's kind of cool, right? ML is just really, really broadly applicable. There's a lot of manufacturing use cases of quality control from images. There's making cool characters inside video games. There's looking at health records and trying to predict if someone's going to get sick.
I mean, there really are just so many different applications of ML. And I think when one gets really kind of cookie-cutter, then you get a company that does a great job with it, and people just buy the model from that company that's offering a complete packaged solution. So, sort of Wild West. I mean, even in terms of modeling approaches, I think people are surprised to learn that a lot of people are still running boosted trees or random forests inside of Weights & Biases.
So, it's actually huge. We should probably do a blog post. Still a big fraction of our user base. In fact, it's even kind of growing, I think, because we got our start in deep learning, and then it's been later that more traditional methods kind of came in. And then we see a ton of different—I mean, I think a big trend over the last year or two has been fine-tuning.
It's gotten so much easier with great offerings like Hugging Face making that easier. So now tons of people do. And then, like, using LLMs. And I think another important point I'd make on all this is customers—sophisticated customers—will typically do all of these things, right? So it's not like—I think there's sort of a sense that, like, oh, if you're more sophisticated, you sort of move closer to current approaches.
But I actually think truly sophisticated customers use the right tool for the right job, and they're not afraid to use logistic regression where appropriate still. So we really try to support all the different ways that you might do ML.
Yeah. And a standard way of thinking about the innovation curve from, like, a Valley and VC perspective is typically the big tech companies, the Ubers of the world, are the more sophisticated ones, more willing to experiment, and that propagates through smaller startups and then to the Global 2000. Is that sort of what you're finding, or are you seeing that the Global 2000 are already pretty sophisticated when it comes to deploying machine learning in production?
Well, I mean, I'll say this. Most of our revenue comes from the Global 2000. So it's not like we only sell to the Ubers and Apples and OpenAIs of the world. We do. I think, yeah, I mean, I actually think that machine learning works so well and so obviously well for a lot of applications right now. I think people are kind of surprised to learn that. I bet every might be strong. I bet 99% of the Global 2000 is using machine learning for something that they actually really care about.
Like, I think there's way more applications than people realize. Like supply chain predictions, like sales forecasting. And you might say, well, why do people build custom models for all this stuff? But it's like, why do people build their own tech stack? It's like everybody has a slightly different set of priorities, slightly different set of metrics they're trying to optimize. People think this stuff is core, so they don't want to outsource it.
So I think that our TAM is—I mean, I'm talking to VC, I guess. I don't know if you believe this, but I genuinely believe—I'm not pitching you—but I think our TAM is basically all the Global 2000. They all should be using Weights & Biases, in my opinion.
Yeah, no, wholeheartedly agree. Any difference per vertical? Anything surprising, like different industries being more at the forefront than you would believe?
Yeah, I mean, I always say this. Maybe it's getting boring, but I feel like people are always surprised by this. I think pharma is investing way more in deep learning than people realize. We've been really surprised by our usage. We sell to most of the big pharma companies, and they're really sophisticated, and it's really, I think, working for them. And I think there's a lag between when you develop a new medicine and when it's out into the world.
So people say, oh, there's no medicines developed in this way that are actually FDA-approved. But it's like, okay, it takes five years no matter how good it is. So I think that's a big growth area. And I think the manufacturing use cases are also really—I think it's bigger than people realize. Like even, we sell into automotive, most of automotive, and we do self-driving cars, of course, right? That's table stakes these days.
But I think most of these companies are also trying to automate their factories, and that's another big undertaking. I've also been surprised that there's a lot of gaming applications. I'm not a super big gamer, but some of our employees get super excited when we sell to these gaming companies that I often embarrassingly haven't heard of, but they do. I think a lot of the game development has moved to ML processes in the last couple years.
Yeah, very cool. So, just switching gears a little bit and focusing on the sort of go-to-market aspect of this, selling a broad horizontal platform is awesome when you get there and when you do have all those Global 2000 customers, but it can be pretty daunting. You mentioned the temptation to verticalize or not. How do you—I guess, how do you think about it? How do you sell? What's your go-to-market motion?
Yeah, and I will say, it took us a really long time to get this working, and there was never a moment where we felt good about it. So I do—it's funny to go back to the early graphs. I remember getting excited because we were at 30 weekly active users and telling my early investors, and they're just like, "Lukas, what are you doing? What are you talking about?" But it's actually kind of funny because the one thing about the graph was that it was basically always going up.
So that gave me some hope, but it was really small for a really long time, and there wasn't—people talk about an inflection point, but that really hasn't been our experience. It's been sort of slow exponential growth that, when you zoom out, looks amazing, but when you're actually in it, it doesn't feel that amazing. And the biggest driver of new users has always been mystery. We call it, like, direct.
It's like they just show up to the website and they just use the product. And I think we've tried to triangulate this a zillion different ways, and I think it's basically word of mouth. So I think we actually have a product that people like, and we know this because we measure NPS in a bunch of different ways religiously. And these days our NPS is like 80, right? So I do think there's a lot of love for Weights & Biases.
There are some angry people, and you can reach out and tell me how much you hate my product, and I'll try to make it better for you. But mostly, when we do random surveys of users, they're really excited about it. So I think that actually has been—it's kind of unsatisfying, but that's been the biggest driver. Lately, SEO has become more and more of a driver for us. That's also been very slow to build. But people can publish content, and now we've been doing it for a couple years, people publish tons of content on Weights & Biases.
And it's slowly like Google has decided to prioritize our stuff more and more. So that's typically nowadays people's—we get more people coming to view Weights & Biases reports than to actually use Weights & Biases. So yeah.
As an aside on this, on the sort of developer-focused go-to-market motion, I saw that the Weights & Biases blog is actually open to others to publish on. I'm curious how that's working, what you're seeing. I just think it's a very cool way of doing it, as opposed to just the company talking to the community, having the community be part of the content production under the company's name.
Yeah, I mean, at first I think we did it because we didn't have a lot of resources to write stuff. And I felt proud when people would do cool stuff on Weights & Biases. When you're really early, you kind of spy on your users. We can't do that anymore. But we would kind of watch what people were doing for the first year. There weren't that many of them.
And I was kind of manning Intercom. So they would write in and talk to me and kind of become friends with the people. I'm like, wow, would you mind just sort of sharing this thing that you already have inside of Weights & Biases? So that's been cool. I mean, it's funny, I always feel a little bit like companies like Streamlit do an even better job of community building, and I feel like we still have a lot to learn.
But I do really love giving a forum to people that are doing interesting stuff. And I think a cool thing about the machine learning community in general is it does feel like there's a lot of publishing sort of outside of the typical journal, a lot of sharing of best practices. I think it's sort of the academic background of it that kind of makes people more excited to just post stuff on Twitter. And from my perspective, it's fantastic, even if it wasn't on Weights & Biases.
I think ML Twitter is pretty cool. And the cool thing, like LLMs, in a way, they're even easier to kind of understand. So an ML paper is often hard to parse the math, but an LLM paper where somebody comes up with a new prompt, I kind of feel like anyone can actually engage with it. And some of these ideas are so clever. I sort of expect that to grow.
Yeah, I really agree. What else do you do to win the hearts and minds of developers? What have you learned over the years works, doesn't work?
I mean, the other big thing for us has been integrations. So we have, I think, over 10,000 third-party integrations. We count open-source projects that import Weights & Biases—that's how we measure that. But we used to really beg people, we still really beg people, to integrate with us. And then we try to make sure they stay good. And that's always been—I always kind of feel like that's part of the product, right?
Because if Weights & Biases errors when it's imported into somebody else's project, people are still mad at Weights & Biases if that's the thing that shows up in the error message. So I think we have this awesome team that basically just kind of manages these integrations and tries to help make them good.
That's a huge number. So developers do create the integration for their product, but the Weights & Biases team does quality control and maintenance of it? Is that how it works?
Well, I'd say that's our dream state. We also submit PRs and get them rejected and yelled at. I mean, it's a little messier in reality. But yeah, we want third-party developers to want to integrate with us and feel like that's a good experience. That's really important to us.
And that actually is another place that drives a lot of traffic because you kind of catch people right at the moment where they are looking for a tool. Yeah.
What else moves the needle? Like quality of documentation, have you found, makes a difference?
Well, it's funny, I always feel bad about the quality of our documentation. Sometimes people tell me it's good and I don't believe them. Sometimes people tell me it's bad and I get really, really anxious, but I don't know how to measure that. I think it really matters. I think it really, really matters. But I don't have any metric that proves it. And I think it'd be interesting to actually look at our NPS and see what people specifically think of the docs.
But some people say it's great, some people hate it.
What else has mattered? I don't know. It's funny, I feel like we—oh, lately, actually, it's kind of funny. We got our start, I was teaching some ML classes, or that was one of the things that started us. I actually think this was an amazingly good experience because I was teaching classes on ML and I was using Weights & Biases. And so what would happen was I would actually just watch like 80 students onboard all at once, and a lot of them would fail.
And that would really stress me out because I'm literally in front of the class trying to get them to use my buggy product and they're mad at me. And so I know exactly. And I would get so anxious. And then you'd sort of realize where people get stuck is weird places. Like, it's not—I think—you just, I don't know, it's hard to know. Like, people don't know to copy the API key.
They don't know they can copy it. Or just weird stuff. Or a lot of people run in Colab, and then the Colab breaks and you don't notice. I mean, there's just a lot of weird things. Or Windows terminals—always Windows terminals—always something messed up about that. So that's been kind of—I think that was a really good experience. And then lately we've actually started putting out more courses again.
And that's been a great source of actually enterprise leads, which is always something we care about a little more because they're easier to monetize than random people. But I love putting out courses because it feels like—
Courses being like online courses where if you go to the site, you can just—
We try to do mid-level technical content because I always feel like that's the gap. I feel like, for super advanced stuff, you can take all Stanford classes for free. And for the really basic stuff, I feel like there's a million courses, like Intro to ML. But I always sort of feel like the developer that wants to get an ML model working in production, trained it once, but it's just like, okay, what do I do now?
That always feels like the gap for me. And so we've been putting out our own course material lately, which I love because I feel like sometimes content marketing, people just want to put out more crap into the world. Yeah. And I shouldn't even say that. I think we've—
No, I think you've earned that.
But, very true. But it's so much more satisfying to actually do something useful for somebody. And I do think there's this funny gap in courses where it's not a sexy thing to teach, like Machine Learning 201 for developers, but we're down to do it as best we can. And it's cool, we can make it all free because the lead generation justifies the undertaking.
Yeah, great. Yeah. To close on the go-to-market motion, this is the developer kind of bottoms-up part, but you mentioned that most of your customers are Global 2000 customers. So have you guys presumably flipped at some point where you also need a top-down sales force that goes and talks to those people? Or I guess, what is the full go-to-market motion?
Yeah, yeah, totally. So I think that's what a lot of ML companies struggle with. And my board was complaining about this for a long time. Maybe they still secretly harbor resentment. I don't know if they're listening. But I think the challenge that we have is, we're not like a Figma where we just totally spread virally within an organization. But the good thing is people do use us in pockets for free quite a bit inside everywhere.
And ML, most of ML is inside of enterprise. That's where most of the interesting work happens because that's where most of the data is. And so we have a lot of usage in most of the, what'd you call it, Global 2000, but it really unlocks it to have a salesperson talk to them. So you can totally buy us with a credit card, and I really want to make it so you can do a managed cloud deploy with a credit card.
We haven't quite gotten that yet, but we've tried to make that one click, boom, you have it. But we've really found that once we have some traction inside an organization, if we go in and we talk to the MLOps team and kind of get them on our side, then we can get companies to standardize on us. So, for example, we'll have credit card usage inside, I don't know, pick your—like Apple, for example.
People think, like, oh, it's impossible to sell to Apple. I don't know, if you have a useful product, they can buy with a credit card, they will buy with a credit card. But then when they get to a certain scale, they're like, wait, I want security features. We need stuff. And then it's just an easier sales conversation. I'm acting like this is also trivial. It's definitely not.
But that is our dream motion. And it really does, if you squint, that is what's happening. You get traction, a salesperson talks to them, four months of random IT bullshit, and then we get a bigger install.
Slash procurement, slash legal.
Yeah, yeah, exactly. Exactly.
Who have you found, for a machine learning tools company, is a good archetype for a salesperson, especially the first few? Are they engineers turned salespeople? Are they salespeople who are more technical? Or do you have a pod approach where you have an AE and a technical person with them?
Ooh, I think a lot about this. I've been thinking about this for like 15 years now. I mean, I took a lot of my good reps from my last company because I think they kind of saw who was good at this. I'll tell you, when I interview an AE, I'm kind of asking myself one question, which is just like, if my good friend was trying to buy Weights & Biases, would I hand them to this person?
And there's so many different ways that someone could be good at that, right? I think if they're nice and humble, and they'll actually listen to what the person's saying and treat them well, that's a great way, and you don't have to be super technical. I mean, some of these reps are amazingly technically deep. I don't know, we have this VP of business development, and I feel like you ask him about a random open-source software, he'll give me the one sentence on what it does.
It's amazing. I always feel bad, like I should be more on top of this stuff. I don't know how he does that, right? But that's actually a pretty useful thing. So I think it's really important that they bring a kind of humility and curiosity to the table, but it's hard to put your finger on. I think I've just seen so many different ways to be successful as a sales rep.
It's not like I have one thing, but I do think I have a pretty good instinct with that question that I'm asking myself of just like, what's going to happen when I hand my random friend who's a lowly ML engineer at some random enterprise—are they going to like that experience?
Yeah, I love that question as a test. I think that's really, really good and interesting.
Nobody wants to golf or anything. I mean, that's for sure, right? Like, a super aggressive salesperson—some of my salespeople are really competitive. I am actually really competitive myself, but they kind of suppress it in a way that I think works with a technical audience. I feel like a lot of engineers actually are pretty competitive, but you don't sort of lead with that.
Like, you don't get crass. Nobody wants to do any of the—nobody even wants to really go out to dinner, right? They just want to get the facts quickly in the same way. Yeah, yeah, yeah.
Yeah. As a leader, in this crazy, crazy time for AI and all the hype and all the things, how do you, first of all, keep track of what's going on, which feels like just that is a full-time job for all of us who care about this domain, and sort of filter through the noise and decide whether to panic, not panic, act, take a breather? How do you make sense of all of that?
Well, I think one thing that's really important is to stay current. I always think about, okay, I need to put pressure where other people are not putting pressure, right? So the board, the company will put plenty of pressure to hit quarterly ARR numbers. It's really important, but there actually are mechanisms that put tons of pressure on that. So I think where I need to apply pressure is like, hey, let's think long-term here.
Let's be disciplined and honest with ourselves. But the more important thing is, make a better product if you can honestly do that, right? And I think the same thing for my own management of myself. It's like no one's really gonna notice if I spend a week and I'm not trying new technologies and I'm not learning. Nobody's gonna tell me to do that. People are gonna tell me to be on their podcast.
They're gonna tell me to be on a sales call. They're gonna tell me to do all these things. But I feel like my job is to kind of fight that. And I carve out time in two ways. One is, every Wednesday I really have no meetings, and I really, really try to hold on to that. And the other is, I learned this from Lew Cirne, the New Relic CEO.
I remember, he's running a public company. I talked to him. I think he told me he took one week off a month to just write code himself and enjoy his own product. And I was just like, wow, if a public company exec can do that, I should be able to do that. And so I really try to take blocks of weeks off, and I just really, really try to fight all the pressure that, as I'm saying, I haven't done it in a little while.
I try to do it once a quarter, but I haven't done it in the last quarter. But I think it's really important. I mean, I also think learning by doing is my only good way for learning. So I just, I don't know, if you tell me about somebody's product and I haven't used it, it's kind of in one ear and out the other.
But once I use the product, I feel like I understand it. So I end up, I have a very inefficient learning process, I think, where it's like I have to use Airflow to understand Airflow. I have to use Dagster to get it. But I think that I really push myself to do it.
Yeah, great. And speaking of products in that general space, whether you've tried them or not yet, taking your W&B hat off for a minute, what do you find interesting, intriguing, again, whether you've done it or stuff you want to explore on your to-do list in that broad generative AI, certainly, but broader ML space?
That's a good question. I think Ray, Anyscale, is a really beautifully done product. It's one of those ones where I admire it, wish I made it. I feel the same way about Streamlit, where it's just like, I feel like it's so beautifully simple, its onboarding is so tight, and it's just so delightful. I think dbt—boy, I wish I made that. That was a good idea.
Damn. What else? I think Keras has kind of fallen out of favor, but it is such a good product, documented so well. I feel like it's just designed with this amazing interface. And of course, we see more Lightning now. And I mean, obviously PyTorch is, I think PyTorch is phenomenal. Everyone thought TensorFlow had won.
Yeah, totally. And I feel like PyTorch came in. I feel like I really am trying to study this because it's such an amazing case study where I don't think they really—I think people tell the story of the silver bullet of the eager execution model. But I think the reality is they just built a product with so much more empathy. I really think that. I don't know, I always liked using PyTorch.
Everyone else felt the same way, where I was like, I think TensorFlow, I should be using it. All their customers would say this, like, oh, probably all your other customers use TensorFlow in production, but I don't know, we just like PyTorch. We don't know if it can be put into production, but we are. And then the PyTorch team's asking us, like, hey, do you have any case studies of stuff in production using PyTorch? And we're sort of like, all of our customers are doing it, and feeling a little guilty about it.
So you probably have a brand perception issue, but you've made something just really nice. And then I think Lightning did a great job on top of it. So I don't know. Yeah, I mean, some of these products are really great. And they're great in different ways.
I love that word you used, empathy, as a design principle. I just think it's so beautiful and elegant.
I know everybody talks about it, but it's so hard to actually do. Because if you actually empathize with your customers, they want such annoying things, you know what I mean?
All right, so maybe to close, what's on the docket for Weights & Biases in the next, I don't know, six months, year, in terms of roadmap releases, anything you can or want to talk about?
Totally. I mean, so prompts is a big deal for us. So we're trying to make LLMOps tooling. Evaluation is something that we really want to explore, and I think it's really important. We're putting a lot of effort into monitoring LLMs and monitoring non-LLMs. That's a big undertaking for us. And then we have this really cool thing that we call Weave, where you can do pretty arbitrary data explorations inside of Weights & Biases, and we'll kind of handle the compute at scale.
So that's a little bit more of our moonshot. But that's also something that we're investing in, kind of opening up our platform that we've been using internally for years to let people build their own ML applications on top of us.
Oh, very interesting. Okay. All right. Well, that feels like a wonderful place to leave it. Thank you so much. Where can people follow you, obviously learn more about the company, socials? Where should they go?
wandb.com, or you can follow me at @l2k on Twitter.
Okay. All right. Terrific. Thank you so much, Lukas. We appreciate it. Awesome.
Thank you so much. Thanks for joining us for The MAD Podcast. We're back here every Wednesday with new conversations with leaders in the machine learning, AI, and data space. And if you liked this show, you can also find a video recording of not only this episode, but many, many more over on the Data Driven NYC YouTube channel. Thanks again, and catch you next week.