Welcome, Gustavo. So you're the CEO of ASAPP, which is a unicorn AI startup based right here in New York. Maybe for people who may not be familiar with the company, walk us through the high level. What do you do? When did the company start? Give us a sense for size.
Absolutely. Thank you all for being here. We started ASAPP in 2014 with a very, very simple mission, which is to end bad customer service. We actually stumbled by accident into that mission. We wanted to start a company that would build AI products to solve real-world problems. And our definition of what made a real-world problem interesting was a problem that was, from an economic perspective, very large. The problem needed to be very broken, and it needed to have tons of data so that we could do something.
With that data. And after many months of thinking about different problems that would meet that criteria, I had a terrible phone call with my cable provider that lasted nearly three hours. And that made me feel very unhappy. And it also made me realize they must be unhappy too, because they paid someone to be on the phone with me for nearly three hours. So I Googled, "How much does the world spend on these call centers that people hate calling?" And the first thing I clicked on was a market study.
That estimated the global spend to be about $600 billion per year. And I said, if this is true, this definitely checks the box for gigantic economic size. And then we learned the problem was very broken and actually full of very, very interesting data. And that's how we got started on this crusade of ending bad customer service by virtue of building AI products that allow you to do that. The company today is about 350 people. We exclusively work with very, very large enterprises.
And the reason we did that is unlike, I would say, perhaps a majority of enterprise technology companies that tend to get started with mid-market opportunities and over time grow to bigger and bigger accounts, it was the monster-sized organizations that had the tens of thousands of agents, the tens of millions of customers, and more importantly, the data between those two that we felt was necessary to bring our product vision to life. Great.
Maybe walk us through the state of the market from a technology vendor standpoint, pre-ASAPP. If you are one of those gigantic companies with lots of customers, lots of agents, what do you currently use today?
Well, what's really interesting, I'll give a little bit of context. It is a fascinating problem where you can think of the problem as a three-legged stool. One leg is the company enterprise. Another leg is the customer that has to interact with that enterprise. And another leg is the agent who works for the enterprise. And in this three-legged stool, all three legs are fundamentally broken. Companies that spend billions and billions of dollars really dislike that economic reality. And they're generally spending it for the privilege of that second pillar, the customer hating their guts when it's time to interact with them.
And then agents have an average attrition rate of about 100% per year, depending on the year. I think in the U.S., most estimates have agent counts at about three to four million agents. That's more than there are truck drivers. And now imagine them attriting every year and the operational nightmare that it means for these companies. So back to the question, the technology, the most interesting thing about the technology: one, it's fragmented. B, most of the legacy incumbents are overwhelmingly underwhelming technology companies.
I mean, the best and brightest engineers have historically not gone and worked at NICE Systems or Genesys, and I'm naming names. By the way, some of these companies have built very successful and large companies, but from an innovation perspective, they're not super attractive. And what characterizes the trajectory of the space is everything's been incremental. You'll remember there was a company by the name of Siebel Systems. Siebel Systems was actually, at one point in time, the overwhelming market leader in the space.
And one day Salesforce showed up and became the new overwhelming market leader in this space. And for all the right reasons: cloud instead of on-prem, nice UI with more features instead of some crappy-looking software. But if you survey the agent that spent eight hours looking at Siebel and now is spending eight hours looking at Salesforce, and you ask the agent, how much more productive are you? The answer is effectively zero. So for roughly two decades, all this space has seen is very little incrementalism.
And that incrementalism is driven by an imperative of the company saving money, often at the expense of the customer experience. So welcome.
So let's dive into the product, and maybe starting at a high level from a design philosophy perspective. You all do agent assistance as opposed to agent automation. So it's sort of a copilot metaphor. Is that the right way to think of it?
We actually do both. When we started, our product vision was pretty simple and straightforward. There are some interactions that you can automate; try to automate them as best you can. And whatever you cannot, and that interaction ends up on the lap of a human agent, try to make that agent as productive as possible. Because if you make that agent 2x more productive, it's the same as having automated half of those interactions away. And by the way, while you're augmenting an agent by suggesting what that agent should be responding or what they should be doing, you start to collect a very interesting dataset of the supervised usage of whatever model predictions you're making.
By the agent using or not using those suggestions. So our approach is automate as much as you can, augment the rest, and the more you augment, the more you can actually automate in a bottoms-up way, not in a rules-based way, which is traditionally how things have been automated in this space. So from a product perspective, we started with that vision and we built it first into an end-to-end messaging platform. So we had the observation that when you look at how we use these little things, it's primarily for asynchronous messaging.
When it's time to interact with a company, either those channels seven years ago did not exist or they were pretty bad and adoption was less than 10%. So we said, why don't we build a modern, beautiful messaging interface that makes the experience delightful for customers and makes the agents way more productive, and we get to automate much more? That was good for a while. The problem with that model is that for companies that already had a messaging platform, no matter if it was a 30-year-old piece of junk, the motion becomes having to replace something.
And especially in the large enterprise context, that is very friction-filled. So you can have sales cycles that are 12, 24-plus months. So at some point we decided to modularize all the AI features or capabilities that were in this product into what we call AI services. And this is essentially, if you think of Twilio selling APIs that are pieces of communication infrastructure, we're selling APIs that are AI capabilities that do different things. So for example, the capability of transcribing a call in real time, there's now an ASR API that you can buy.
This thing that suggests to the agent what to respond at every turn of the conversation is an API that you can put onto whatever platform. So today we have seven different AI services, and if someone wants to buy the whole thing, the end-to-end platform, we can also sell that.
Okay. So just to play it back: so you do voice, but also digital, and then your API, meaning that you can integrate your AI into other people's products, but you're also a platform that business users or agents can interact with.
For messaging, we're a platform. For voice, we don't get into a telephony business, so we integrate with whatever telephony a company uses.
Okay, great. So maybe bring it to life for us. If I'm an agent and I interact with ASAPP's technology, what do I do? What's happening?
Yeah, I'll give you an interesting example on the messaging side. For example, if you're chatting with a company, it takes, in most industries, the agent roughly 20 seconds to type a response to whatever utterance you've sent their way. If I can predict what you should respond, and instead of having to type that sentence, you just click on it, it goes from 20 seconds to roughly a second. And then the next logical question is, that sounds wonderful. For how many of my agent responses can you actually predict the right thing, where they're going to go from clicking something that takes a second versus typing something that takes 20 seconds?
And today, in our most mature customers, roughly 80% of everything an agent does is click on those suggestions that the system's generating.
And that is Auto Compose?
That's Auto Compose, yeah.
Okay. And then you have Auto Assist, Auto Summary, Auto Transcribe. What do all of those do?
So Transcribe is speech recognition, and it's primarily speech recognition for the contact center, which has very, very different acoustic properties than general-purpose speech recognition like you can buy from Google or Microsoft or Amazon. A lot of those models initially were trained on data from the assistant devices, and these are devices that sit in your living room. And the acoustic properties of a living room are such that they're not so noisy or complex. A contact center is quite tricky because there are two somewhat random variables.
On one hand, you have the agent that is oftentimes in a very loud and noisy environment. Then you have the customer, which at any given point can be driving home with highway noise, can be walking down Fifth Avenue. So solving for that specific problem makes the problem a little bit trickier, and why we decided to build our own rather than use third-party ASR.
Coaching AI is our newest product. And again, I keep coming back to the three-legged stool and all legs being broken. If you ever call a company, it usually goes like this. You dial some 1-800 number. Before you get to speak to anybody, they say, "This call may be monitored for quality assurance purposes." What that means is that call is recorded. And the company, either due to regulation depending on the industry or just from a desire to quality-manage interactions, is recording all these calls.
And at the end of the month, they have this 25-year-old software from some of the companies I mentioned that allows you to manually listen to those recorded calls. And as you're listening to them, manually score them on whatever set of policies you've determined make good or bad quality in your company. So as a result of that pretty burdensome and silly way of doing quality assurance, on average, less than 1% of interactions end up going through that process. So Coaching AI, essentially in real time, you feed it whatever policies you have, it understands them, and in real time scores all your interactions based on those policies for 100% of the agents.
So it's leapfrogging current quality management quite significantly.
As I think about AI for customer service, I always wonder how one measures the results because, as you said several times, you have your three-legged stool. So I guess it's agent satisfaction, but ultimately what matters is our satisfaction as end users. Like, how do you make sense of it?
Yeah, I think that there's really two primary drivers of value, and depending on the company, they'll think about them somewhat differently. One is cost. What is my overall cost for having an interaction or serving a customer, however defined? The other one is customer satisfaction, however defined: NPS, CSAT, whatever. And those two variables have always been a trade-off. And a simple example is, pick your favorite cable company. I can't name them now because a bunch of them are customers.
So I got to be very diplomatic with the statements that I make. And they've been trying to pinch every penny imaginable out of a multibillion-dollar call center operation. And as a result of that, they've made that cost optimization decision at the expense of customer experience, meaning some of these companies are famously bad at treating you. On the other hand, a company like Zappos is famous for being quite delightful at treating you. You can call Zappos today and say, "I'm going to Vegas tomorrow."
Actually, I've bought no shoes from you, but I'm going to Vegas tomorrow, and I was wondering if you have a favorite bar. And they'll actually help you. So obviously that is very delightful. It is extremely expensive. So most companies have to figure out, am I going to be cable company X or am I going to be Zappos? And we're in that trade-off. Do I want to be? I think the power of bringing artificial intelligence products to this domain is that you can break away from that trade-off.
And for the first time in a very long time, doing the right thing for the company's bottom line happens to be doing the right thing for the customer. Because if you have delightful automation instead of clunky IVR, frustrating automation, or an agent that is three times more effective at dealing with you and more informed and efficient, blah, blah, blah, you're doing both. You're tackling the cost problem while increasing customer delight.
So I'd love to dive into the technology part of this to the extent that you can, or you're willing to share. So in particular, you've been doing this since 2014, which was two years after the resurrection of deep learning, or the acceleration of deep learning after 2012 ImageNet. So you've been at this for a few years, and in particular before Transformers and then before GPT-2. So what did you build then, and how did you think about the evolution and bringing, I guess, the best of what's come out?
So one interesting thing about ASAPP is early on, when we came up with this vision to automate and augment and do that seamlessly, and as you augment, leverage that data to increase the automation, we realized quickly this wasn't a product vision where you can grab components from different pieces, put them together, and build a product. There were a bunch of questions, primarily fundamental research questions in natural language processing, some in theoretical machine learning, that we needed to sort of move the needle forward if we wanted to make this product vision come to life.
And that's why very early on we started building an AI research organization, and we were very fortunate in the people that we were able to attract and the critical mass that we've built in our organization. So to answer your question, because we had that team from 2015 to effectively now, the vast majority of the quote-unquote under-the-hood AI capabilities that we have in our products have been built internally. Because we have this research organization, we ended up doing fairly low-level work where, in early 2017, most of our models were trained on what's called a single recurrent unit, an SRU, which is a neural network architecture that we developed in-house and then we published.
And at some point it was sort of competing with Transformers. And then we hired the guy who ran all NLP research at Google, who the Transformer guy worked for, and we started training Transformers. So it was interesting—
It's a badass move. So, like, we hired the boss.
Well, actually, the guy who invented it is a good friend of mine, and he's a cool entrepreneur now. What we do, too, is for most of our products, we have to train our own models for a number of reasons. The most important reason is, take AutoCompose, what we discussed briefly. So you're having a conversation, you say something, our model predicts, "Agent, here's what you should respond." Go click it. That model gets updated in a conversation that might have, say, 100 back-and-forth messages, thousands of times.
Every time there's a new keystroke, it pings the model and says, based on this new piece of information, this additional keystroke, what do you think is the best response? So in an average conversation, there's thousands of pings to this model. You can say, well, couldn't GPT-4, for example, make those predictions? It actually could, and it'd probably do a very decent job out of the box. The problem is, your conversation will cost you $50 to $100 instead of something that we sell for, I don't know, 30 cents.
The second problem is, if you've all used ChatGPT, sometimes it takes 20, 30, 40 seconds to respond. If I go to the agent and say, "Hey, wait 40 seconds while I come up with the best response for this conversation," I'm not reducing handle time, I'm adding handle time. So the imperatives are cost and latency and accuracy. By the way, if you end up training the same transformers in vertical or domain-specific ways, you can actually get similar or higher performance. So today, for most of our products, we train our own models.
We're starting now to look at using open-source checkpoints and then highly customizing them. And we've even looked at paying for some things with OpenAI. Where we see this going is what OpenAI and others are doing is really a fantastic and spectacularly aggressive commoditization of NLP and some other machine learning capabilities. And unless you really think your stuff is better for religious reasons, then one should be pragmatic and say, it takes me this many people to do this. Well, I can just pay for using an API for these other things.
So what we're currently doing is, depending on the products—and we have seven different products—and our models historically were between half a billion parameters to 10 billion parameters. Now we have much larger ones, but we're kind of building technology that abstracts from this and says, "I'm going to be an orchestrator, and depending on what I need, I'm going to ping different resources where I'm optimizing cost and latency and accuracy." And what we see is a significant amount of effort, some on our part, to figure out how you can get higher or similar model performance to some of the state-of-the-art models with much smaller models. And when you deploy a new customer, do you then train all the models on their data specifically?
Yeah, because we work with these very large companies, we've from day one been set up in a way where we don't commingle data. So we do get better because we have more customers, and when we figure something out by virtue of working with those customers, the models are the same for everybody, but they get individually trained on the data of that given organization.
Are all those customers cloud, or do a lot of them—
100% of them are cloud today.
Okay, interesting. I read somewhere that you have a concept of agent model as opposed to language model. What does that mean?
Well, if you think of what an agent does, they really do two things. They communicate, they're having a conversation with a customer, and they do things based on whatever that conversation is. You might be having a conversation about moving to a new address, and then the agent needs to go and update that address on some CRM or whatever system. So they have a verbal or a language workflow, which large language models can do very good jobs at predicting language and things of that nature, but they really can't do very much in the action space, especially in that action space that's not API-driven, but is literally a UI workflow that an agent's doing.
So we began training models that are essentially embedding all the workflows, the actions that agents are taking, so that we can make predictions of, "Respond this," and, "Go take these other three steps on that system over there." And that's essentially what we mean by agent models, because language alone doesn't solve the problem.
Interesting. All right, switching to go-to-market. So you mentioned you work with some of the world's largest enterprises, JetBlue, American Airlines, Ernst & Young, and a lot of others. For the entrepreneurs and salespeople, go-to-market people out there, any lessons learned selling to those giant customers and those long sales cycles? As a young startup, how do you make sure you don't—
Well, the first one you touched on, which is, you introduced me as saying you were in stealth mode for a very long time. That's a very, very stupid idea. Do not be in stealth mode. There's very seldom good reasons for doing it. So I strongly advise against it because reality is, if you're selling to the Global 2000 or whatever, odds are there's, I don't know, two, three, five people at each of those companies who are potential champions or relevant influencers in that pursuit.
So there's a universe of, I don't know, 3,000 to 10,000 people that better know who you are and what you do and why they should take a conversation from you. And if you're in stealth, you're making—unless you have some magical viral moment of people talking about you by virtue of being stealth, which I think is what we were trying to do in 2015 and '16, which worked for a very limited amount of time. Reality is, at a minimum, go-to-market is as important as everything you do on the technology side, and it might be, in many cases, more important.
What does a go-to-market organization look like today? Are you sort of a classic sales, marketing, customer success? Anything there?
So we have Frank Slootman from Snowflake on our board. So customer success is touchy-feely with someone who wrote a book, and there's a chapter that starts with, "Every time I join a company, I fire the entire customer success department." So we renamed it, essentially. No, I'm joking. But we're really nascent. So today, ASAPP's about 350 people. I would say 80-plus percent of our organization is research and engineering and product, or overall technology roles.
So the go-to-market side of the house is one that we started building much more recently than the historically predominant technology side of the house. Andy there runs our financial services. So we've recruited people like him, who have long tenures at companies like EMC and all these classic enterprise powerhouses at selling. And they come with pedigree and understanding of what it takes to win deals with large organizations. In many cases, you can get that plus relationships. That's a good thing, but reality is, I think every startup will go through a period of maturity in which they go from predominantly founder-led business development-type sales motions to actually the imperative of having to build a sales machine, which is a machine that produces without the founder perhaps being involved in every sales pursuit.
And I think that's the aspiration that all of us should have, is to be able to build— So you mentioned earlier that you built a research-first organization for AI.
Any lessons you can learn about how you do that, especially today in a context where AI talent is so hard to get? How do you recruit them? How do you convince them? How do you retain them?
Great question because it's such a relatively small universe of people that are in such high demand that it makes a question Quite relevant for those who try to do it. And by the way, there's another variable, which is some of these people in large tech have athlete-type compensation in some cases. So I think it starts with purpose, right? I've had conversations with people that were considering, say, a fresh PhD graduate from a top program somewhere that's considering, do I go to Google Brain or do I come join your startup?
And something that I've many, many times said is. Google is probably going to pay you more. You're probably going to be surrounded by excellent people and you're going to do interesting work. What my problem would be with that choice is whether you go there or whether you don't go there, nothing will happen to Google. Absolutely nothing. And it doesn't mean that you don't go there and you write the next transformer paper or whatever. But reality is your impact is extraordinarily small despite amazing work.
Whereas in any startup, our startup, an individual can have a singular impact on the trajectory of the company. So there's some people who want the comfort and all the good things that come with being in large technology companies and others that like the purpose, the mission, and the impact that they can have, I think end up gravitating to our place. So I think that addresses sort of how do you go about getting them? The second question you asked is, How do you keep that?
Which I think one is obviously sustaining the purpose, but ultimately culture. And culture, I think, matters greatly. If you had asked me in 2014 where it was just me in a room, what do you think of culture? What are your values? I would've frankly said like, I don't know what you're talking about. Like, I need smart people that don't wanna sleep and let's go.
Yeah. And, and, and turns out that one of the areas where I probably mature my thinking the most is precisely that. If you, if you, If you manage to be successful at getting someone bright and talented and hardworking to walk through your doors and say, I'm here. Well, the second most important thing is how do we organize ourselves? And in software, everything we do is a result of the people we have. There's no dimension of physical constraint. Like for this building to be successful, they need square footage.
In our world is what's in here, what's in Andy's head, and our ability to align our heads with whatever shared objectives we have. So how do you define the culture and how do you enforce it? I think it's, it's equally important. That's Hiring great people.
And that's a natural tension when you have a bunch of very smart AI researchers between research and I guess applied research and people wanna write papers and open source, but you need them to build stuff that you can sell. How do you manage that and how do you measure the output of a research organization?
Yeah, I can confidently answer because I have scar tissue on that topic. And I think it starts with, you have to be extremely upfront about what you want from people. If you want to be an academic lab in industry, by all means, go for it. If you don't want to publish papers and you want to be building products, be upfront about it. So I think just being clear about what you want people for is a fundamental requirement. The second one is, and this is something that, that I think all companies generally fall into the trap of.
The kind of the hierarchy of what team am I in or what my role is or what my title is. And the more that you can suppress it and, and, and kind of remind people we're here, there's no such thing as individual success. There's no world in which I'm successful in ASAP and Andy's not successful in ASAP. And by the same token, if he's successful, it means we're adding a lot of customers and I'm also more successful. So the recognition of this is a, it's a team effort and a team sport, and, and we all play different roles on that team.
Uh, I think it's, it's one of the best pieces of advice I was given by the CEO of a very, very large technology company was you're gonna come up with your version of this culture thing and you're gonna articulate it and you're gonna get in front of your company and then you're gonna say it and, and hopefully you were well prepared and it comes across in a compelling way and you think you're done, but then you're gonna have to say it again, again, again, again, again, again, again.
And when you think you've said it way too many times, Say it again. And I think it's fundamentally true that no matter how good and well-intentioned and motivated people are, the reminder of why we're here needs to be constantly, that drum needs to be constantly beaten.
You have an incredibly illustrious board. So you mentioned Frank Slootman, the CEO of Snowflake. But before that, and I believe at the very beginning of the company, John Chambers, a legendary CEO of Cisco. Uh, is an investor and a board member. John Doerr, legendary VC at Kleiner Perkins, uh, famous for many things, including, uh, leading the Series A at Google. So first of all, especially for those first 2, how did it come about? I mean, 10 years ago, I assume you were not that, you know, many years after college or school.
Like, how does a young entrepreneur manage to convince those industry legends to come on their board?
Well, one, a gigantic amount of good fortune and luck. And 2, when presented with that good fortune and luck, back to that purpose, what are we doing and why does it matter and how are we going to ensure that we're successful at it? And I think we were from an early stage capable of articulating 2 things. One, the immediacy of this problem of how enterprises interact with their customers. Is arguably the largest problem in enterprise technology. From a dollar perspective, there's more money being spent in call centers than there's money spent in cybersecurity, storage, and databases combined.
So it's just a gigantic amount of money. It's fundamentally unsexy, right? So it's kind of like, really? Call centers? Yeah, it's gigantic. There's more call center workers than truck drivers in the United States. So very large. The second thing which helped with someone like Doris Prudy technically sophisticated and not necessarily interested in just some discrete problem, no matter how big it is, is the idea that building AI products that can be deployed to a large number of people, and by virtue of how you build them and who uses them, those products can be improved in somewhat autonomous ways, ends up being a really interesting long-term vision for us where we're starting with this call center problem.
But really what we're doing is building products that automate and augment a given set of workflows. And at some point we might decide we want to transfer everything we built to this second problem that has somewhat similar attributes or a third problem. And our belief has always been that whoever does that over a long period of time is probably gonna have an opportunity to build a very, very large company.
So maybe to stay on this topic, but to close, and then I'll, uh, open to questions. Uh, any lessons learned? From them, I imagine it must be an incredible experience, but also a pretty intense experience to work with that group.
Yeah, I remember our first board meeting with those people. I showed up and I wasn't thinking much. I've never had a board meeting up to that point. So I really called a VC friend of mine who wasn't in the company. So I was like, what is a board deck, essentially? Can you have a template that you can anonymize and share? I can spend an hour or much more probably sharing anecdotes. There's really three dimensions of help that I think a good board and supporters can give you.
One is being a sounding board that recognizes you're driving the car. I'm not driving the car. I'll tell you what I would be doing if I were driving it, but even if you disagree with what I'm saying, I recognize you're driving it. So God bless you. The second one is tactical advice, and that's where somewhat illustrious people—I'll share one anecdote. A few years ago, Apple released something called Apple Business Chat, or Apple Business Messaging, or now it's called Apple Business Chat.
Essentially, the idea is you can go on iMessage and have a conversation with a company instead of a person. And when they announced this, they had select partners that you could—they weren't giving the software to enterprises. They kind of deal with Salesforce and three other companies. And we had a few very large customers that, the day the announcement came out, said, well, we have a problem because you're our provider, but you're not on that list, and we really want this Apple thing.
And that really happened two or three days before a board meeting. And I asked John Doerr, do you think you can help me with someone at Apple? Because there's this Apple Business... He's like, what do you need? Can you drop me a note? And then 12 minutes later, I had an introduction to Tim Cook saying, can you help this kid with, like, this thing? I'm like, we were there. So that is very helpful. Very, very helpful. And then the third thing, and this is going to sound kooky, but there's something about showing up to a board—or actually, forget about the board meeting itself or not—it's being accountable to these people.
If you kind of think about it, it should be a very energizing thing when you recognize the privilege you've been given and the responsibility that you have. So I think there's a motivating factor that probably might supersede all other benefits.
All right. Incredible. Thank you so much for sharing.
Thanks for joining us for The MAD Podcast. We're back here every Wednesday with new conversations with leaders in the machine learning, AI, and data landscape. If you like the show, you can find the video recording of this episode, along with many more, on the Data Driven NYC channel on YouTube. You can find all the important links in the show notes.