So I thought a fun place to start would be to talk about the big recent news on the funding front. So you just raised $1.4 billion, I believe. So could you maybe talk to this and take us behind the scenes about how the round came together?
Yeah, so first of all, it's a real honor, and we really appreciate the vote of confidence. I mean, the AI space is obviously hot today, but it's still not an easy funding environment for funding these days. So we don't take it lightly. And with that said, there was such incredible interest in the company, and I'll speak about AI21 later, but it's just a very unique company in this space, not just working on deep technology, but having end-user-facing applications and a platform that serves enterprises.
So I think investors looked at AI21 and definitely saw an opportunity to have a substantial player in this space. And so, the dynamics of the round, I mean, Pitango and Walden Catalyst were just great investors and partners from the early days of the company. And then we've kind of assembled a few more financial investors that joined the round, and also a couple of strategic investors, particularly Google and NVIDIA, that also joined the round as strategic investors.
And the idea behind this is not just investing in the company, but actually forming strategic partnerships with both of these companies.
How important is fundraising strategy, do you think, at this moment in the generative AI market? Obviously, headlines have been dominated by massive fundraisers, and this is certainly a very large one, but is it a little bit of whoever raises the most wins because you need to spend money on GPUs and attract super-expensive talent? How do you think about it?
Yeah, I think one should think about it strategically. There's a lot of excitement, and I think the excitement is for a reason. What we're experiencing right now, it's just a huge platform. It's a huge technology shift. So it's going to have impact, and it's going to have economical gains. But I think also we should be, I mean, we as company leaders, we should be responsible and think: how do we grow the company healthily and build business around it?
Like, we build technologies that are transformative, but we also build businesses, and we need to constantly calibrate the progress of the business side to the valuations and to the amounts that we raised, because each step in this road just creates an expectation for the next step. So I think, again, as young companies in this space—and young relatively to the incumbents—we should be very responsible here and think about how we grow these companies. And in terms of GPUs and compute resources, and I'll speak about it later, I think there's a more practical approach.
I mean, we can build these massive models that are very generalizable and can do many things, but we can also be more thoughtful about how we train even the largest models. But we also can be very thoughtful about how we deploy these capabilities in production and how do we kind of make smaller models that are much more economical and are optimized for use cases that the market is actually ready for. And that kind of aligns with the philosophy of how do we grow the business gradually, how do we kind of fit the technology and its capabilities and its cost where the market is and what's the level of maturity.
So I don't think it's a healthy way to think about it, that you'd win if you raise more. I think you'd probably want to deploy the capital very smartly and make sure you're capturing enough market in the early days so you could be a substantial player longer term.
Okay, very, very good. All right, so that's the most recent news that you announced maybe a couple of weeks ago. So congrats on that again.
I want to go back now to the other end of the spectrum, like the very beginning of the company. So as I was prepping for this, I actually read that the company started in 2017, which for the world of NLP was a big year, but you were very much at the beginning of that big NLP wave. So how did it all come about?
Sure. So it all started—we started a company, Yoav and myself. I had a technical background. This was my second company. The first one was an analytics company in the networking space that was basically acquired by another Israeli company called Cellwize that was eventually acquired by Qualcomm. And in 2017, I was extremely curious about AI. I didn't have any background in AI, but I was very curious about this space, and actually I had a few ideas. And I kind of randomly—it's an interesting story by itself—but I met Yoav, who's my partner.
And Yoav's background: he was a professor at Stanford for almost 30 years. He ran the AI lab and actually started—this is his fifth company, and all of his previous companies were acquired. And when his last company was acquired in 2015, they decided, the whole family, to move back to Israel. And that's how we met, and we joined forces to start AI21. And shortly after, Amnon Shashua joined us as the third co-founder.
He's a founder of Mobileye, right?
Exactly. And Amnon is known for being the founder and CEO of Mobileye, the autonomous driving company. We all started the company around this very grand vision, or at least an observation, of where AI is heading from a historical perspective. Because back then in 2017, and actually even more so today, AI was all about deep learning. It's almost a synonym. And we believe, and we still believe, that deep learning is amazing. And with that said, it's a necessary but not sufficient component.
And if you truly want to realize that vision of having these very smart assistants that help us make decisions and help us in our daily routines, like we imagined a virtual assistant for finance or a virtual assistant for legal or a virtual assistant for medical advice, then we really need the level of reliability. We can't just use statistics. It has to deal with some problems that are more deterministic by nature. And these problems were basically solved or dealt with back in the '80s, where AI was mostly about symbolic and expert systems.
And we say that the true next step of AI is actually: how do we merge these two approaches together? And these days it's called neurosymbolic AI. And so we started the company because we said, hey, the future is going to be neurosymbolic. We want to build a company around it. And along the way, we want to find smart ways to commercialize the technology and basically provide value for individuals and companies along the way, with the understanding that we really want to move the needle here in terms of the technology.
Okay, very good. So fast forward to today, I'd love to spend a good amount of time talking about the several products that you offer. So as a high-level tour, my understanding—and stop me if that's not correct—is that you can think of the business as three different components, obviously all interrelated.
When you look at AI21 as a whole, there are really three components, as you mentioned. One is the core technology, the foundational models and the systems that we build. And then we have two product lines. One is applications that we build ourselves and deliver to end users, applications that help people better read and write. And the second one is an API platform for developers and for enterprises that helps them take these models and systems and embed these capabilities within their workflows and applications.
So this is the three-headed beast: the core technology, the applications, and the platform play. And going back to the neuro-symbolic piece, the ultimate goal is: how do we make these systems actually reason? And there are several weaknesses. I think the most common one, and the one that people mostly speak about, is hallucinations, that these models tend to make up stuff. And I think this is just one area of weakness.
I think sometimes when you want to prompt the model, you want to use these models to get to a conclusion. You want it to reason, you want it to maybe combine several pieces of information or give you an answer based on different data pieces. Then you see it falls short. It just miserably fails in many cases. And that's just a showcase that these models are not good at reasoning, at least not reliably the way we would expect them.
And there are other components like planning. Some of our thinking or our thought process is not done by just, you can think about this as a single turn. Sometimes we process things and we decompose them and we run some sort of a plan to execute and solve the problem. And these models are obviously not designed for that type of activity. And so there are these core fundamental issues that need to be solved. And I think we've made a lot of— we as an industry made a lot of progress in making these very powerful models that are useful.
They cross a usefulness threshold. But now I think this is the time, and I believe ourselves and many others in the industry will start to compose these models in various ways. And this will probably be a shift, since two years from now, I guess we won't be excited and we won't be speaking about large language models. We'll be speaking about AI systems that use maybe several language models and they orchestrate this problem solution in such a way that is more reliable and more deterministic.
And we'll see, and we as AI21 will have very exciting announcements this year about this general direction.
It's a fascinating topic. And I believe there's been this whole debate between Yann LeCun and Gary Marcus on potentially the advantages and drawbacks of deep learning and how to combine them with other things. And I think there was this whole research track of Josh Tenenbaum at MIT as well around this. But the layman's view from my perspective is: I heard one of the key difficulties was, how do you actually combine those two things and know when to give priority to one versus the other?
And the early results of some of the people I heard had tried was, in fact, you still got better results with just super-powerful deep learning, even when you tried to build in this stuff. In other words, the hybrid system performed not as well as just pure deep learning. I guess, what do you say to that, to the extent you can talk about any of this?
Yeah, to be modest here, I think the jury is still out, but I have a sense that if you truly want to deal with the long tail and the problems in life, and if you do want to deal with natural language robustly, the long tail is extremely long and diverse. And my sense is that the way to move forward, the way to make these systems more reliable and transparent in a way, is not going to be by creating a more powerful model, like more parameters or tweaking the data.
And there are definitely performance gains that you can keep on squeezing, but it'll eventually fall off the cliff and we'll face the same fundamental problems we are facing with the strong models today. And we're already starting to see diminishing returns. You take the architecture, you increase the size of the models, and you see that at these levels of scale, we're dealing with diminishing returns. We start seeing it. So I think it shows that in order to fundamentally resolve the problem, we need to take a different approach.
I think there were actually quite recently success cases where the hybrid systems actually reached better results than the end-to-end types of systems. But again, we'll see. I think there is another aspect to it. I think we need to be somewhat pragmatic and also consider the economic considerations here, like taking a huge model and serving it at scale. It's costly and it's so inefficient. And even with the pace of hardware progress and advancements these days, I still don't see it as a practical way to move forward.
So I think this is also an important consideration.
Yeah, this is really a super interesting discussion. As I was prepping for this and listening to different things and reading different things, I don't know to which extent people truly appreciate that yet about AI21. And people, I think, tend to lump you guys into the, oh, well, that's a direct competitor to OpenAI and Anthropic and whoever, which is probably true from a business standpoint. But the approach from a technical standpoint is actually very differentiated because, as far as I know, all those other players are going for—
Let's build ever bigger, and also more efficient, but bigger large language models. So, very interesting. Okay, thanks for sharing. So how does that translate into this family of models? I read you have Jurassic, which is now Jurassic-2, and I think you had multiple models, but you're down to three. Maybe talk about that.
Sure. So we have Jurassic-2, our family of general-purpose language models, constantly improving almost on a weekly and sometimes monthly basis. And it's instruction-following, chat-supporting language models. And these are in direct competition with the other models in the space, like OpenAI's and Anthropic's. And they have some advantages. Some use cases they better support. One of the latest studies that was done by AWS showed that actually our models are top-performing at RAG, at retrieval-augmented generation.
And this is a very common, typical use case in the enterprise. So even for our general-purpose language models, we're trying to steer them in directions that would be helpful or useful for our customers. And the Light, Mid, and Ultra just kind of— it's a trade-off between the size of the models and their performance under various tasks and costs, basically.
And cost and latency, that's one side of the equation. And the other side is size and quality, basically.
Yep. And so all of those are proprietary models and you serve them, right? So, okay, the use case is I'm a developer in an enterprise and I want to have access to the actual model itself. I just call it.
And then, just to play back what you just said, depending on my needs, how do I know which one to use? I mean, obviously the parameters around latency and price, but are there certain tasks where I know I'm going to need the bigger one versus the small one?
Yeah, I guess typically, a Light is great at extractive or classification tasks. A Mid is for short-form generation or RAG-type use cases. And Ultra is one that is very good at longer-form generation, taking long context and referring to it, and having the ability to attend to more information and be consistently cohesive with it. So I think basically people should try. I mean, when they have the use case, my first recommendation to customers is, even on a small scale, do your best effort to create an evaluation set that represents the problem or the distribution of the problem you're trying to solve.
Once you have that in place, you can get a sense for our models and also other models in the market and see what works best for you. Another thing I should mention is that we are taking a very neutral approach. So we are delivering our models in many different ways. We have our SaaS platform, that is actually multi-cloud, and we also serve our models through various types of integrations. So with Google Cloud, you can basically consume our models through the Marketplace and in the future through Vertex.
And in AWS, we offer our models through SageMaker for customers who'd like to deploy these models inside their VPCs, and also in Bedrock. So we offer this optionality, and I think this is actually an area where I was surprised how much customers care about having the portability and capability of working with a single vendor that can run under different clouds, but not only cloud, other environments as well. And we just announced a few months ago our partnership with Snowflake, and there will be others.
And I think that once your LLMs are integrated within a wide set of ecosystem partners, that creates a huge advantage for many customers who are taking this into consideration.
Great, great. All right. So that's the foundation models part. Let's talk about the first category of application, the task-specific models. So just again, reviewing my notes, you talk about text editors, retail software, knowledge and support platforms, content management systems. So those are discrete business applications. So how does that work? And presumably they're fed by Jurassic-2 models underneath. Maybe just go through how it all works.
Yeah. So the market has just shifted from sporadic exploration to massive experimentation in the last nine months. And we've learned a lot from the market. And what we saw is that 95% of the use cases in the enterprise are actually pretty specific. I mean, there's barely any areas where you'd want an open-ended chat system that can address any query that you have in mind and be that generalized. Most enterprises, when you think about how to incorporate this in a workflow, how to incorporate this in an application, have a much narrower use case in mind.
Very powerful, very valuable, but narrower. And then we came to the conclusion that maybe using the general-purpose model as is may be overkill. And we should take a different approach. And we think about it as a matrix. You have the tasks as the columns and the industries or domains as rows. And then you have summarization of financial reports or text generation with certain types of constraints in the retail industry for producing product descriptions, right?
So you have these Lego blocks, where each piece is really good at a particular task. It's also a customizable piece. It's not a closed box. You can then customize it to certain stylistic guidelines or other knowledge that these models need to incorporate. But the thing is, technically, we don't deliver these as just models that we fine-tune on specific data. I think that's a very simplistic approach. We actually wrap these models into systems that really take care of the reliability part.
So when they generate a result, there's a whole verification process, and there's also verification for the input. So you cannot just prompt-engineer it the way you do with a general-purpose model. The system expects a certain type of input, and the customer expects a certain type of output. So you can have verification components in place. And that creates a very powerful—I mean, from testing we've done, this task-specific approach just wins over every general-purpose model when you test it in a specific setting.
And we also find it as a very effective way to market, just because customers understand it better. And you can think about the large language model as a raw engine, but then these Lego blocks as systems, as cars, right? They're much closer to a ready solution that customers can actually start embedding within their systems. So we're seeing a lot of interest around this approach, and we'll see. I mean, we'll have to see where the market is heading.
I think that the opportunity is just huge. I don't think it's a winner-take-all kind of market, and I don't think it's a single approach that will be a winner here. I think different approaches will work for different types of buyers. And maybe as an analogy, and I know a lot of people give these analogies, it's sort of what happens with the database world: not a single vendor and not a single approach, I guess.
And what's the third part then? Let's talk about contextual answers, and what would those do?
So contextual answers is actually one of our task-specific models.
Okay. So that falls into the second bucket.
And maybe just describe what it is, and then we'll jump into that third bucket. Yeah.
As I mentioned, it's a module that—and again, it's a system. It's not just a large language model. It's a system that is optimized for question answering. So let's say you have a large corpus of documents or a large repository of documents, or even structured or semi-structured data, and you want to query these. If you want to implement it by yourself, you need to compose different components. You have to have a large language model, and you may need to tweak it.
And you need a vector database, and you need to embed the target documents or data that you're interested in. It sounds easy, and a lot of people are starting to build these demos by themselves, and the demos look great. But then when you go to production with these, you face all sorts of problems, from performance and also quality—the quality of the results and reliability of results. So Contextual Answers is actually a box that includes all the components, and it's optimized for question answering.
You give it a set of documents or a knowledge base, the system gets a question, and it generates an answer. And the answer has to be based on the information, or reason, so to speak, on the information that it was given. And if the information doesn't exist, it will tell you, "I don't know." So it is optimized for not making up stuff. And also, once it generates an answer, it will tell you where this is coming from. Why did I answer this?
So it kind of reflects the reason or the source of information for the answer. And this is a very typical use case for a lot of companies: question answering, grounded question answering. We build these Contextual Answers systems that basically, very easily, in just a few lines of code, you can get to a place where you have that system up and running and optimized for that use case.
Okay, great. And now that third bucket.
So yeah, we discussed the models and the platform where we serve the models, and the application we built called Wordtune, and we released it about three years ago. It's a reading and writing assistant. So you can use it to summarize anything you read and ask questions about anything you read, or write things. It will help you draft, it will help you ideate, it will help you rewrite. And it's not delivered in a chat interface, and we did it deliberately because we don't think a chat interface is definitely the most optimal way to consume and produce information.
And so it's really a suite of reader and writer apps, and also an extension that is basically acting as a system that works wherever you go. And it's a pretty popular application. It has more than 10—so we launched it three years ago, and it already has more than 10 million users, and it's continuously growing. And it also has a very solid business model in place. So you have some daily quota of usage, and if you want to use the system intensively, you need to upgrade to the premium tier.
And it will keep evolving. We think there's a lot of room for innovation in what we call AI experiences. It's not just about the technology and how to improve the quality of the results. It's also how do you deliver it in a compelling and useful way for the customers. And that's something we're putting a lot of emphasis on. From the early days when we started the company, we put a lot of emphasis around it.
Okay, great. So let's talk about use cases in the enterprise a bit more. You just used that expression, that sentence, a few minutes ago, which I loved, which is that the industry has gone from sporadic exploration to massive experimentation, I believe you said, which is a wonderful way of putting it. From your perspective as somebody who's in the trenches every day, what are some of the top use cases? And maybe talk about some of your customers and some of the case studies there.
Sure. So it's really across industries, from cool startups to healthcare companies. And it's really diverse and interesting ways of using these models. I think I'm most excited just to see how these are employed to enhance productivity of knowledge workers, of all of us, basically. So we work with a few financial institutions, and they essentially are using these tools to make the information—so a lot of these financial institutions have repositories of research, huge repositories, could be millions of documents.
And then there are all sorts of roles in the organizations that need access to that information. And typical search won't do it. Sometimes you have a question and the information resides across many documents, and you need a system that can pull together all the pieces of information and stitch them together into a coherent and relevant answer. And it is quite magical. I mean, we've seen some of the analysts say, "Wow, this would take us like three or four hours."
Now with this tool, we're doing it in just a couple of minutes. And because we never had the chance to think about these specific keywords to search or these specific lenses to look at it, we wouldn't have even reached that conclusion. And so I think what makes me excited is that it's not just the productivity gains in terms of time, it's just the new possibilities it introduces, just new access to information that wasn't there before.
But also, of course, it can be measured by time savings and productivity. The other thing, relatedly, is in the medical space. Think about drug discovery and all the research there. Again, it just unlocks the opportunity to gain new insights using these tools, extract new information, or just deliver new information. You can think of summarizing information in a particular way that is very useful for the researchers and analysts, that just surfaces new types of insights.
So, again, many, many types of use cases in many industries, but I'm most excited by the ones where it just immediately clicks, that you can even intuitively feel the value. You barely need to measure it. It's just so obvious. Yeah, very exciting.
Where do you think we are in the arc? So, from sporadic exploration to massive experimentation. So clearly what's next, hopefully for all of us who love this industry, is massive deployment at scale in production everywhere. But where are we? Are we starting, at AI21 but also in general, to see people actually deploy generative AI truly in the enterprise at some scale? And if not, when do you think that is coming?
Yeah, I think so. As I said, we're now in the massive experimentation phase. I think there are some very early deployments in production. There are deployments, but they're very small-scale relative to the amount of experimentation that is being done today. I think in six months we'll see gradual rollouts to production for many use cases. And we'll also see people—in fact, we've seen it with a few of our customers and partners—we'll also see the issues, the struggles, the problems that, just from a very basic functionality perspective, we would need to face when you go to production, when you start to roll it out on a massive scale.
Again, costs and resources are scarce, and compute resources are an issue today. So how do you deal with that? And the whole notion of reliability that we discussed, I think it's still a challenge for many of these deployments. So I think we'll see that the market is moving towards more maturity and going from demo to demoware to production in the next six months. And then in the next phase, the six months after, my assumption is that we'll see more and more rigorous types of measurement to understand the impact and the value of these use cases on their businesses.
And so that's, I guess, the next phase. And some of the use cases may not be economically viable, and others surely will. But we'll get that better signal, I think, about a year from now. So that's how I think about the various stages and phases of the market. Great.
Maybe let's spend a couple of minutes on go-to-market and distribution. So you alluded to, you have a partnership with Google, Google Cloud Marketplace. I also read that you have a BigQuery integration, I guess probably similar to what you did with Snowflake, and then a partnership with Amazon. Listening to what you described during this conversation, that's a lot of surface area to cover, right? You have the core foundation models, and then you have the applications, and then you have the building blocks.
How do you go to market? Do you have a sales force?
go bottom up using the Wordtune kind of consumer motion? How does it all work?
Yeah, so we have a very large footprint in the market, and it keeps growing. So we definitely use it as a way for us to drive the enterprise motion, but that's more of an organic or bottom-up sort of motion. But we invest heavily and extensively in the top-down, more traditional high-touch motion. And our strategy is we want to work closely with our partners. We feel it's highly synergistic.
And to work with them towards customer success. We're not here to make flashy demos and hype. We are here to bring real solutions for customers. And I think the way to go and the way to scale would be through our partners, who are the hyperscalers and the system integrators and software vendors. It's a whole ecosystem, and that's kind of the key part of our go-to-market strategy.
Great. All right. So to close, I'd love to maybe take a step back. Obviously, AI21 is a global company with global ambitions. At the same time, it was a company that was started in Israel. And I know that there is tremendous interest in the Israeli AI ecosystem from investors and operators, myself included. To the extent that it's doable, I'd love for you to give us a quick tour of the ecosystem. What should we look up, research?
Who should we be familiar with in terms of companies or universities or researchers, and/or anything you're excited about in the ecosystem, any concern that you have? So whichever way you want to take it, I would love to be educated on the sort of 101 of the Israeli AI ecosystem.
Sure. I think generally speaking, Israel is just an amazing technical talent pool: very creative, very chutzpah mentality. And that serves us very well. I think if you look at the Israel tech ecosystem from a technological perspective, you see that it adopted many times and very fast. Back in the '90s, it was all about chips and hardware, and then thereafter networking and telecommunications. And then it's making these transformations.
And I think if you pinpoint right now, Israel is very well known for its cybersecurity talent, very unique and good at cybersecurity. But we are feeling that there is a movement towards the next wave, which is basically AI. And to be fair, there's also been AI talent here for a while now. There are a lot of computer vision companies. Mobileye is a great example. Of course, there are other computer vision companies here that were built in the last decade or so.
So the DNA and the core expertise in deep learning and so forth is definitely here. Another kind of phenomenon that I think people should be aware of is that when you looked at NLP research, there wasn't a lot of NLP research done in Israel. If you think, like, six years ago, nowadays there's a lot, but six years ago, there weren't many places conducting deep research in NLP. And then there was a wave of great students that went to study in the States, at the top universities, and did their PhDs or their postdocs in the leading universities.
And then in the last six years, they came back to Israel, and some of them are teaching in the leading universities, like Hebrew U or Tel Aviv University or the Technion, and others, of course. And others also kind of blended into the industry, in various companies. And so I think that know-how that they brought is really propagated all across. And when there's a new wave, Israelis really adapt themselves and study this space very deeply.
So I think the combination of knowledge coming from universities and the appetite to kind of get into this new wave—we'll see. I think we'll just see massive talent in the AI space in the next couple of years.
Well, that feels like a wonderful place to leave it. Thank you so much, Ori, for spending time with us. This has been super interesting. Where can people find you online, learn more about you, learn more about the company?
AI21.com and @OriGoshen on Twitter. Follow me, DM me, happy to chat. And thank you for having me. This was fun.
Absolutely, very much appreciated. A lot of fun. Thank you, Ori. We'll see you next time.
Thanks for joining us for The MAD Podcast. We're back here every Wednesday with new conversations with leaders in the machine learning, AI, and data space. And if you like this show, you can also find a video recording of not only this episode, but many, many more over on the Data Driven NYC YouTube channel. Thanks again, and catch you next week.