But no, this is a real serious podcast happening right now. Talking about foundation models feels like it has been, and continues to be, at the heart of the action. From a State of AI 2024 perspective, what do you make of the current state of foundation models from a research perspective?
The major shift is: first, we started with single modality, just language input, text tokens, output text tokens. Then we started to do two modalities, maybe text plus image, termed vision-language models. And that's been really powerful for enabling this kind of renaissance of robotics, otherwise termed embodied AI. And then now we're maybe adding a third dimension, which is audio, because oftentimes images and video come with joint audio, which probably describes words. And so all these modalities fit well together, and models trained with all these modalities can be more powerful.
One shift is this growth to multimodality. Llama has it, so Meta's Llama has it. Anthropic has it, not all the modalities. OpenAI as well. On that topic, I think what's been interesting is China was not in this fight 12 months ago and now is very much in it. There are models like Alibaba's Qwen, and then the spinout from a quantitative hedge fund, DeepSeek, which publishes code models and others. And they've been actually a very lively contributor to open source.
And on some of these vision-language model benchmarks, they perform incredibly well. And that's interesting because, A, China hasn't been, or is certainly not portrayed in the West to be, the most open ecosystem, number one. And number two, they've been victims of sanctions from the U.S. government and other countries that are trying to rate-limit access to either GPUs or EUV, these lithography machines from the Netherlands that are used to make these chips. So despite all this, they're still coming out with good stuff.
And it's interesting that they decided to be very strong open-source contributors. We had Clem from Hugging Face on this podcast a couple of weeks ago, a few weeks ago, and we were talking about this. And it's sort of unclear why that happened. There must be some kind of very smart geopolitical reason.
If he doesn't know, I probably don't know. He's the king of open source. But yeah, it's interesting. And Western companies are using them. I think the data as of a week or two ago from his platform, Hugging Face, is there's probably over half a billion Llama derivative models that have been downloaded. I think Qwen is growing rapidly as well. So that's one part of this debate. The other one is, of course, this constant fight between OpenAI, Anthropic, Google DeepMind, GDM, as being the main contenders.
I think on the one side, you could say the gaps seem to have narrowed in capabilities. When you're looking at these various benchmarks, they sort of look increasingly esoteric to the non-expert perceiver, but the gaps seem to be pretty small. This is outside of vibes and what they feel like. And I think there are consumer preferences for one or the other. But while the gaps have gotten small and Meta has these ginormous systems they're giving away for free, which are interesting in their own right, a sort of satirical take on all this is you could have basically looked at the state of the landscape—who's number one, who's number two, who's number three—12 months ago, and then just deleted Twitter, not read any machine learning news for 12 months, and then basically turned it on again.
And basically seen the exact same thing. In fact, today I think there was some report from Ramp Data, which looks at credit card spend from their corporate customers and breaks down which model their customers are using. And they made the claim, "Oh, look, there's a sort of fragmentation and more models are being used." And so, look, there's happy competition, but really, when you sum the makers of all these models that are getting used, it's still like 80% OpenAI. That hasn't changed.
So it doesn't look like fragmentation to me. It looks like domination. So, yeah, that's, I think, one of the main things. And that probably also goes to vibes. It's sort of like Pepsi versus Coke.
The theme of this conversation.
Yeah, it kind of tastes the same, but it's very hard to convert a Coke drinker to a Pepsi drinker. It just doesn't happen.
You were just talking about Llama. I think somewhere in the report, you have this really funny sentence. You call Zuckerberg the de facto messiah of open source. What do you see around Llama? What do you think? Why are they doing this, and what impact is it having?
Yeah, it's probably one of the best ROI trades of a public company in a long time. And basically the chart shows that from the launch of metaverse, companies saying, "We're going to invest, I don't know how many bajillion dollars in metaverse," to basically, "We're going to stop doing metaverse," when the stock price was just bleeding. And I think various shareholders had written letters to tell them to stop, et cetera. That was like negative $600 billion or something like this in market cap for Meta.
Two trillion—sorry, three trillion—market cap appreciation, which you could argue, who cares if anybody's using this stuff? That in itself is amazing shareholder value creation. Yeah, there's been a lot of downloads. A lot of companies are using the Meta backbone, or at least the pretrained model as a base for downstream tasks. And I think that's been very useful and will probably continue to be useful. I think there's some implementation details about whether it's dense or sparse or whatever, but I think the question I ask myself is, okay, if there are half a billion downloads on Hugging Face of this open-source model and they're pouring a bajillion resources into this, why is it that OpenAI and Anthropic's revenue keeps ripping?
Why is it that it appears that developers and companies still vote for closed models that have more convenience, or are just reliable and fast? You don't have to worry about it. It'd probably be a bad analogy, but why is it that individuals who live in the West who are sort of part of a middle-class-or-up demographic buy iPhones? I don't remember adding somebody's phone number recently that's not an iPhone.
And so that just provides you with a full-stack, nice experience. It just works. It always gets updated. And that seems to me very much like OpenAI. And then the Android is like, you can fork it if you want. You can do some funky stuff with it if you want, but people don't really do it that much anymore. And then the challenge with open source too is, who's going to run it? How are you going to get big distribution?
Who's going to run inference for this? Who's going to host it? And then when you look at Google, they have their own effort for their own models. They're probably going to prioritize that. Amazon also has Amazon AGI. They pseudo-acquired Adept for this and a few others. So I'm not sure if they're going to host it. And then they have the Anthropic relationship. And then Microsoft is in this constant, I don't know, frenemy situationship with OpenAI. So they're probably not going to host it.
We're not going to invest very much in it. So who's left? Probably Databricks. And Databricks and Snowflake, probably more Databricks, will be this enabler of open source. But yeah, it's tricky.
I guess Zuckerberg said a few months ago that he was not going to turn this into an enterprise business for Meta, which, by the way, they did at some point, right? They had a Slack competitor for a few years.
Yeah, Workplace. So it would not be completely unheard of for Meta to run an enterprise business.
Yeah, I don't know, it's hard to give him strategy advice, but it seems like it's a shame they don't have a cloud business that they could at least serve this stuff to customers if they wanted to. But I do think ultimately that just comes back to what's Meta's core business? Ads. And I think in his Meta Connect presentation not that long ago, he gave some stats around how consumers click through at a higher rate. And it was material.
It was like somewhere in the range of 7% to 10% click-through improvement on generative ads, and that they've served like tens of billions of ads in the last month. So that's the ROI right there.
You were mentioning a second ago, revenue at foundation model labs is ripping. That's also a big difference, or a significant difference, from last year to this year. It seems like foundation models, which a lot of people were thinking would never make much money, actually are making money.
Yeah, that's a major one. I mean, I'm probably included. I think anybody who tells you, "I'm going to launch a product, it's going to make billions of revenue in one year," I mean, it's hard to believe, right? Because it never happens. But it did this time. So that's one. And then the second thing is that these models would be so expensive to run, like, there's no margin in it. And that one seems to be changing a little bit, whether it's just the sheer cost drop from the most expensive models, like a year ago, to now.
0.06 per million tokens. Not the same level of intelligence, but not that far off. So I think most people I talk to who are close to the coalface working on inference improvement say that you will likely have similar quality of intelligence for much lower price and much smaller models as we kind of improve things like knowing what pre-training data to feed it, sort of the idea of curriculum learning, like very much the analogy to going to school. You don't get taught PhD-level physics when you're like 10 years old.
You get taught the simpler stuff first, and then things like how to refine post-training, like what kind of examples should you give it? And just given that, I think we probably forget to include this in general discourse, but until ChatGPT and AI became super hypey, there were not that many contributors who actually worked in this field. I think there was some data from various firms we included in prior reports, but it was like, I don't know, 100,000 machine learners in the world, a million machine learners in the world, like really small numbers.
And now I think you have everybody who spent years doing ad optimization, infrastructure optimization, DevOps, all this stuff of making software run really fast, who are now making AI run really fast. And these people are really smart, and they will undoubtedly find issues that AI architecture people will not have been experts in. And so I think that's also what's driving some of the cost reduction in addition to, of course, good old-fashioned price wars.
You were quoting somewhere a chart that actually might be from Ramp as well that shows some improved stickiness of AI applications in the enterprise.
Yeah, it was something like, just directionally, retention after one year in the 2022 cohort was like 43%, I think. And then the cohort from 2023, after 12 months, was like 65%. So that was pretty material. And then we also showed the quarterly billing. So, how much individual Ramp customers were spending on AI products every quarter. And that's roughly doubled, or at least grown by 50%, which is like an argument against, like, hey, this is all demo where people are not really using it for real.
And then in the next slide, or very close to it, we have data from Stripe, which I was looking for last year: examples of this. And now I think it's pretty clear so far. And it basically looks at the 100 most promising SaaS companies on Stripe and the 100 most promising AI companies on Stripe, founded before 2020 or after 2020. And then it charts how long those companies took from their first sale on Stripe to reach certain revenue targets. And it's something like, I think 2020 is obviously not GenAI, but pre-2020 it took like 12 months or something like this.