At the time when Gemini started at Google, that's when I left. So I wanted to create a small research environment that reminded me of the early days of FAIR or Google Brain. So, a very small team, elite, no distractions. Sorry to say that: no product managers, just research scientists, no emails, just locked in a room with the machines and focused on science. And in particular, the goal for me was really to keep working on fundamental research and keep pushing the field forward and training students and so on, because I felt very grateful to have been able to do research in such an open environment.
It was also obvious to me that, and for this I think I agree 100% with Yann LeCun, what made AI dynamic and get from ImageNet in 2012 to where we are today is open research. That's because it's kind of a worldwide collaboration, and everybody benefits from the progress of everyone. So, for me, it was important for the field itself to keep this going. And so we decided to create a nonprofit with the help of Eric Schmidt, Xavier Niel, and Rodolphe Saadé.
So, for the anecdote, the codename was Sphere because that's the name of the restaurant where we discussed the project. And we then understood that we could never trademark the name Sphere, obviously. So we just asked ChatGPT for "sphere" in a few languages. And sphere in Japanese is Kyutai. And there was AI in it. So we were like, okay, that's the name of the lab now. And so that's how we created Kyutai. The first person I reached out to was Alex Desfossés, who is also now co-founder of Gradium.
He's our chief science officer, because we had done our PhDs together at Facebook. And then when I joined Google, we kind of became rivals because we were working on the same stuff at the same time. And every time we'd meet one another, we'd just not talk about anything because, like, what are you working on? Oh, nothing. Okay. And, yeah, we had a small team, but with big expertise in speech. Again, it was an opportunistic decision that we made.
So I looked at the kind of stuff we could do around voice. And for our lab, okay, we had 1,000 GPUs. That seemed like a lot in 2023. It was still already at least one order of magnitude lower than bigger labs. So we wanted a project where we could make a difference despite the fact that we were four people. It shouldn't need too much compute, and it should be very innovative such that, just by being smart about what we did, we could make a difference.
And so we focused on conversational AI and real-time conversation because I had seen from the inside at Google that nobody was daring to touch this topic because it was so challenging, and it really seemed extremely far-fetched to be able to cast the task of dialogue into an LLM, right? People were just working on TTS and speech-to-text. The interactive stuff was really not there yet. So we thought, okay, we're going to work on it, and we're going to make it full duplex. Might as well do something really innovative.
What we didn't know at the time is that OpenAI had already been working on speech-to-speech conversation for a while. But in six months, with a core team of four people and then six, we were able to ship a model trained from scratch called Moshi that is still, to this day, the only full-duplex model. You can talk to it. It's a bit dumb because it's archaic in terms of intelligence, but the latency is still, one year and a half after, the best in the world by far.
And I think what was very interesting was that a very small team could make such a difference because then we shipped the first speech-to-speech translation system, and then streaming speech-to-text, and then streaming text-to-speech. And our models have been used across all industries. We are always proud to hear that a lot of big companies are using them. A lot of small companies are using them. The Maya and Miles demo from Sesame was built around our open-source models. I think what I really like about voice, and what makes Gradium a project that I deeply believe in, is that it's one of the modalities in AI where a very small team can make a difference.
I don't really see benefits from having extremely large organizations with a lot of people and resources because you don't need that many resources. You don't need 10,000 GPUs to train a speech model. You don't need 1,000 people. The ability to go fast, iterate fast with the right people, to me, is far superior to the advantage of having a big organization.
And then let's talk about Gradium. So Gradium is a commercial spinoff?
Yes. So, as I said, our open-source models were very successful. They are still downloaded millions of times a month. And we started having companies reach out, from all sizes, small and very large. They saw the potential. Sometimes it was a bit weird because I was talking to people leading extremely large teams, and I was explaining how I could train, by myself, streaming speech-to-text that was better than all alternatives. There was something we were doing very well, but at the same time, our open-source models remained limited.
They were fundamentally prototypes. For us, they were not even the actual contribution. For us, the contribution was the invention and the related research publication. For us, it was kind of like an artifact accompanying the paper, these models. But people wanted such models to be multilingual, higher quality, all the things that make an actual product. And so we considered kind of outsourcing this part, working with another team that would lead the product and so on. And honestly, after a few interactions, I realized nobody could carry such a project except us.
Nobody can believe in something if they have not developed it from scratch. We were in conversations with companies that wanted to create partnerships to improve our open-source models, and I would look at the specifications they wanted and realize it's something we could do in a few days, and they had been struggling for months. So, yeah, it seemed that we were also the best team to do that. And from a personal point of view, I was kind of addicted to academic prestige, having best paper awards, all this stuff.
It was always nice, but it was never enough. Every time I did a paper, I was happy for like two hours, and then I was already thinking about the next version. And at the same time, it eventually felt a bit weird that I could not imagine that, at the end of my career, I would have worked on something so applied and so close to actual real-life applications and not done them myself, not contributing to them directly. So I think in the end, even in terms of achievement, academic success is one thing, but nothing is near in terms of impact to the fact of having people using your models in the real world.
And in a way, I think it's something that is kind of generalized now in the industry, and it was also something that made me think. I looked at all the people I respected the most, the main one being, to me, a living legend, Aaron van den Oord from Google DeepMind. All these guys, they were making such nice contributions scientifically, and then they decided to focus on products. And I want to think it's because they also realized that academic prestige is one thing, but making a real impact is having your models being used in the real world.
And so, for me, it's the ultimate impact we can have. We still do science. In particular, Kyutai keeps doing open source and open science and so on. For me, the upside about it is mostly to be able to train the next generation of AI researchers and keep the field alive, as I said, because I think to have a healthy and vibrant AI field, you need to have scientific dissemination, so scientific exchange between institutions. And there also, the Chinese labs are doing remarkable work and are kind of forcing everyone to stay open to some extent, because otherwise it also hurts the ego, I think, of the people who are in the labs that don't publish.
So that was also something where I was very opportunistic about. So my strategy was like, if we publish in a world where the others don't publish, they will get so pissed off at us claiming all the inventions that it will make them join us eventually. And it's true that it's kind of hard because some people, they want to be in the place that is the cool place where the cool stuff happens. And it's not only about compensation and so on.
Now, I would say everybody working in AI and doing a good job in AI is going to get good economic outcomes. So then, glory is also very important, and it can be scientific or it can be just being proud that you are making the best products. But I think it's also an important part.