I have this idea of how to extend it. One thing led to another, and I got an offer to join.
You just cold emailed them? You found their email address and just cold emailed them? That's awesome.
Yeah. Well, on papers, you have the email address under the author name. I could just reach out to them to say, "Hey, super cool experiments. Did you try this? What if we did this next?" And they replied to me and said, "Hey, why don't you come down to Mountain View and do some work with us?" So I got hired as an intern on Łukasz Kaiser's team, and I was sat next to Noam Shazeer. And that was sort of the start of it.
The end of it, if I skip all the way to the end, I was leaving Google after the internship had finished up, and they were throwing a little goodbye Aidan party, whatever, like some sweets and stuff. And my manager Łukasz is like, "Okay, everyone, Aidan's going back to his PhD. Aidan, how many more years have you got left?" And I had to be like, "Oh no, I need to finish third-year undergrad." And Łukasz was like, "What? We don't hire undergrad students."
I think I got in through an administrative mistake because my manager thought I was a PhD student. So that's sort of how I got there. The process of the Transformer, I mean, it was incredible. The velocity that came together was—I haven't seen anything like it since. So I showed up, the project that I was supposed to work on, it's a paper that is out and it was released at the same time, called One Model to Learn Them All. And it was like an omni-model.
So you can feed in text, audio, images, everything, and it could output the same. So super multimodal. Again, this was eight years ago, and so primitive compared to what we have today, but certainly prescient to where things were going. But as we were working on that, Łukasz and I built this framework called Tensor2Tensor. And it was used for doing big training jobs, distributing over a bunch of different GPUs, making things super efficient. Because we were sitting next to Noam, we convinced Noam to use it.
So Noam joined Tensor2Tensor. And then Noam was in conversation with folks over at Google Translate, which was like the second group of people who were working on these projects. And we realized folks were kind of working on the same thing. We were all looking at text-based autoregressive models that heavily leveraged attention or were much more pure attention, as opposed to these previous RNN models, LSTM models, which were quite complicated and kind of, in some ways, ugly. So we wanted to strip all that back and just create the most simple, efficient, attention-based language model.
And so then we just decided to team up and join forces. And that happened about a month into my internship.
So that was purely organic, like Noam happened to be around, and then you had some conversations with other folks. Was that how Google Brain operated? Meaning Google Brain allowed people to just organically form groups and teams?
Yeah, totally. That was it. It was a group of people who were researchers with full academic freedom to do whatever interested them. And you would sort of congeal around projects or ideas. And so that's what happened. We were just chatting to folks, saw a good idea, and then teamed up to take it on.
It seems that it's a defining characteristic of successful research organizations. At some point recently, we were chatting with Daryl Aquila, now of Contextual, and he was talking about FAIR at the time, and it sounded like it was a little bit like that as well. Do you think that was a moment in time when all those labs were authorized to, or allowed to operate that way and given free rein to explore anything? Is that still true today? Has it changed?
I don't know. I've been out of Google for long enough that I'm not sure how the culture has shifted. I would say that the economic relevance of this work is very different to when I was interning eight years ago. So I imagine things would have to shift out of necessity, especially because of the product implications, the amount of resources that are being thrown at these projects and these models. It's much more consolidated, I would imagine and expect. Certainly the way we run things at Cohere, it's much more like a product organization.
Like, you have very clear work streams. There's scope to experiment and try to find new alpha, but towards the ends of the product, right? So it's focused and more narrow. But back then, it was very greenfield. And so you could work on whatever you were excited about. And yeah, I do think that is like a crucial component of successful research organizations. And it worked. It produced incredible technology.
So going back to the story, the eight of you got together, and then how long was that process of writing the paper?
Super, super fast. So probably about a month in, we decided to consolidate and all work on the Transformer together. And it just became a mad dash towards the NeurIPS conference deadline. So NeurIPS is like the biggest AI conference for academics where you submit your papers to. And so we were just all-out sprinting. And it was a lot of, honestly, it was a lot of throwing shit at the wall and seeing what sticks. So many different things were tried, so many little bugs kind of hackily patched.
One example is pure attention architecture. The model can't tell the difference between positions of the elements in the same way that an LSTM could, because it could consume each one one by one. And so then Noam just came up with this idea. I remember the day I was sitting next to him, and he was talking to me about it, of just throwing these sinusoids into the embeddings and having that represent the position. And it stuck. And I think we've moved on a little bit, but shockingly, we're still quite close to that strategy today.