So lots of tools, lots of frameworks. So if I'm a person who wants to become a data engineer, what job do I need to have first? Am I a data analyst, or am I a software developer, or am I a businessperson? How do I become a data engineer?
Yeah, I see a lot of people move laterally into data engineering, often from either data analyst or software engineer. Data analyst, because I think there's just the sheer quantity of data analytics jobs. I think that's just the way it is. In terms of skill sets, it's usually things like SQL, Python, or some programming language. You should understand how a data pipeline works. You should understand how to put it together. Same thing with data modeling: you should understand how to model for OLAP or data warehousing models versus transactional systems.
So in terms of programming, do you need to be really good at programming? And what do you need to know? Do you need to know Python? Do you need to know other languages? What do you need to do?
Yeah, I'd say Python's definitely a good place to start because it is so prevalent. Obviously, there's probably some arguments to be made about how it abstracts a lot away. So you don't have to think about the difference between a list and an array as much as you would in other languages. But it's also just very prevalent. When it comes to data engineering teams, that's likely what you're going to be using. Obviously, in some places they still use Scala and other languages for one reason or another.
Python is probably the most prevalent. In terms of how proficient, I'd always recommend people don't feel like you have to cap yourself. I think sometimes people ask me how much programming you have to know. I'd know more than just for loops and the basics. I'd hope you can at least write code that operates in more complex systems. I think that's the biggest thing. Maybe you don't have to know data structures and algorithms to a T, but you should be able to, if I need you to write code against an API, be able to do that and iterate over it.
And it shouldn't take 30 minutes to process 100,000 rows of data. So I think there are aspects around that where it's like, yeah, you should be able to at least move data around via Python. And then obviously you're going to be spending some time in the terminal or command line as well. So know how to navigate.
And so you mentioned SQL. So how proficient do you need to be at SQL? Is that the core of the job?
I'd say that's where a lot of companies use SQL as their de facto transform tool. Obviously, there's some UI-based tools, but I think SQL has probably become the de facto tool in most cases. So I'd say the actual syntax of SQL doesn't take significantly long to learn. There's a few things that maybe I see every once in a while people struggle to learn, such as window functions, if you want to get specific. But that won't take you long.
What's going to take you longer, and this almost happens at every company, is getting a sense for how the data is shaped. Like, okay, when I interact with it via SQL, how is it shaped? When I join, where am I going to run into accidental places where I can't actually or shouldn't be joining things together? Or what do I need to filter out? Or how's the logic set up? So I often find that it's more about the data and less about SQL.
Because again, SQL you can learn quickly. What's going to take time is just getting a sense for how data operates and flows, and datasets in general.
And you mentioned data modeling. Do you want to define what that is and maybe go into some detail there?
Yeah, so there's a couple different ways you can essentially approach data modeling. Most people are probably accustomed to what they learned in their database course, which is transactional systems, which tend to be normalized to third normal form, depending on how far you want to take it. Obviously, now we can be a little more loose with that because there's so many ways you can handle data that's maybe a little less structured. But from there, that's generally not very good for analytics.
You run into various problems, and also you're generally trying to centralize maybe multiple different sources into one location, which is your data warehouse. And so we've kind of come around, and there's a few different ways you can model for what we call analytics. So people often reference that as OLAP, or online analytical processing. The way most people start is through this standard kind of data warehouse, which is facts and dimensions. And that's just creating data or putting data in such a format that, again, it's just easier for someone who's maybe less technical to come at and be like, okay, I can kind of see the fact tables in the middle.
If I need anything descriptive, that's in the dimension tables around that. And so that's what we're really trying to do at the end of the day, is get data from one format that maybe is a little harder, requires more joins, requires some understanding of business logic, because sometimes there's layers of logic in the application that maybe aren't shown in the data itself, and we're trying to capture some of that and put that also in the data model. That way, when an analyst comes to it, they can just look at it and be like, okay, I have a sense of what's going on.
I think that's the real goal, is we're trying to make it usable for analytics. Just as a quick point, the other thing is it usually, depending on the system you're running it on, is better for analytics in terms of performance for various reasons as well.
So what else do I need to know as a budding data engineer? Do I need to know Spark or Kafka or any of those frameworks? What would you recommend next? Once I have my language, I have my SQL, what do I do next?
Yeah, once you've kind of built that base of SQL, Python, whatever your programming language is, and understand how to work with the terminal and command line, I'd say then you can start getting comfortable. I'd pick a few tools that are popular and companies are hiring for. I think that's obviously what you want to do. So things like Spark and Kafka, I think, are a good place to start. Also getting familiar with interacting with the cloud, however you want to call it, AWS and GCP, I think are always a good place to start.
Obviously Azure exists. I have my own qualms with Azure every time I use it. So I just like AWS and GCP slightly better. But yeah, pick up some of those tools, get comfortable, and then start building some projects on that. I think you should have enough that you can start testing out projects. You're going to run into more tools as you're doing that. You're going to be like, oh, I want to maybe test out dbt or SQLMesh or something like that for transforms.
You can play around with that. But yeah, I think initially just picking up Kafka, Spark. Those are still so heavily in use that even if they go out in the next three to five years, it's going to take time to migrate and things of that nature. And then pick some clouds just to get comfortable with and interact with. Great.
And I think you mentioned Airflow earlier in the conversation, which is an orchestration framework. Is that something you need to learn as well early on, or is that more of a stage three kind of thing?
Yeah, I think that can definitely be more of a stage three, because Airflow is probably one of the most, if not the most popular, open-source solutions in terms of usage. We also know it best, so we tend to—I say we know it best; I'm sure there are other ones people know—but it's very well known, so we know the issues we dislike about it. So it's kind of this interesting space, right? Or it's in this interesting time right now where it's like, we have all these qualms with it, but also it's heavily used.
You're going to run into so many different data orchestration tools and data pipeline tools that you can pick some, and you'll run into a different one when you finally get hired. You're going to be using SSIS in some cases, and you're going to be like, oh my gosh, this is so old. But there are still companies using things like that, or Azure Data Factory, which isn't old, but these exist. And so I don't think you necessarily have to pick one.
But you can pick kind of one of the popular solutions for pipelines.
And beyond the technical stuff, what would you recommend data engineers learn to be successful at the job?
Yeah, I'd start getting comfortable understanding and talking to the business and understanding what their problems are and what you're actually doing. Because I think early on in your career, I think it's totally cool to just focus on tech and explore it, figure out what you like, figure out what you don't like, figure out how to improve performance of queries—do that. But you eventually need to start poking around and being like, okay, but why am I building this data model? What is it actually helping?
Who is it? Where's it going? What is this data actually being used for? I think asking those questions, the earlier the better. But if you're just focused on tech for the first year, that's fine. But once you can just start asking those questions and start trying to understand beyond just, like, okay, I'm building a data pipeline, but what is this for? Because I think that makes you a better strategic partner as you kind of grow.
And then that lets you open up new opportunities in terms of, like, you're not just this person that gets told what to do. You can start interacting with the business and being like, I think we should do XYZ, or these days, I think we should build this type of a data model, or we should start collecting this data over here, because that would help us on this side of the business and answer questions on churn or whatever it might be that your business is trying to answer.
How do people do that? Like, in your experience, do they just go around, make friends in the company, and ask what people do, what people want? Or what's a practical way of being able to do that?
Sure, I think that's never a bad idea. I think if you've ever talked to Ethan Aaron, he talked about how he would go around and do, like, a happy hour internally before he would do his low-key happy hours.
So alcohol is the solution?
No, I mean, it's not the solution unless you're, I guess, an O-chem person. But you can find a non-alcoholic version. That was just a story. Honestly, whenever you take on a project, I think that's part of it, right? Just get involved and care. I think if you can do that beforehand, just what we did at Facebook is we would, every six months, go around and talk to our partners. We'd talk about what we did maybe the last six months and try to understand what problems they're running into and where they need help, whether it involves data or not, because we would try to understand, even if they don't know that data or something we can do can help them, can they tell us something about what they're doing in their workflows where we could maybe interject?
Ourselves and be like, hey, could we do something here to help you guys? We think if we did XYZ, that'd be helpful. So having those outreach moments is always good as well.
And then the flip side: if I'm somebody that manages a data engineer, so I guess two scenarios. One, I'm a data team manager, so I'm a technical person managing other technical people. And then, as a second scenario, if I'm a business person and somehow I need to interact with or possibly manage a data team, what do I need to do for the collaboration to be successful?
Yeah, I mean, I'd say one thing is probably, hopefully, you can spend some time understanding some of the language, especially if you're working with more early-career data engineers, just because they're probably not going to bridge that gap initially. So you need to kind of help them figure out where maybe they're talking or using language that, if they're interacting with a stakeholder, could be just a little too deep—how to kind of handle those situations and better bridge the gap.
Because I think that's one of the spaces that, like, I noticed even for my own self when I was especially early on. It's like you start getting too deep in the weeds and kind of losing people. And so I think that's a big place that always can be helpful. So if you can call out someone when they're like, hey, you're going a little deep. Is there a way you could say this that they could understand instead of thinking about it from just your perspective?
Like, how would you talk to someone who's never heard of even what a data warehouse is? There are cases where maybe they don't know what a data warehouse is, or when you say data pipeline, maybe they don't understand. Or when you reference Spark, they're like, why do I care what Spark is? Right? And really just helping them understand how to bridge that gap.
You wrote about this quite a bit, where basically there is that constant risk for miscommunication or misunderstanding. And to play it back, if I got this right, I mean, what you're sort of suggesting is that the technical people should learn how to talk business a bit more, and the business people should learn how to talk technical a bit more. Is that fair?
Yeah, I think there's definitely that aspect. I think in general, I found that if you're the data team, you're going to probably have to do a little more of the lift, just because it's maybe a little easier for you to make that leap, or maybe you just have more time. I don't know why, but it just feels like generally, in my experience, that the data side tends to be the one that has to bridge that gap. You also have, I think, an interesting view of the data where it's a—not data, sorry, the business.
You have this derivative view that is the data itself. So you can see the data, and it represents the business to a degree. So you might see spikes before the business even sees it. So if you can understand it, that can help you proactively just come to the business and be like, hey, we saw this spike. And we didn't just tell you we saw the spike; we dug into it. We think it's from whatever, store A or product B or feature A.