MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    Understanding Data Engineering in 2025 | Ben Rogojan, Seattle Data Guy

    Ben Rogojan is the Data engineering consultant at Seattle Data Guy. We cover why data engineering turns disparate sources into reliable datasets, how Iceberg lets companies separate data storage from compute engines, and why AI accelerates basic tasks but still needs engineers to prevent faulty models and exposed data.

    01/23/2025

    Hosted by Matt Turck · with Ben Rogojan, Data engineering consultant, Seattle Data Guy

    data engineeringApache IcebergAI engineeringdata infrastructureSnowflake
    Listen now
    YouTubeApple PodcastsSpotify
    1h 2m · 16 chapters
    Contents

    Transcript

    Why 2025 will be huge for data engineering

    1:20
    Matt Turck1:12

    Hey, Ben, welcome.

    Ben Rogojan1:14

    Hey there, Matt. Thanks so much for having me on.

    Matt Turck1:27

    So we are going to talk about all things data engineering in 2025, and it sort of feels like 2025 is going to be a huge year for data engineering, given everything that's happening in AI.

    Ben Rogojan1:48

    Yeah, it kind of feels like more of the same, in the sense that we repeat the same lessons. I think people kind of go through the same iterations. New people come to the space, or possibly maybe it's more the business side, get really excited by the prospect of what AI or data science or neural networks, whatever iteration we're in, can do. And then we kind of get down the line and we're like, oh, we need to get certain things in order first to kind of get to a point where we can actually use data well.

    Ben Rogojan2:10

    So yeah, I think it's going to be a big year. I think we're set up to try to actually figure out how these things work. We've had two years to figure out some of the basics and now see where things go.

    Matt Turck2:35

    Yeah, it sort of feels like that moment in the hype cycle where everybody got excited about AI, and then last year, especially in the enterprise, particularly big enterprises, there were POCs and a little bit of like, okay, what are the use cases? And if this year is the year of implementation, that's the year when the rubber really meets the road of, oh, how does that actually work? Where do we get that data and all the things? And so we are going to try to make this episode interesting to lots of different people interested in data and AI.

    The story of the Seattle Data Guy

    2:55
    Matt Turck3:05

    So both business people and much more technical folks. So hopefully we can cover some key concepts, some definitions, and all the things, as well as going into more technical stuff so that there's a little bit for everyone. But maybe before we get into this, a little bit of your background, your story. So you're Ben, but you're also known as Seattle Data Guy. So do you want to go into all of this?

    Ben Rogojan3:27

    Yeah, sure. No, great. Thank you for the light intro. So yeah, like you said, Seattle Data Guy. Now I live in Denver, so it currently snowed, I think, the other day. So it's white outside. That's new for me from Seattle. I've been kind of in the data engineering space for nearly a decade at this point. Started at a hospital doing a combination of data science, data analytics, data warehousing work. I think, like most people around 2012, 2015, I got bit by the data science bug, and I was like, I want to do data science.

    Ben Rogojan4:00

    It's the sexiest job of the 21st century. So I got into that, but eventually found data engineering and really enjoyed it. From there, I worked for a healthcare analytics startup that did everything from fraud detection, some stuff with opioid monitoring, and a few other things around population health. Again, around data engineering and building kind of a massive data warehouse, essentially, of pretty much, like, I think it was over half the U.S. in terms of who existed in that data warehouse. And then eventually did that at Facebook as well, data engineering.

    Ben Rogojan4:19

    All throughout that time, I've been consulting as well and putting out a ton of content. So I think I started putting out content somewhere around 2017. So I've been writing for a long time at this point, and then started putting out videos. So a lot of people might know me from different times. Every once in a while, I run into someone who's like, oh yeah, I saw your stuff on Medium, didn't know you had a YouTube, or I saw your stuff on Substack.

    Ben Rogojan4:42

    But yeah, so I've been really focusing at this point for the last few years on data infrastructure, helping companies figure out what and how to set up their data, how to find value from it. And really, that's kind of the space I've been playing in for the last little bit.

    Matt Turck5:07

    Yeah, very interesting. And you built a really big brand in the space, the whole Seattle Data Guy. I mean, you got over 100,000 subscribers on YouTube, over 100,000 on LinkedIn, all those networks. Maybe that could be interesting for people that are thinking about, in data engineering or software engineering, building brands and content and that kind of thing. Like, any lessons learned or how you think about it, like the pros and cons?

    Ben Rogojan5:26

    Yeah, I think the main lessons learned that I've picked up over the last few years is, like, don't expect the first thing to work. I originally started writing a blog before even Seattle Data Guy that was around food and culinary stuff, because before I worked in the data world, I had a stint in fine dining. And it did okay, but it never really took off. And you can see that as a loss, or it kind of helped me prepare for the next few things.

    Ben Rogojan5:54

    So, one, not everything you do has to immediately turn into something of value, something that produces dollars. It's just like learning programming or anything like that. Have side projects that are just there for the sake of learning, because you never know when it's gonna turn into something. From there, it's a lot about consistency. So just keeping up with it. And then just always looking for ways to improve yourself. Try to look around, see what other people are doing, see how to do what you're doing better.

    Ben Rogojan6:25

    And just look for little micro-improvements you can make. It doesn't have to be perfect day one. And that's kind of nice. I think a lot of people like the idea of having, like, 100,000 followers. It can be terrifying. Yesterday, when I put out my newsletter, I accidentally said that Databricks had bought Tabular. Or not Tabular. I said that Databricks had bought Iceberg. Iceberg was the next step. I meant to say Tabular. And I put that out there.

    Ben Rogojan6:38

    And then I was like, oh, I didn't see it until I was filming a video version. And I was like, oh, how did I not catch that? So those mistakes go out to now 100,000 people versus just 10, which is a much nicer situation.

    What exactly is data engineering?

    6:51
    Matt Turck6:54

    Yeah. And I got to say, as somebody who's been following your content for years, you're very consistent and you produce a lot of content, which I think is, as you said, key to a lot of these things. So let's start from the top. What is data engineering?

    Ben Rogojan7:07

    Yeah, I mean, my view, or at least a big portion of what data engineering is to me, is taking all this data and making it usable for people and now machines. I think that's something I liked out of Joe Reis's recent talk with data modeling. We're really just trying to take data from multiple disparate sources that maybe don't naturally talk to each other and put them into a format that other users, whether they're analysts or data scientists, can easily work with, and have that in a repeatable fashion.

    Data, AI, and ML: where do they overlap?

    7:41
    Ben Rogojan7:41

    We don't just want to have it work once. We have to have this system, these data pipelines, these tables that exist that people can rely on. And so that's really kind of the core of it. It's like we're just trying to make data more usable for people who maybe don't want to have to go through the core systems, and you can have those standards and things of that nature set up.

    Matt Turck7:46

    And so what's the overlap with AI engineering or ML engineering?

    Ben Rogojan8:10

    When we say AI engineering or ML engineering, I assume it probably means more taking those tools, like the tools that someone else probably built, right? Like someone else built OpenAI; we're probably now implementing it. Same thing with ML engineering. A lot of that, you probably had someone else build some of the models and you're just building it. There's a lot of overlap because oftentimes we will build things for maybe ML or AI teams. We do some of the collection aspects of it in terms of collecting the data, putting it in one place.

    Ben Rogojan8:34

    Depending on how your company is set up, you might just access the data directly from something like a data lake. And if you're an ML engineer or an AI engineer, you might just be doing it very raw. Some companies do have maybe the ML engineers access data more directly from maybe the data warehouse, which was something I saw at Facebook. In some cases, people would be accessing data closer to the source if they were in ML or AI, but there's plenty of ML and AI people who are relying on the data engineers to kind of build the core datasets for them to then go from there.

    Ben Rogojan9:07

    I think the other place you often see overlap—I wrote an article about this, or had someone else write an article and I shared it—they talked about one of the ways one of the teams at Riot Games was deploying their ML models, and they were using Airflow. And I thought that was interesting because, to me, Airflow is a very data engineering tool. But you see that you're still having ML engineers in some cases deploy using things that are data engineering-esque, in the sense that you're doing a batch model or something where you're training it maybe overnight or at a certain interval, and you're just using some tool, again, like Airflow to run that.

    Data analyst vs. data engineer vs. data scientist: what’s the difference?

    9:23
    Matt Turck9:49

    And while we're still on this topic of job definitions, you mentioned data analyst and data scientist. And again, that may be abundantly clear for anybody that's in the field, but it's always kind of surprising to me, once you leave that exact world, how people just cannot tell the difference between those different jobs. So what's a data analyst versus, again, a data engineer? And what's a data scientist versus a data engineer?

    Ben Rogojan10:08

    These roles still get very much amalgamated into one. I still see people, and I still talk to people, who are like, yeah, we're hiring a data scientist to build data pipelines and things of that nature. And so, in many ways, if you're at a small enough company, if you're in specific orgs in a company, you might just be doing everything in that one role. But in a perfect world, data scientists are likely doing things that involve a little more research, are a little heavier on the statistics side.

    Ben Rogojan10:46

    So they're trying to find and create models or insights using different tools that maybe analysts would use. Analysts tend to, I think, lean more heavily purely on SQL and maybe digging deep into those specific aspects using SQL, whereas data scientists might be using tools that are a little more statistically heavy. They're using models like something like k-nearest neighbors or something to try to find insights, versus just maybe an analyst generally is telling you and being very descriptive, like, hey, here's what happened. We can tell you why it happened and tell you maybe a little less in terms of what could happen in the future, how you can maybe build models to predict around that.

    A day in the life of a data engineer

    11:20
    Ben Rogojan11:20

    So you're more just describing what's happened and maybe have dug into specific cases of why it's happened, can kind of give a report on that, whereas data scientists might be on the other side building models off the top of that information that can kind of predict maybe in the future how to avoid that or how to better segment your marketing spend in the future to improve that. So I think that's some of the main differences you'll see.

    Matt Turck11:42

    So if you're a business person in a company or if you're somebody who's considering a career in data engineering, and both of those kinds of people are wondering what a data engineer does every day, what's a day in the life? What are some of the tasks that you tackle during a standard day as a data engineer?

    Ben Rogojan12:03

    I'd say you're going to spend some time trying to figure out exactly what the business needs, or if not the business, likely the analysts and data scientists, kind of talking to them, trying to understand what problems they're trying to solve. Personally, I think it's better to also get closer to the business, understand, okay, well, why are the data scientists and data analysts asking me to build this table? Once you've gained an understanding of the business, maybe some of the requirements they're actually asking, you're going to spend time planning and data modeling.

    Ben Rogojan12:29

    So, taking whatever that raw data is, whatever it looks like, trying to create it into a format that, again, analysts and data scientists in the future can parse and actually use in their models or in their reports. From there, obviously, you're going to build data pipelines that actually push to those tables and then actually kind of put it all together. And on top of that, you're probably going to be spending some time fixing broken data pipelines here and there.

    Ben Rogojan12:50

    Because things break all the time, no matter how much you try to kind of get around it. People make changes to tables. We've got some methods we're trying to, I think, implement, depending where you are, whether it's data contracts or something of that nature, to try to avoid some of those breaks. But to this day, people are still having, like, a data type issue break a pipeline that then causes everyone to have to kind of spend a little bit of time fixing them.

    Data engineering: Silicon Valley vs. everywhere else

    12:58
    Matt Turck13:19

    And so you worked at Facebook, and you also wrote in one of your posts about the reality of data engineering at companies outside of Silicon Valley. Can you maybe compare and contrast? If I'm a data engineer, do I do different things if I'm at Facebook or Meta or Google or Uber versus what I would do in a company outside of Silicon Valley?

    Ben Rogojan13:44

    Yeah, I don't think necessarily the goal changes. I think there are aspects of the fact that there are different tools. One, when you work at companies like Facebook and other big tech companies, their infrastructure is very mature. They've generally done, at least most of the ones that I've either seen or worked at—in the case of Facebook—they've done a good job of creating a centralized set of tooling that everyone's using and everyone's on, which makes your job very easy. When I worked at companies outside of Facebook, you have to spend a lot more time just setting up and configuring whatever tools you need to manage.

    Ben Rogojan14:13

    At Facebook, you essentially have your internal Airflow that a different team runs, and you can just push files to it, and it picks up the data pipelines automatically. Whereas in most other companies, you might be the person that has to even manage your Airflow instance, which is managing all your data pipelines. You go into a different company, they're going to have six different data storage or data platforms. One team's going to be using Databricks, one team's going to be using Snowflake, one team's still somehow using a SQL Server instance.

    Ben Rogojan14:39

    You've got all these mergers and acquisitions happening, which also cause further layers of, okay, we brought this other team in, how do we integrate their data into the rest of what we're doing? And so it can be a little more—I think it can feel a little more chaotic. And there's less support if you're a data engineer outside of Silicon Valley. And so you're having to solve problems that maybe a data engineer at a big tech company isn't having to solve.

    Ben Rogojan15:08

    Obviously, in big tech companies, the data is big, and so you have those problems to solve. But on the other side, you've got these more process and organizational issues when you're working in non-big-tech companies that, because again, they kind of change and move, and one VP wants to be in charge of some initiative. Every VP wants to have an AI initiative, I'm sure, right? Right now. So everyone's going to have their own. Yeah. So everyone's going to have their own tool that they've picked.

    Ben Rogojan15:26

    And now you've got ten different tools under one house, where I think, again, big tech is a little bit better. I'm sure they have their own issues in terms of initiatives, but they're a little bit better at picking, hey, we're all picking these tools, and we're just going to run with them.

    How to become an AI engineer

    15:27
    Matt Turck15:47

    So lots of tools, lots of frameworks. So if I'm a person who wants to become a data engineer, what job do I need to have first? Am I a data analyst, or am I a software developer, or am I a businessperson? How do I become a data engineer?

    Ben Rogojan16:14

    Yeah, I see a lot of people move laterally into data engineering, often from either data analyst or software engineer. Data analyst, because I think there's just the sheer quantity of data analytics jobs. I think that's just the way it is. In terms of skill sets, it's usually things like SQL, Python, or some programming language. You should understand how a data pipeline works. You should understand how to put it together. Same thing with data modeling: you should understand how to model for OLAP or data warehousing models versus transactional systems.

    Matt Turck16:35

    So in terms of programming, do you need to be really good at programming? And what do you need to know? Do you need to know Python? Do you need to know other languages? What do you need to do?

    Ben Rogojan16:57

    Yeah, I'd say Python's definitely a good place to start because it is so prevalent. Obviously, there's probably some arguments to be made about how it abstracts a lot away. So you don't have to think about the difference between a list and an array as much as you would in other languages. But it's also just very prevalent. When it comes to data engineering teams, that's likely what you're going to be using. Obviously, in some places they still use Scala and other languages for one reason or another.

    Ben Rogojan17:26

    Python is probably the most prevalent. In terms of how proficient, I'd always recommend people don't feel like you have to cap yourself. I think sometimes people ask me how much programming you have to know. I'd know more than just for loops and the basics. I'd hope you can at least write code that operates in more complex systems. I think that's the biggest thing. Maybe you don't have to know data structures and algorithms to a T, but you should be able to, if I need you to write code against an API, be able to do that and iterate over it.

    Ben Rogojan17:54

    And it shouldn't take 30 minutes to process 100,000 rows of data. So I think there are aspects around that where it's like, yeah, you should be able to at least move data around via Python. And then obviously you're going to be spending some time in the terminal or command line as well. So know how to navigate.

    Matt Turck18:01

    And so you mentioned SQL. So how proficient do you need to be at SQL? Is that the core of the job?

    Ben Rogojan18:26

    I'd say that's where a lot of companies use SQL as their de facto transform tool. Obviously, there's some UI-based tools, but I think SQL has probably become the de facto tool in most cases. So I'd say the actual syntax of SQL doesn't take significantly long to learn. There's a few things that maybe I see every once in a while people struggle to learn, such as window functions, if you want to get specific. But that won't take you long.

    Ben Rogojan18:49

    What's going to take you longer, and this almost happens at every company, is getting a sense for how the data is shaped. Like, okay, when I interact with it via SQL, how is it shaped? When I join, where am I going to run into accidental places where I can't actually or shouldn't be joining things together? Or what do I need to filter out? Or how's the logic set up? So I often find that it's more about the data and less about SQL.

    Ben Rogojan19:01

    Because again, SQL you can learn quickly. What's going to take time is just getting a sense for how data operates and flows, and datasets in general.

    Matt Turck19:07

    And you mentioned data modeling. Do you want to define what that is and maybe go into some detail there?

    Ben Rogojan19:31

    Yeah, so there's a couple different ways you can essentially approach data modeling. Most people are probably accustomed to what they learned in their database course, which is transactional systems, which tend to be normalized to third normal form, depending on how far you want to take it. Obviously, now we can be a little more loose with that because there's so many ways you can handle data that's maybe a little less structured. But from there, that's generally not very good for analytics.

    Ben Rogojan19:57

    You run into various problems, and also you're generally trying to centralize maybe multiple different sources into one location, which is your data warehouse. And so we've kind of come around, and there's a few different ways you can model for what we call analytics. So people often reference that as OLAP, or online analytical processing. The way most people start is through this standard kind of data warehouse, which is facts and dimensions. And that's just creating data or putting data in such a format that, again, it's just easier for someone who's maybe less technical to come at and be like, okay, I can kind of see the fact tables in the middle.

    Ben Rogojan20:32

    If I need anything descriptive, that's in the dimension tables around that. And so that's what we're really trying to do at the end of the day, is get data from one format that maybe is a little harder, requires more joins, requires some understanding of business logic, because sometimes there's layers of logic in the application that maybe aren't shown in the data itself, and we're trying to capture some of that and put that also in the data model. That way, when an analyst comes to it, they can just look at it and be like, okay, I have a sense of what's going on.

    Ben Rogojan20:51

    I think that's the real goal, is we're trying to make it usable for analytics. Just as a quick point, the other thing is it usually, depending on the system you're running it on, is better for analytics in terms of performance for various reasons as well.

    Matt Turck21:08

    So what else do I need to know as a budding data engineer? Do I need to know Spark or Kafka or any of those frameworks? What would you recommend next? Once I have my language, I have my SQL, what do I do next?

    Ben Rogojan21:30

    Yeah, once you've kind of built that base of SQL, Python, whatever your programming language is, and understand how to work with the terminal and command line, I'd say then you can start getting comfortable. I'd pick a few tools that are popular and companies are hiring for. I think that's obviously what you want to do. So things like Spark and Kafka, I think, are a good place to start. Also getting familiar with interacting with the cloud, however you want to call it, AWS and GCP, I think are always a good place to start.

    Ben Rogojan21:58

    Obviously Azure exists. I have my own qualms with Azure every time I use it. So I just like AWS and GCP slightly better. But yeah, pick up some of those tools, get comfortable, and then start building some projects on that. I think you should have enough that you can start testing out projects. You're going to run into more tools as you're doing that. You're going to be like, oh, I want to maybe test out dbt or SQLMesh or something like that for transforms.

    Ben Rogojan22:21

    You can play around with that. But yeah, I think initially just picking up Kafka, Spark. Those are still so heavily in use that even if they go out in the next three to five years, it's going to take time to migrate and things of that nature. And then pick some clouds just to get comfortable with and interact with. Great.

    Matt Turck22:32

    And I think you mentioned Airflow earlier in the conversation, which is an orchestration framework. Is that something you need to learn as well early on, or is that more of a stage three kind of thing?

    Ben Rogojan22:50

    Yeah, I think that can definitely be more of a stage three, because Airflow is probably one of the most, if not the most popular, open-source solutions in terms of usage. We also know it best, so we tend to—I say we know it best; I'm sure there are other ones people know—but it's very well known, so we know the issues we dislike about it. So it's kind of this interesting space, right? Or it's in this interesting time right now where it's like, we have all these qualms with it, but also it's heavily used.

    Ben Rogojan23:16

    You're going to run into so many different data orchestration tools and data pipeline tools that you can pick some, and you'll run into a different one when you finally get hired. You're going to be using SSIS in some cases, and you're going to be like, oh my gosh, this is so old. But there are still companies using things like that, or Azure Data Factory, which isn't old, but these exist. And so I don't think you necessarily have to pick one.

    Ben Rogojan23:22

    But you can pick kind of one of the popular solutions for pipelines.

    Matt Turck23:29

    And beyond the technical stuff, what would you recommend data engineers learn to be successful at the job?

    Ben Rogojan23:52

    Yeah, I'd start getting comfortable understanding and talking to the business and understanding what their problems are and what you're actually doing. Because I think early on in your career, I think it's totally cool to just focus on tech and explore it, figure out what you like, figure out what you don't like, figure out how to improve performance of queries—do that. But you eventually need to start poking around and being like, okay, but why am I building this data model? What is it actually helping?

    Ben Rogojan24:11

    Who is it? Where's it going? What is this data actually being used for? I think asking those questions, the earlier the better. But if you're just focused on tech for the first year, that's fine. But once you can just start asking those questions and start trying to understand beyond just, like, okay, I'm building a data pipeline, but what is this for? Because I think that makes you a better strategic partner as you kind of grow.

    Ben Rogojan24:15

    And then that lets you open up new opportunities in terms of, like, you're not just this person that gets told what to do. You can start interacting with the business and being like, I think we should do XYZ, or these days, I think we should build this type of a data model, or we should start collecting this data over here, because that would help us on this side of the business and answer questions on churn or whatever it might be that your business is trying to answer.

    Matt Turck24:48

    How do people do that? Like, in your experience, do they just go around, make friends in the company, and ask what people do, what people want? Or what's a practical way of being able to do that?

    Ben Rogojan24:58

    Sure, I think that's never a bad idea. I think if you've ever talked to Ethan Aaron, he talked about how he would go around and do, like, a happy hour internally before he would do his low-key happy hours.

    Matt Turck24:59

    So alcohol is the solution?

    Ben Rogojan25:22

    No, I mean, it's not the solution unless you're, I guess, an O-chem person. But you can find a non-alcoholic version. That was just a story. Honestly, whenever you take on a project, I think that's part of it, right? Just get involved and care. I think if you can do that beforehand, just what we did at Facebook is we would, every six months, go around and talk to our partners. We'd talk about what we did maybe the last six months and try to understand what problems they're running into and where they need help, whether it involves data or not, because we would try to understand, even if they don't know that data or something we can do can help them, can they tell us something about what they're doing in their workflows where we could maybe interject?

    Ben Rogojan25:52

    Ourselves and be like, hey, could we do something here to help you guys? We think if we did XYZ, that'd be helpful. So having those outreach moments is always good as well.

    Matt Turck26:19

    And then the flip side: if I'm somebody that manages a data engineer, so I guess two scenarios. One, I'm a data team manager, so I'm a technical person managing other technical people. And then, as a second scenario, if I'm a business person and somehow I need to interact with or possibly manage a data team, what do I need to do for the collaboration to be successful?

    Ben Rogojan26:32

    Yeah, I mean, I'd say one thing is probably, hopefully, you can spend some time understanding some of the language, especially if you're working with more early-career data engineers, just because they're probably not going to bridge that gap initially. So you need to kind of help them figure out where maybe they're talking or using language that, if they're interacting with a stakeholder, could be just a little too deep—how to kind of handle those situations and better bridge the gap.

    Ben Rogojan27:02

    Because I think that's one of the spaces that, like, I noticed even for my own self when I was especially early on. It's like you start getting too deep in the weeds and kind of losing people. And so I think that's a big place that always can be helpful. So if you can call out someone when they're like, hey, you're going a little deep. Is there a way you could say this that they could understand instead of thinking about it from just your perspective?

    Ben Rogojan27:25

    Like, how would you talk to someone who's never heard of even what a data warehouse is? There are cases where maybe they don't know what a data warehouse is, or when you say data pipeline, maybe they don't understand. Or when you reference Spark, they're like, why do I care what Spark is? Right? And really just helping them understand how to bridge that gap.

    Matt Turck27:49

    You wrote about this quite a bit, where basically there is that constant risk for miscommunication or misunderstanding. And to play it back, if I got this right, I mean, what you're sort of suggesting is that the technical people should learn how to talk business a bit more, and the business people should learn how to talk technical a bit more. Is that fair?

    Ben Rogojan28:10

    Yeah, I think there's definitely that aspect. I think in general, I found that if you're the data team, you're going to probably have to do a little more of the lift, just because it's maybe a little easier for you to make that leap, or maybe you just have more time. I don't know why, but it just feels like generally, in my experience, that the data side tends to be the one that has to bridge that gap. You also have, I think, an interesting view of the data where it's a—not data, sorry, the business.

    Ben Rogojan28:34

    You have this derivative view that is the data itself. So you can see the data, and it represents the business to a degree. So you might see spikes before the business even sees it. So if you can understand it, that can help you proactively just come to the business and be like, hey, we saw this spike. And we didn't just tell you we saw the spike; we dug into it. We think it's from whatever, store A or product B or feature A.

    Will AI replace AI engineers?

    28:46
    Ben Rogojan28:46

    We think that's causing this problem. Instead of waiting for them to tell you, hey, we're seeing a decline in churn, you can already have proactively looked into it.

    Matt Turck29:24

    So, speaking of the sort of reality of what data engineers do on a daily basis, obviously one big topic that's very current in lots of people's minds is the potential automation of engineering and coding in general by AI. And there's been a lot of discussion around typical software engineering, certainly, like a million startups doing this. There's a handful of startups also trying to do automation for some data engineering stuff. What's your take on this? What's the potential? What's the risk?

    Matt Turck29:29

    Is AI friend or foe? Is it overhyped, underhyped overall?

    Ben Rogojan29:53

    I definitely have tried to integrate, and have successfully integrated to some degree, AI into my various workflows. It's not to say that I've got OpenAI in everything, but it is to say that, yeah, it makes things easier. There are tasks I used to have to do. I was trying to do something as simple as create a CREATE statement where I wanted to rename the columns and already go through and unabbreviate all the column names. And I'm like, I could go through and try to figure out all the abbreviations and what they mean, but I'm like, I'm pretty sure if I put this into any form of LLM, it'll figure it out for me, and it does it considerably faster than I would.

    Ben Rogojan30:29

    So there are aspects where it will accelerate. I think there are tasks that are especially very basic that it's been helpful for, and I think it will be helpful for. I think in terms of risks, I don't think the goal is always to be faster. Code speed isn't always the goal. I think being clearer and being more understanding of what we're trying to design is usually the goal. It's like, what are we actually trying to build? Why are we building it?

    Ben Rogojan30:51

    So if we can get better at that, then I think a lot of these tools maybe pose some risk. But I think because there's all this technical knowledge I think you need to have to even be able to interact with some of these systems, in the sense that you need to know what to ask for, I think that still means a technical person needs to be involved with a lot of this work. So a data engineer is still going to need to be involved.

    Ben Rogojan31:16

    Otherwise, you risk people building the wrong model or maybe exposing data that they shouldn't be exposing. All this, maybe it's baked into the LLM, maybe it's not. And I think there's just a lot of risk if you don't have someone look at it. So maybe you need less data engineers in the future, but I still see that someone needs to understand what's going on. A good example of that is I did a video with Josué. I don't know if you've seen his content, but he does a lot of Databricks content.

    Ben Rogojan31:45

    And we were kind of going through and testing out Databricks Genie. And of course, maybe it could have been better configured. There's a lot of things about how it's set up that could play a role. But it ran into an issue writing a query that was not that long, I'd say, because there's like 5,000-line queries out there that exist. So I think this was maybe in the few hundreds, maybe 200 lines, and it had an issue in it somewhere. And had Josué not been technical, he wouldn't have been able to understand why that issue existed and how to correct it.

    Ben Rogojan32:12

    And he was able to correct it. And obviously then, from there, Databricks was able to correct it. So it takes time to do that. So I think we're a few years away from just replacing data engineers, which has been the goal for what feels like a decade or more now, since SSIS came out. They're like, okay, how do we get rid of these people that keep telling us to slow down with data?

    Matt Turck32:36

    And that's the coding aspect of it mostly. But do you think, I guess everything is code ultimately, but do you think there are some chunks of the chain that could be automated? Like, I don't know, the whole world of ETL, ELT. Why do you need to have a bunch of companies writing connectors to other sources, moving data around? Do you see a future where that gets automated?

    Ben Rogojan33:03

    I can, but it's weird because that's also something that I've thought about for a little bit, where it's like we've spent, it feels like, billions and billions of dollars every year just trying to extract data and put it into things like Snowflake or some other data warehouse. And it still feels like we're not great at it. Sure, good, but you'd assume we'd be better in some cases. It'd either be faster or cheaper. And in some cases, it's still not. I still run into people who, for certain ELT tools, they're like, it is way too expensive.

    Ben Rogojan33:28

    It's costing me way more than Snowflake. So I do still run into that where I'm like, wow, that's wild. You'd think that by this point it wouldn't be, but I still have people who are like, yeah, I'm only pulling from one data source, but because it's so many rows or something, they're charging us $20,000 a month for it. But in theory, you could hope that in the future you could write an LLM that, what do you call it, could write the API connectors better.

    Why is the data world so complex?

    33:42
    Ben Rogojan33:42

    Because that's usually the big thing: the API connection tends to be the hardest.

    Matt Turck34:12

    Actually, taking a step back, through this whole discussion, there's been a lot of conversation about complexity, lots of tools, and all the things. And as you know, I do every year this big market map called the MAD Landscape, where there's every year even more gajillion new tools. Why is it so complicated? From your perspective, why do you need to put together those chains of tools? And sometimes it works, sometimes it doesn't.

    Matt Turck34:21

    You mentioned broken pipelines. That seems to happen all the time. Why are we here?

    Ben Rogojan34:44

    That's something I've definitely been thinking about more recently, because when I first started in the data world, the common approach for a lot of companies I worked with would be, okay, we'll set up some cron jobs to run some SQL scripts, essentially, or some Python scripts that call SQL scripts, or PowerShell scripts that call SQL scripts, and we'll just time it all correctly so that they run in a specific order or set up dependencies somehow. And then what a lot of people would end up doing is just essentially building Airflow internally.

    Ben Rogojan35:01

    And so we knew the problem. We'd run into this problem where it's like these things don't talk to each other, and that causes issues. So then when Airflow, I think, came out, I think that's why it was popular for a lot of people. It was like, okay, cool, now we can write scripts and not have them bump into each other, or have them make it very easy to backfill and all these things that we find painful as data engineers.

    Ben Rogojan35:30

    And then somehow with the MDS, the modern data stack kind of push in 2020, '21, '22, we kind of ended up here again, where it's like, okay, we have one tool that does extract. Okay, we have one tool that, from extract, does transform. Okay, from there, you kind of can keep parsing it down. We have another tool that does data quality. Part of this works because I think it grew the pie, where now we have more companies that can actually access these tools, because before it's like, okay, you want Informatica, let's talk six, seven figures in contract costs.

    Ben Rogojan36:02

    Now it's like, okay, instead, let's talk a grand or two a month. Then that seems more reasonable. So I think you've expanded that. But at the same time, yeah, we kind of still built this very disparate set of tools. So I think some of it is our own—we're trying to constantly iterate and make things better. But by doing that, we take a few steps back occasionally. And so now I think we're relooking at things, and people are like, okay, how do we get out of what we've built so that we build some simpler systems?

    Ben Rogojan36:30

    I mean, I have clients who literally, if I bring up more than two tools, will look at me and be like, can we do this with one? Is there anything we can do to simplify this? So I do think people are feeling it and they want less. But I think it's just, again, kind of this natural bundling, unbundling that happens as people try to make a better version of just one part of a tool. And then you'll kind of hopefully bring it back at some point.

    Ben Rogojan36:40

    Oh, also, just to add another point, I think the reason you see so many tools is, I mean, VC funding was pretty high. So you can blame yourselves.

    Matt Turck36:43

    I knew VCs were behind this.

    The functional consolidation of the data world

    36:53
    Ben Rogojan36:54

    They're not alone, obviously. It's like VCs, I think. We data engineers and technical people, we love new tools, we love shiny objects. And so we'll test anything out.

    Matt Turck37:24

    A little bit on that topic: do you view a potential future where the consolidation we've been talking about for many years happens? And I'm more thinking, so there's certainly the sort of corporate consolidation, M&A, and all of the things, but I'm more thinking about functional consolidation. I guess the modern data stack was an attempt at sort of consolidating, having a sort of default set of technologies that work together. So do you view more of that? Do you view the hyperscalers, whatever, Microsoft Fabric, being able to provide, or Databricks trying to provide, all the things for all people?

    Matt Turck37:36

    Is that a likely future for you?

    Ben Rogojan37:56

    I think if they execute well, it's possible, right? I think that's been the big thing. Like, when I use Snowflake, for example, they have some integrations with things like HubSpot and other tooling. And it's nice. I've used it for a couple of clients. It makes it considerably simpler. But then there's always just little things like, okay, Snowflake has tasks, but how I set it up, they're kind of getting a UI now. So they're kind of getting to the point where it's like, okay, now you're getting into this dbt-, Airflow-esque space where maybe, yeah, you could be everything.

    Ben Rogojan38:18

    And Databricks has already kind of been there, where they have their workflows that you can automate, whether it's notebooks or something else. So they've had that. So I think it's going to come down to execution, where it's like, if they can do it well, sure, people will go for it. But if they don't do it well enough to where it's like, look, I prefer to use Airflow, of course we're going to stick with Airflow, or whatever tool you're using for your orchestration.

    Ben Rogojan38:33

    And if you don't have a good extract tool, you're going to use whatever they can find out of the box if they don't have enough engineers.

    Big data stories from 2024

    38:34
    Matt Turck38:50

    All right, let's talk about where we are sort of today, like in 2024, taking stock of what's been happening in the last 12 months, and then we'll go into 2025. But in 2024, what were some of the big stories for you?

    Ben Rogojan39:01

    Yeah, I mean, I think we had a few large acquisitions, right? Rockset got purchased out by OpenAI, which I thought was interesting. For, I don't know, was it like nine figures, eight figures? I don't remember where it was.

    Matt Turck39:05

    Yeah, I think that was something like $500 million, right? Or something like that.

    Why Iceberg is a game-changer

    39:28
    Ben Rogojan39:28

    Yeah. So, large acquisition there. Same thing with Tabular. So there's some interesting purchases, I think, right? We've kind of now set, in a weird way, Iceberg as the de facto—if we're going to use an open table format, that's likely the one people will, I assume, pick, since everyone's now supporting it, right? AWS is trying to do their own version.

    Matt Turck39:45

    And maybe, still in an effort to make this interesting to a broad group of people, do you want to talk about Iceberg—what it actually is, what it does, why it's a big deal—and then the Tabular acquisition by Databricks? So Tabular being the main company behind Iceberg.

    Ben Rogojan40:06

    Sure. I think the big thing is that it sets a standard for how you're going to end up storing data and interacting with it, which, instead of Databricks stores their data in Delta, Snowflake stores it kind of in S3, but also in their own formats, now you're storing it in a singular format, which then allows all of these other solutions to essentially interact with it. Like, okay, well, now Snowflake can just have their method of handling Iceberg, and you don't have to think about it.

    Ben Rogojan40:36

    And same thing with Databricks. You can just have one data store and then pick which data compute engine you want to use, which is actually something that Facebook, when I was leaving, was already doing. They had kind of an HDFS sort of underlayer of where they stored their data. And then from there, you could just say, like, I'm going to use Spark, I'm going to use Trino—well, Presto internally, but Trino for people externally. So I'll use Presto. We still occasionally had Hive jobs.

    Ben Rogojan41:00

    So if someone wanted to have Hive, it lets you pick those. So you're not locked into a specific vendor so much. You are able to kind of pick which compute you want, which makes a lot of sense when I see things or when I talk to a lot of people where it's like, hey, we use Databricks for maybe the more expensive compute, and then we push it to something like Snowflake for our data analytics layer for BI work. So now, instead of having to do that back and forth, you just have one place you're storing the data, and you just have a catalog that needs to sit on top of it and tell everything else what's going on.

    Matt Turck41:34

    Which is a big deal in the industry because obviously the big strategy for Snowflake, but also Databricks, over the last few years has been to say, hey, you have one repository where you just dump all your data, and that happens to be us, Snowflake, or that happens to be us, Databricks. And then you're sort of stuck. So the industry has reacted in a way that basically says, no, you can have your data anywhere, and then you pick your compute.

    Ben Rogojan41:57

    I think that's fair. And that's been kind of the recent fight between, I'd say, Databricks and Snowflake: trying to convince you that their compute is the compute to use for use cases or different workflows. It's really been a fight for workflows versus a fight for contracts, because in the past it was like, okay, you sign a $10 million contract for Teradata, you're kind of stuck with it, right? There's no choice now for the next three to five years. Now it's like, well, you could use Snowflake or you could use Databricks.

    Ben Rogojan42:05

    It's just whichever one convinces you it's better for whatever workflow you're doing.

    Matt Turck42:38

    I saw you write in one of your posts, which I thought was super interesting, especially in contrast to this whole titanic fight between Databricks and Snowflake, is that a bunch of the projects that you worked on had to do with SQL Server and on-prem migration. So while those two juggernauts fight it out, it sort of feels like there's a chunk of the reality where people are like, oh, well, yeah, this whole cloud thing could be interesting for us. Is that fair? Is that what you see sort of in the trenches?

    Ben Rogojan43:06

    Yeah, I mean, there's still people, again, I think I made a joke earlier, that have a SQL Server instance running underneath the desktop, essentially, and are just now thinking about migrating. I think I had three or four migration projects from SQL Server specifically to Snowflake. I think it's a pretty natural transition when they're going to the cloud. So yeah, there's all sorts of people and all sorts of tools. I think I also posted once, the more you think about Snowflake versus Databricks, the happier they are.

    Ben Rogojan43:22

    It limits your choices, but there's so many, right? You can use BigQuery. Not a big Redshift fan. I've not come back after getting burned in, like, 2016, 2017.

    Matt Turck43:24

    Why did you get burned?

    Ben Rogojan43:45

    Oh, I think most people just, you go to the cloud, you expect certain benefits, and then there are little things. I think they fixed some of these, but, like, they didn't have a MERGE statement. You're still kind of stuck on a certain amount of storage. And so you just run into little things where you had to do VACUUM to clean up data that would get fragmented, or if you deleted it, it would be only partially deleted. So there's lots of little things where I think, because of who they, if I recall, they purchased a solution that was based on Postgres and then just moved it to the cloud.

    Ben Rogojan43:59

    Because of that, they were just so attached to the initial technology that they couldn't really change that much.

    Matt Turck44:26

    And I think you also wrote about one of the things that you saw in 2024 was fractional data teams and consolidation in tooling, which I guess seems to imply that there's been general budget restrictions, at least from the clients that you work with. And maybe there's some bias here because you're an external consultant, so maybe they go to you when they are in that situation. But what do you see and hear about the reality, again, in the trenches, of sort of like budgets and effort on data engineering so far in the last few months?

    Ben Rogojan44:58

    Yeah, no, I mean, and I appreciate that you call it the bias. Sometimes I think about that myself when I'm like, is everyone's data a mess? Or is it just the fact that you call a consultant when that's why you call the consultant? I do think there's that aspect for some teams, that spending $200,000, or whatever it's going to be, to pay for even a single data engineer is a big ask, right? Like, if your company's in the several million, $100 million, like, it's not nothing.

    Ben Rogojan45:26

    I like to think that the reason you have a data team is often because you've already run a decent business, right? Like, if you have, let's say, $500,000 to invest in data, which is what it might cost you depending on the size of your company, that means you've already run a business that has that much surplus, right, on top of your actual net profit. So you're running a good business as is without the data team. And so you're hoping that if you bring in a data team, you're seeing that 10x return on that.

    Ben Rogojan45:53

    And if companies aren't feeling like that, they're going to try to look for other ways to get it. Like, they still want to answer questions. All companies want to answer questions. I mean, if it's a $10 million or $5 million business, they're going to figure out a way to do it. So I think that's one reason you're seeing some fractional teams become popular, or at least more popular in my view. So it is kind of a little bit up to us to actually make it worth it.

    How startups manage data in their early days

    46:02
    Ben Rogojan46:02

    We're also seeing that the ZIRP era is over, and people are just hiring less in general.

    Matt Turck46:24

    So that's what you would recommend? Like, if I'm a Series A or seed startup and data is really important to me, but that's not everything I do, you'd start with consultants, fractional data teams. And I guess, what is the first thing you'd do? Like, you'd get Snowflake or something like that?

    Ben Rogojan46:26

    I'll try to avoid the "it depends" answer.

    Matt Turck46:34

    Well, you can go into "it depends." It's okay to be nuanced for something like this.

    Ben Rogojan46:53

    I do think if you're early, starting out, there's no problem in probably bringing on some sort of consultant to do maybe more of the data engineering work, and then bringing on a full-time, maybe, analyst to kind of work on top of that is what I'd imagine would be good. Have someone that can set up the infrastructure and then hopefully get it running, and then maybe have them come back occasionally. And then I do think it's valuable to have an analyst who understands the business, who gets integrated, who's actually involved, and not just kind of this someone that's a mercenary, so to speak, that's going to be gone and maybe wants to help you, but at the same time, it's just not as integrated in the business long term.

    Ben Rogojan47:30

    So I think that's, if you want to talk about a good baseline, have someone set it up. Does it have to be Snowflake? That just depends, to me, on the size of your data. I think what usually I see happen is, initially, people just report off whatever their transactional system is. So if you're using Postgres or something and you've got some application layer, you can report off that for a good amount of time to get your baseline metrics. But at a certain point, once you either need to get more funding or maybe once you get more funding, people want to see more metrics, right?

    Ben Rogojan47:57

    Like, VCs want to see more metrics. Your team probably needs to see more metrics for you to do better. Or maybe you just have a larger amount of data coming in because you're doing more transactions, you're having more users on your system. Then it starts making sense to look into, okay, do we need a Snowflake? Do we need Databricks? Do we need BigQuery to kind of handle some of this and handle some of your processing? And you'll probably start running into various issues.

    Ben Rogojan48:17

    And the issues I usually see is people run into, we only have a software engineer and they're kind of having to do this querying stuff. We don't want them to do that. We want them to be building features. So we're going to bring in someone to do that instead, maybe push it somewhere else. The reports are going too slow, or however we're doing it, it's too slow. So we need something that's more fine-tuned towards analytics, or we're trying to integrate more datasets, right?

    Ben Rogojan48:43

    Like, we want to bring in five or six different datasets. And right now we have an analyst kind of piecemealing it all together. So those are usually the causes that I see people say, like, okay, we should probably consider building more infrastructure, which might mean we need to bring a consultant or a full-time person based on kind of how much budget we have to build that, and then the analyst can work off of a more centralized kind of reporting layer.

    Seattle Data Guy’s favorite tools

    48:44
    Matt Turck48:55

    You alluded to some of this, but what are some of your favorite tools? So, not an Azure fan, not a Redshift fan, but conversely, what do you really like?

    Ben Rogojan49:18

    I do think I tend to have things I dislike more than I like. There's a lot of tools. Snowflake makes it really easy, and I'd say BigQuery does too, but Snowflake makes it really easy. If you do not have a ton of people to manage things, I think it's hard to go against it. Similar with BigQuery. I think Databricks makes sense to me still more from an ML standpoint. I know they like the data warehousing side. I know that's where they're going to pitch themselves.

    Ben Rogojan49:46

    But I just tend to see, again, like I referenced earlier, the migration from SQL Server to Snowflake has worked really well, from what I've seen, for a lot of clients. Where the way I translate "it goes really well" is, like, when we migrated, they were only doing a certain amount of reporting. And by the time we ended migrating, they're already bringing in more data, they're already asking more questions because it's just easier for them to interact with. So I use that as kind of my baseline, where I'm like, okay, I'm seeing people just do more.

    Bold predictions for 2025

    50:09
    Ben Rogojan50:09

    And that, to me, means that it works out well for them. So I always have a soft spot for Airflow just because that's kind of the first tool I ran into that made my life easier. Because, again, I came from this world of cron-scheduled jobs, and it just made things easier. So I think that always has a soft spot.

    Matt Turck50:25

    So, 2025, let's talk about what we have in store for this year and a few predictions. So let's go with your top five predictions.

    Ben Rogojan50:39

    I think some of this I might have referenced, but I think, again, Iceberg's kind of cemented its stronghold, so to speak, right? In terms of where it will be used, where people are going to look for an open table format, that probably will be the one that they'll pick. Like I referenced earlier, I think at large enterprises, it gets really hard to pick a single tool or to pick a single initiative that's going to be run by one VP that everyone's going to agree to.

    Ben Rogojan51:11

    So I think in those places, you're going to see people pick other tools or maybe somehow have five different instances of Iceberg, or maybe one's using Databricks, one's using Snowflake, one's using AWS, and maybe they'll all talk to each other, maybe they won't. And I think the other side is a lot of VPs, directors, and even just data engineers don't all want to spend their weekends tinkering with all these tools. Some people love it. Some people love trying out all the new tools.

    Ben Rogojan51:35

    The example I gave in my recent article was I talked to a VP that I remember saying—he told me, he was like, "I like skiing on the weekends. I don't want to deal with something that breaks because we built some custom tool. So I want something that works. I don't really care if it costs an extra $5,000. Great, I'm going to pay for it." So I think there's that aspect of it.

    Matt Turck51:36

    Prediction number two.

    Ben Rogojan51:56

    Prediction number two, perfect. The next one I said: SQL isn't going anywhere. And I kind of said that for more than one reason. One, I feel like someone's going to tell us it's time to revamp Data Lake V1 again, where we'll just dump everything in there and we'll let an LLM figure it out, which I don't think will work. So I think that's why we'll still see SQL kind of be around. Maybe someone will convince us, and there'll be enough new people in the space and enough business people in the space that will be convinced that it's like, hey, again, we can get rid of data engineers in a possible way, or some layer of data modeling, and we can get faster access to data.

    Ben Rogojan52:29

    Yeah, let's do that, right? Saves us several million dollars. But I think we'll kind of end up in a similar space. And that's why even I made a bet in that space, more from my own angel investing side, where I'm like, yeah, I think some tools make sense, but others don't.

    Matt Turck52:59

    In the same vein as some of the discussion earlier about the AI automation, there is certainly a group of company features that offer the promise of enabling business users to effectively run SQL, but enter a prompt in natural language, and there's something that translates it into SQL and then fetches the results. Is that something that you've seen? Is that something that you think is interesting or not?

    Ben Rogojan53:24

    I mean, I obviously think it's interesting. I've seen the natural language-to-query kind of startups exist for a long time, and it's just been very hard. I don't think we're going to see that take over anytime soon unless you have your data in such a well-formatted way that, like, yeah, of course, it's easy to do. Like, you're very clear on maybe the terminology or something of that nature. So there's no confusion on what you're asking for, maybe. But even then, I don't know if I'm going to start seeing several-hundred-line queries be perfect from an LLM.

    Ben Rogojan53:53

    Maybe if it's, like, a very specific question, sure. But once you start digging deeper, I think that's where the problem comes. And that's usually the kind of barrier between self-service even now, is like, you'll answer your first question, and then there'll be like 10 more questions. And that's when suddenly they're like, okay, can I just put this into Excel because I'll answer it myself?

    Matt Turck53:55

    All right, prediction number three.

    Ben Rogojan54:08

    Yeah, I think we'll start seeing some AI go from press releases and press release-driven initiatives to real life, right? There are probably companies, and I know of some that I've talked to, where they've been building things. For the last few years, they've been quiet, they've been building things, they've been trying to build a layer of data infrastructure that will support doing things like deploying whatever they're doing in AI very easily versus having to cobble together something and make it very hard to even deploy something as simple as an ML model.

    Ben Rogojan54:49

    I think you're going to start seeing that become realistic. We talked about this before. People have done their POCs. People are trying to figure out where they work. If they've built a good foundation and they have some POCs that work, you'll start seeing that be realistic. I think we'll see, hopefully, things more than just customer service be the use case. Because that always feels like the first place people jump, is customer service. It's why we all hate calling anything at this point.

    Ben Rogojan55:05

    It's like, okay, I got to call the bank. I got to go through this, what they call a chatbot at some point, or some people call them phone trees. But okay, now I just get to talk to OpenAI instead, but it's kind of the same thing.

    Matt Turck55:07

    Prediction number four.

    Ben Rogojan55:30

    I think we'll continue chasing the same holy grails. I feel we want to see new things, but I think we're still trying to figure out things like self-service analytics or some of the other holy grails. Even being data-driven kind of feels like this holy grail that companies try to reach out for and, in some cases, get there, but in many cases just spend a lot of money and they're never really satisfied with the end. Some of that's just due to lack of data quality, lack of data governance, things of that nature.

    Ben Rogojan55:56

    Maybe they never defined what they really wanted in the end. They heard self-service and the salesperson painted an idea, but no one defined, okay, this is what we mean by self-service internally, and that's how we know a project is successful or not. So I think we'll just continue chasing some of those things because I still see companies struggling even there.

    Matt Turck56:25

    And why is that? We were talking a bit earlier about why there are so many tools and brittle data pipelines and all the things. Same question here. We've been doing this data thing for decades now. Fifteen, 20 years ago, it was like the whole big data thing. I mean, it sort of feels like there should be a standard playbook by now. I mean, as an industry, we've accumulated enough institutional knowledge that we should know how to do this. We should know, okay, well, this is what self-service analytics looks like, and maybe you can do it in the specific circumstances of your company, but at least the goal should be clear.

    Matt Turck56:39

    So why is that? Why are we not here?

    Ben Rogojan56:59

    I think there's probably a few reasons. I think, one, we lose a lot of just brain—like, a lot of people leave, right? You have a lot of people who leave the industry before a certain point. And so you have a lot of new people come in. I say this as someone who, when I started, got very wooed by the idea of self-service analytics, right? I came in, Tableau was like self-service analytics.

    Ben Rogojan57:22

    And I was like, that's great. So I think that pushes a lot of the industry. Same thing with, again, on the business side, they hear a term, they're excited about it, they think it's going to work, and so we kind of look past all the stuff that actually matters. So things like, okay, but do we have the processes in place? Do we have the right people in place to actually make this happen? I think we're getting better.

    Ben Rogojan57:42

    I think we're starting to consolidate some of this knowledge. And I think that's the other part, is for a long time, all this knowledge was just everywhere, right? Yes, we had it all, but it was never consolidated into a playbook. It just kind of exists, right? Joe Reis only put out his book recently on data engineering, and I think the reason it did so well was because it was kind of the first book that consolidated information.

    Ben Rogojan58:07

    At a high level, it covered a lot of everything that you need to know. And I think that's why it was popular. So I think that's what we kind of need now, is just a playbook that becomes the thing that everyone's like, okay, this is how you run different teams. And I think there's always the business as well, where the business side pushes data teams one way or the other. And if you don't have a strong data leader that can either push back or make sure that you're doing things in a way that makes sense, you kind of can get trampled by the business side and what they need, and their kind of fickle nature where it's like, okay, now we're focusing on this, now we're focusing on that.

    Ben Rogojan58:25

    And that kind of can make things hard.

    Matt Turck58:31

    All right, and drumroll, prediction number five for 2025.

    Ben Rogojan58:56

    Yeah, I think we'll probably start seeing some vertical-specific data solutions come out, whether that's tooling or maybe just—one example I gave in my article was Tuva Health, which is trying to make it easier to just standardize healthcare, specifically claims data, which is something that I worked on a lot when I came into the industry. So that's why when I saw it, I was like, oh, this makes a lot of sense. Claims data in particular is decently standardized. I say that with quotes in the air because every company has their own system.

    Ben Rogojan59:24

    Some companies still have some older formats of data that people probably haven't seen, like positional files, where it just tells you column zero to 55 is this field, column 56 to 70 is a different field. But they've kind of tried to standardize some of that and build a vertical-specific kind of tool. I think that's something we'll hopefully see more. We've got all these tools that I'd like to think do maybe their job well, right? Like, okay, extract data really well, but they don't immediately drive value, right?

    Ben Rogojan59:51

    You still have to spend the next few months actually building something and translating that into some sort of business value. Just because you've got the data in your data warehouse doesn't mean it actually works. So I think that's kind of the next aspect that I'd like to see, and I hope to see, and we're starting to see. The big thing that usually holds that back is the fact that it's not customizable enough. So that's one reason I think some layer of open-source-esque access to the code will make sense, because people will always have these little, like, well, my data is shaped slightly differently.

    Ben Rogojan1:00:09

    I need a CASE statement for whatever, for how I'm defining this one product category. So it has to be this mix of flexibility plus standardization.

    Matt Turck1:00:21

    Okay, so maybe to close, where can people find you? We mentioned your YouTube, we mentioned your Substack. Where can people find you?

    Ben Rogojan1:00:37

    Yeah, like you said, YouTube, Substack. You look up Seattle Data Guy, I will show up there. Also on LinkedIn. But yeah, that's definitely where you can look me up if you need help or if you're just asking data engineering questions. You'll probably find some sort of answer in the content I've put out.

    Matt Turck1:00:54

    And so you produce a lot of great content. What would you recommend? Where should people start? So if people are interested in this general conversation and want to take next steps, what would be like two or three, whether it's video or post or what have you, that they should start with?

    Ben Rogojan1:00:56

    Oh, that's going to be a harder question to answer.

    Matt Turck1:00:59

    Which of your children do you like most?

    Ben Rogojan1:01:19

    Also, it's like, who are we talking to? If you're talking to someone who's maybe in a data leadership position, I have a few articles around that. One's like, "Don't Lead a Data Team Before Reading This." That's just meant to kind of call out some things that I think are important. And there's a few other ones, like "Thinking Like an Owner," that it's like trying to find your data projects. So that's also on Substack. But if you're starting more and you're just trying to understand data engineering, there's a ton of content there.

    Ben Rogojan1:01:41

    One that you can look up is just the Data Engineering 100-Day Crash Course, which is just meant to be a bunch of content that should, I think, either be free or mostly free that you can just kind of go through and be like, okay, here's everything I need to know, at least from a high level, that gets you started in the space. It's not going to make you a data engineer, but by the end, you should at least understand what you do know, what you don't know, and kind of where you maybe need to dig a little deeper.

    Matt Turck1:01:58

    Well, Ben, a.k.a. Seattle Data Guy, thank you so much. This was terrific. Lots of really interesting thoughts and content and insights.

    Ben Rogojan1:01:59

    Really appreciate it.

    Matt Turck1:01:59

    Thanks.

    Ben Rogojan1:02:01

    Yeah, thank you. Thank you, Matt.

    Matt Turck1:02:22

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.