Making data useful

How do we turn messy data into profit?

Making data useful is difficult. The whole point of data engineering/science/analysis/etc is to turn all of the endless piles of information we collect into something that people find valuable. I've spent the better part of the last decade in my career as a software engineer building systems that do this, and I've noticed that it is remarkably difficult to do well. There are many points where the process breaks down, and no silver bullet for fixing them. I've spent a lot of time trying to organize my thoughts on the topic and wrote this down one night:

Turning chaotic and unorganized information into useful insights

This is one path that the piles of chaotic information we hoard can take to become useful. One problem is that executing all of these steps well requires at minimum a software engineer, a data scientist, and a business analyst.

Without an engineer your data pipelines will be unreliable and incorrect, without a data scientist your analysis will be misleading, and without a business analyst the results will be meaningless and never reach the people who need to see them. That's easily $0.5M/year in headcount alone, a huge expense even if you manage to find a team of open source superstars who can do all of this with free tools and minimal infrastructure.

Another problem is that the steps along this path are not neatly delineated in any way, making them hard to outsource or share. You need the whole team working closely together start to finish to clear a path for your data to flow through the organization.

There are tools that can help make some of this smoother, but it's a fundamentally complex problem to optimize in a cost-effective way. It's really interesting to think through where the pain points are, and how technology can help address them. This is what's taking up a lot of my bandwidth lately, so if any of it resonates with you I'd love to chat sometime!

Related

Storage is cheap, but not thinking about logging is expensiveJan 21, 2025What “Supporting Our AI Overlords” and “Semantic Spacetime” Tell Us About the Future of Data InfrastructureSep 25, 2025Data downtimeFeb 18, 2025Could Disposable UI Solve Data Science's Low ROI Problem?Jan 8, 2025dltHubJan 6, 2025Software engineering is table stakesJan 4, 2025what-is-data-engineeringJan 2, 2025DataTalksClub data-engineering-zoomcampJan 2, 2025

Kira Howe is a software engineer writing about building with care in the age of AI.

More about Kira →

Follow