Data downtime (via)

I work with a lot of different kinds of data, and I'm very interested in the processes around how we transform the piles and piles of messy information that are so ubiquitous these days into useful data. I'm learning about data observability on Coursera and just came across this article that I think articulates many of the biggest problems in data engineering and data science right now really well. In particular this point:

Data downtime — periods of time when data is partial, erroneous, missing, or otherwise inaccurate — only multiplies as data systems become increasingly complex, supporting an endless ecosystem of sources and consumers.

This hits home. It's so easy to just pull a dataset out of anywhere now, but we rarely give any thought to whether the data in it make sense. Virtually every dataset I come across has duplicate and missing values, obviously incorrect values, and doesn't line up with its metadata. It's a huge problem, and beginning to untangle it is a very complicated problem but one I'm super passionate about.

Related

Sep 25, 2025What “Supporting Our AI Overlords” and “Semantic Spacetime” Tell Us About the Future of Data InfrastructureJan 21, 2025Storage is cheap, but not thinking about logging is expensiveJan 6, 2025dltHubJan 4, 2025Software engineering is table stakesJan 2, 2025“Data engineering is the process of designing, building, and ”Jan 2, 2025DataTalksClub data-engineering-zoomcampJan 2, 2025Data Engineering 101Dec 31, 2024AWS Certified Data Engineer Associate exam

Kira Howe is a software engineer writing about Clojure, data, and the craft of building small, durable software.

More about Kira →