TERBIS
All articlesData

Stop cleaning all your data: the trap of generic data quality projects

August 17, 2026

Quality matters, but not everywhere

It's true that AI places higher demands on data quality than most things that came before it. A model that makes decisions automatically amplifies errors in the underlying data instead of dampening them, as we covered in the previous article: the closer to automated action, the less tolerance for junk data.

That truth often leads to the wrong conclusion in practice. Many organizations respond to the demand for better data quality with a generic data quality project: an initiative meant to clean, standardize, and assure quality broadly across the entire business, before anyone even knows which decision the data is actually supposed to support. It sounds responsible. It's also, in most cases, wasted money.

Dark data: the part no one talks about

The reason has to do with how much of a company's data is never actually used. Splunk's global survey of more than 1,300 IT and business leaders found that 55 percent of organizations' data is what's known as dark data, meaning it's collected and stored, but never analyzed or used for anything. IBM, citing the same underlying data, reports that 60 percent of respondents estimated half or more of their data to be dark data, with a third estimating the share at 75 percent or more.

Dark data, by Gartner's definition, is information collected and stored in the course of regular business operations but generally never used for any other purpose: old log files, previous versions of documents, information about employees who left long ago, form responses that never got connected to an analysis. It's there out of old habit, because storage became cheap enough that no one ever had to say no.

Put the two things together: if 55–60 percent of your data is never used, and a generic quality project by definition sweeps across all data, you're largely cleaning a room no one ever enters. It's like deep-cleaning the whole house before a dinner party: vacuuming the attic no one's set foot in since the 90s, scrubbing the garage with the flat-tired bike, raking the forest out back, while your guests spend the entire evening in the living room. The work is still a cost. It just creates no value at all.

Why it feels right anyway

Generic data quality projects are appealing for three reasons that are all understandable, but wrong.

They feel fair: focusing quality work on only certain systems can feel like picking favorites internally, while "we're cleaning everything" feels neutral and easy to justify to a steering committee.

They feel future-proof: the idea that "we don't know what data we'll need tomorrow, so it's better if everything is high quality" is intuitive, but it ignores that quality work isn't a one-time effort. Unused data degrades again as soon as no one maintains it, which turns "future-proofing" into an ongoing commitment with no end date.

They're easier to sell internally than the alternative that actually works: a narrow, use-case-driven initiative feels small and unglamorous compared to a program branded something like "Data Quality 2.0" spanning the entire business.

What actually works: work backward from the decision

The right starting point isn't the data, it's the decision or action the data is meant to support. Just as in the data-to-action model covered earlier (descriptive, predictive, prescriptive), the quality bar differs dramatically depending on what the data will actually be used for. An internal dashboard tolerates far worse data than a system that makes automated decisions.

That gives a simple test question before starting a quality initiative: which specific decision, which specific model, or which specific system is this work for? If the answer is "generally, for the whole business," that's a sign the work should be broken down before it starts, not scaled up.

This is also exactly what the Data Asset capability (one of the five capabilities covered in the article on data value) is really about: not that all data is high quality, but that the data actually in use can be found, trusted, and applied to the right problem. That's a far narrower ambition than cleaning everything, and it's one that's actually achievable.

The conclusion isn't to ignore quality

The point isn't that data quality doesn't matter. The point is that quality work with no link to a specific use case is a falsely reassuring way to spend money on the 55–60 percent of your data that was never going to be useful anyway. Next time someone proposes a generic data quality project, the best question isn't "how much will it cost," it's "which decision gets better because of this, and why aren't we just cleaning the data that decision actually relies on?"

Sources: "The State of Dark Data," Splunk (global survey, 1,300+ IT and business leaders); IBM, "What is dark data?" (citing the same Splunk data and Gartner's definition of dark data).

Share

LinkedInX

Cite this article

Norström, A. (2026). Stop cleaning all your data: the trap of generic data quality projects. Terbis. https://terbis.se/en/articles/fallan-med-generella-datakvalitetsprojekt