Categories
Architecture

Should Bad Data Be Allowed Into Your Silver Layer?

When combining and curating data in your Silver layer it’s not uncommon to find entries don’t always fit quite right. It could be missing references, incomplete data, or broken validation rules.

Should that bad data be allowed into your carefully crafted Silver layer? And if not, where should that data go?

Let’s look at some options for handling the bad data and how they compare.

Reject it?

The simple solution is rejection – ignore the record and move on. This could be explicit filtering, or implicit such as an INNER JOIN excluding unknown entities. The two key issues with this are:

Firstly, outright dropping a record removes visibility. I’ve seen calendars which were populated annually and issues arose of missing sales at the start of a new year. Silently discarding records meant transformations looked fine, but data quality had taken a nosedive.

Secondly, if incomplete data is removed, are we confident that a complete record will arrive later? If not, we’re creating a risk of ignoring data which may be at least partially useful.

Sometimes – for very bad data – rejection may be a fair solution. In most cases though, the lack of visibility and no guarantee of a follow up are too great of a risk to take this approach.

Quarantine it?

The first alternative to consider is quarantining the bad data outside of the curated dataset. This retains the record for visibility, but it doesn’t dilute the final data quality. It’s similar to a dead letter queue.

This option solves the visibility issue as we now retain a copy of the bad data. We can build data quality processes around this to actively quantify, monitor, and notify us if an upstream change or a bad release impacts data quality.

With bad data quarantined, we now have options for remediation.

Holding out for valid data is still an option. However we can introduce a workflow around quarantined records so they can be corrected and reprocessed into the curated set. Options for correction can vary dramatically between automated, technical, or business processes. For another post.

Quarantining bad data is a solid approach to ensure all data is processed and ends up somewhere. It solves the issues of straight rejection, but can require some form of manual intervention to review and resolve the data quality issues.

Keep it?

Another approach would be to retain the good parts of the bad data and allow those into the curated set. This may reduce data quality but provides completeness.

Let’s consider sales data which refers to an unrecognised product. We could create placeholder product records (typically called inferred members) with the data which we do have. Quarantining those records would lead to sales figures being understated. Although we may not know what was sold, visibility of sales values could be preferable.

Placeholder records should be signposted accordingly to indicate they’re incomplete or came via a different route, such as a data quality flag or source identifier. This provides similar visibility to quarantining along with the ability to quantify and monitor.

The key difference to quarantining is that you trade off quality for data completeness. The dataset will be complete, but data quality will be degraded until the reference data arrives or a correction workflow like above is in place.

Placeholders are a great way to get as much data as possible into the curated model sooner, provided reduced quality is acceptable. The drawback is that if reduced quality is ‘good enough’, it may be hard to embed a correction process.

Wrap up

Data quality isn’t binary. Here we’ve looked at different approaches for handling bad data. Depending on the definition of ‘bad’ data, these won’t always be applicable, but as a rule of thumb:

ApproachOutcomeQualityBest when
RejectionDiscardedUnaffectedData isn’t worth keeping
QuarantineRetailed separatelyUnaffectedQuality is paramount
PlaceholdersRetained in SilverReduced, until resolvedCompleteness matters more

Personally, my preference is to quarantine impacted records. Segregation, clear visibility, and an enforced remediation workflow deliver a quality outcome. That’s not always easy to embed into an organisation though. When concessions are needed, placeholders can be a solid compromise where completeness is more important than quality.

Our Silver layer should be clean, but won’t always be spotless. What’s important is a common understanding of what quality means in that context. And when it falls short, we should be able to quantify the gap and, where needed, remediate it.

Leave a comment