← Back to Articles

The Data Problems Nothing Warns You About

Take the process you know is broken: the one running on a spreadsheet, a mail thread, and somebody chasing people. The data it needs already exists inside your company.

The question that decides whether the fix works is not whether that data exists. It is whether the detail survives the trip. And the answer is almost never visible from where you sit, because the failures that matter do not throw errors. A process that duplicates records still returns rows. A monthly roll-up that destroys the variance a model needed still produces a table your dashboard reads happily. Nothing turns red. You find out when the finished system is confidently wrong about something a person would have caught.

Readiness is set per use case. For your one workflow, the data has to clear three bars:

  • Accessible. Reachable by a program, without a person exporting a spreadsheet
  • Sufficient. Enough volume and history for the approach the work actually needs
  • Preserved. The detail the system depends on survives every step between the source and the point of use, with nothing rolled up or dropped on the way

A dataset can miss any of the three without raising a single error.

A data manager finding a problem inside an archive box
The step that quietly drops the detail.

The Three Invisible Failures

Three failure classes sit between data that exists and data a system can act on. None of them raises an error, which is why they survive every audit built to look for errors.

Flowchart: source records with full detail pass through aggregation and rollup, then a warehouse with a duplicated layout, to what the system sees, which is what is left. Notes beside each step read: everything it could have used; variance discarded, irreversibly; one entity, counted twice; ceiling set upstream, debugged downstream.

Storage you pay for and never read. Volume grows when deletion is nobody's job and retention is nobody's budget line. The diagnostic is blunt: ask which tables the system actually reads, and then see whether anyone is able to answer the question.

Aggregation that discards the signal. The roll-up that makes reporting fast deletes the detail a model needed, and this failure is one-way. Once the raw detail is gone from the retained record, the ceiling is set upstream and no downstream tuning will raise it.

The cheapest data to store is the data you never collected. The most expensive is the detail you deleted before you knew you needed it.

The same entity, stored twice, disagreeing with itself. Two records, two identifiers, conflicting attribute values, and no shared key that resolves them into one. The duplicate entity gets counted twice, the evaluation inherits the leak, and the accuracy number that comes back from testing measures the wrong thing with great precision.

The list is not theoretical. On one engagement, a client stored large volumes of unneeded, polluted data, and nobody had noticed: the storage layout duplicated records, and the aggregation step feeding the models discarded information that mattered. Fixing the foundation cut storage costs by more than 60% and made the models perform 2% better. The full account of what happens to data nobody owns sits with the ownership argument.

"We came to Omar with messy data and a deadline we'd already missed once. He turned it into a working system in days, kept us updated throughout, and it held up in production, reliable in a way this field rarely is." — Head of Data, hedge fund

The same principle decided the outcome on a property valuation engine, which prices homes within 10% of closing price for 90% of cases. What made the bar reachable was not clean data. The spread between the photo-derived, satellite-derived, and comparables-derived estimates is the signal the confidence bound was calibrated against, and a store that averaged the channels at ingest would have left nothing to calibrate.

Audit The Path, Not The Warehouse

The audit starts at the workflow and walks backwards. Trace every field the system will read to the source that produces it. The expensive finding is a field that exists but reaches you through a manual export, a roll-up, or a system whose owner left the company.

Then ask what the path deletes, what it collapses, and what it stores twice. Every aggregation, deduplication, rounding step, and dropped column is a decision made without knowing what you would need, and each must be defended or reversed on purpose.

Write ownership down as a contract: the source system, the one person accountable for it, the schema, the freshness it promises, and the validations that must pass.

Five Checks On One Source

Run these on the source your workflow depends on, not the whole estate: completeness (null rates in the fields read), uniqueness (duplicate records that double-count an entity), consistency (one entity under different identifiers across systems), timeliness (feeds arriving when expected), and validity (values inside the ranges the source claims). Add one legal question: what personal data is in here, and may you use it this way?

The checks are gates, not a report. A check that logs a warning and lets the load proceed anyway is a check that documents the failure it was supposed to prevent, so wire them to block and let the first run fail loudly. What that first run rejects is the remediation scope, already priced, at the cheapest moment the gap will ever be visible.

First Steps

  1. Name one workflow and list its fields. Not a data strategy — the fields. Every column the system will read, traced back to the source system that produces it today.
  2. Ask the path what it destroys. Walk from source to point of use, marking each step that aggregates, deduplicates, rounds, or drops. Each is a decision to defend or reverse.
  3. Wire the five checks as blocking gates on that one source. Let the first run fail loudly, and cost what it rejects before anyone commits budget to the build.

Prove The Path Before The Build

If the workflow is named and the data path is what nobody can vouch for, the next move is not a company-wide warehouse program. It is a narrow proof against the real sources: one path that preserves what the work needs, five checks wired as gates, one named owner per source, and a queryable store. A scheduled script loading into PostgreSQL clears this bar for a first build; heavier tooling can wait until scale makes it necessary.

Proving that path on your own data, against your own sources, before the build is budgeted, is exactly what the first call is for. The bar for what it has to hand back is set out in what to demand before you sign. The foundation that comes out the other side is small, owned, and specific, and it is the one everything after the first build stands on.

NEXT · TO PRODUCTION

Find the real blocker.

What is slowing delivery, and the fastest path through.

15 minutes · no charge · with Omar