You have a workflow worth fixing, you have the data it runs on, and no way to tell whether the second is good enough for the first. The two usual answers are both expensive. Wait for a broad "AI-ready data foundation" and the workflow stays manual for another year. Start on scattered inputs that were never stable enough for production and the build turns into data archaeology around the integration milestone, with the budget already spent.
Build-ready data sits between those extremes: accessible enough, representative enough, and stable enough that one specific workflow can move into delivery. It is a property of the workflow, not of the enterprise. And the property that decides it is not volume or cleanliness. On the property valuation engine ML LABS built for a PropTech platform, the data was abundant and the model was the easy half: the error sources that decided whether the system could be trusted were staged listing photos in postcodes where comparables ran thin, renovations the comparables panel had not caught up with, and satellite imagery going stale over active construction. Producing the estimate was straightforward. Knowing when not to trust it, and routing those cases to a human appraiser, was the build.

The Three Build-Readiness Tests
Data is build-ready when it passes three tests, in order, on the fields one workflow reads.
- Access test: the data can be pulled programmatically, on demand, without manual exports or one-off requests that then break on the second run.
- Signal test: the data carries enough signal for the workflow under realistic conditions: production-representative volumes with real noise, not a curated sample.
- Stability test: the source, schema, and quality profile are stable enough that the team is not firefighting upstream changes that invalidate finished work.
The access test is the cheapest and fails most quietly. If the workflow depends on manual exports or a spreadsheet somebody assembles, there is no build path: there is a person, and the system inherits their availability. What cannot be re-run cannot be operated.
Signal Beats Cleanliness
Reachable data can still be unusable. If key fields are mostly null, if labels arrive too late to act on, or if identifiers do not match across systems, the build looks viable until the model fails in ways that trace to the inputs. Research on data quality and ML performance (Mohammed et al., 2022) finds that which flaws are present matters more than how many.
That standard is measurable before a model exists, and one number makes it concrete. On the clinical automation ML LABS delivers for HeartSciences, the threshold that stops the work is extraction producing more than 15–20% of records with missing critical fields, or the same event coded differently across sites. Past that line, the automation produces more exceptions than it processes, as the billing automation case study says.
The valuation engine is the counter-example that matters, because its data was messy and yet it was still build-ready. Each failure mode had a measurable signal, and the spread between channel-specific estimates widened before the ensemble estimate became unreliable, which is what the confidence bound was calibrated against. The system shipped automated valuations within 10% of closing price in dense metro areas for 90% of cases, and routed the rest to a human appraiser rather than guessing at a number.
"We needed a valuation engine that could price properties within 10% of closing price in dense markets, in seconds instead of days. Omar built it, and it lifted our real estate insights and client engagement." — Innovation Director, the client firm
Stability Is An Upstream Property
A build also fails when the source moves underneath it. A study of production ML deployments at a major cloud provider (MLSys, 2022) found data drift causing accuracy drops of up to forty percent, even with routine retraining. A dataset that passes today's quality bar is not stable unless the process generating it is stable. The test is whether the team can name a small set of checks that should stay true over time: completeness on critical fields, freshness for scheduled loads, consistency for shared identifiers, validity for values in known ranges. If those checks cannot be named, the build is early.
When The Source Is The Project
If the source process lives in paper, in a vendor tool with no dependable access path, or in a system whose owner will not grant programmatic access, the data path is the project, and saying so early is worth more than a model that compensates for broken inputs.
Build-ready data is not perfect data. It is data the team can reach on demand, trust selectively, and monitor continuously.
There is a quieter version of the same blocker, and it is the expensive one. Data can be reachable, plentiful, and structurally wrong at the same time: aggregated in a way that discards the detail a model needs, or duplicated so that storage grows while information does not. Inside a hedge fund, correcting a foundation of exactly that shape cut storage costs by more than 60% and made the models perform 2% better, and none of it was visible as an error; what a live AI system costs when nobody is reading it is the full account.
First Steps
- Map the fields the workflow reads. Trace each to a source and record how it is accessed, refreshed, and owned. A field with no owner and no access path is the finding.
- Measure the critical fields, not the warehouse. Completeness, consistency, freshness, and validity on the core decision's fields, against a live sample, never a curated extract.
- Decide, and say which. If access and quality clear the bar, start building. If they do not, the data path is the first project, and it gets scoped as a project.
Judge Readiness At The Workflow Level
The bar is short enough to check this week: the fields are named, pulled programmatically, measured against live data, and each failure mode has a signal attached.
Where the data path is the open question, the fastest way to settle it is to prove it: a fifteen-minute call turns one workflow into an architecture, a build plan, and a price, and the build that follows runs against your real data, which makes readiness a demonstrated fact rather than an argued one. The readiness scan checks the rest of the operation in ten questions. When several workflows depend on the same broken path, see the data foundation an AI build actually needs; the messy-data build above is the property valuation engine in full.
References
- Mohammed, S., Budach, L., Feuerpfeil, M., Ihde, N., Nathansen, A., Noack, N., Patzlaff, H., Naumann, F., & Harmouch, H. The Effects of Data Quality on Machine Learning Performance. arXiv, 2022.
- Mallick, A., Hsieh, K., Arzani, B., & Joshi, G. Matchmaker: Data Drift Mitigation in Machine Learning for Large-Scale Systems. Proceedings of Machine Learning and Systems, 2022.



