A proposal is in front of you. It has a number on it, a timeline, and a team behind it who believe in it. The question nobody in the room has answered is whether anyone independent has checked the three things that decide whether it can work at all: the decision it is meant to change, the data it will run on, and the systems it has to plug into.
The failure statistics are not in dispute. A study of AI project failure (RAND, 2023), built on interviews with 65 experienced data scientists and engineers, puts the overall rate above 80%, twice the failure rate of non-AI IT projects. The root causes their practitioners name are not modeling causes at all, but misunderstood problems, inadequate data, missing infrastructure, and a fascination with the technology rather than its purpose.
ML LABS has built and run heavy-workload systems in production over fifteen-plus years, across healthcare, telecom, PropTech, and finance. The three causes below map onto that taxonomy, and each is settled long before anyone opens a notebook.

The Three Root Causes
Each cause acts at a different lifecycle stage and needs a different intervention. They also compose badly: an unclear objective guarantees the wrong data will be collected, and the wrong data guarantees that integration will surface problems nobody budgeted for.
1. Unclear Business Logic
A vague objective, "predict customer churn" or "optimize pricing", goes to a data science team, which optimizes what it said. That is how a company gets a churn model that flags customers who already left, or a pricing model that undercuts its contractual floors.
The fix for this is a target, and a target is a specific object rather than an intention. This is schematic, not a client's spec: "we need to predict demand" commits to nothing, while "next-week SKU-level forecasts, accurate enough that purchasing can cut overstock write-offs by roughly a fifth against last year's baseline" commits to something a model can be checked against, and something for which a vendor can be refused payment.
That last clause is not rhetorical. ML LABS writes the targets into the contract before work starts, and acceptance is judged against them, not anyone's impression of the result.
2. Data-Production Mismatch
Data in a lab is curated. Production data is whatever the business actually recorded, including the parts that nobody looked at. On the property valuation engine that ML LABS built for a PropTech platform, the sources of error that decided whether the system could be trusted were exactly the ones no curated dataset contains: staged or seasonally misleading listing photos in postcodes where comparables ran thin, renovations that the comparables panel had not caught up with, satellite revisit gaps that left construction sites stale in the spatial channel for weeks. Each failure mode had to be caught by a measurable production signal and priced into the confidence bound. The system hit its accuracy target, within 10% of the closing price for 90% of cases in dense metro areas. Producing the estimate was the easy half; calibrating when to trust the estimate was the harder build.
Data also fails silently and expensively. A hedge fund was storing large volumes of unnecessary and polluted data without noticing: the aggregation mechanism was discarding information that mattered for model training, and the storage structure was duplicating what it kept. Correcting it cut storage costs by more than 60% and made their models 2% better; the full account sits in what ownership of a live system actually buys.
3. Integration Validated Too Late
A model in a notebook is not a product. Research on hidden technical debt (NeurIPS, 2015) showed the ML code is a small fraction of a production ML system, and the rest, data plumbing, serving infrastructure, configuration, monitoring, is where the schedule lives.
The clinical messaging layer ML LABS built for the cloud ECG backend at HeartSciences shows it plainly. The AI result itself was never the hard part. Whether a result reached the ordering clinician depended on what one hospital's EHR would accept: patient identifiers with or without leading zeros, an ordering provider reference pointing at a retired NPI, an observation identifier using a LOINC code the EHR had never mapped to a local result type.
Integration is not a phase after the model. Validate consuming systems before the model produces anything worth consuming; a hardcoded return value is enough to prove the path.
Front-Loading The Expensive Questions
Front-loading is not a virtue, it is arithmetic. A business-logic gap caught in a conversation costs a conversation; caught after model development, it costs the model and the budget that paid for it. Three gates catch each root cause while it is still cheap.
- Map the decision, not the data. Write down the decision the model informs, the action on its output, the outcome that counts as success, and every system that consumes it.
- Prototype the whole path, not the model. Build the smallest end-to-end system that exercises the real integration path against production-representative data, with the model returning constant values if necessary. Integration failures and data defects surface here, where they cost days to fix instead of costing quarters.
- Deploy against a real slice. Route a subset of live traffic under full monitoring, and expand only when the business metric moves, not when the model metric does.
Buy The Decision Before The Build
The cheapest failure is the one that never gets built. A major US TV network had been quoted a full software system for a workflow that did not require one. ML LABS ran a scoping session, and its deliverable was a finding that the system should not be built.
"Omar delivered in two weeks what our team estimated would take six months. The scoping session alone saved us from a $200K mistake." — AI Program Manager, a major US TV network
Settling the decision is what the first call is for: whether the thing should exist, what it must be checked against, and what the data will support. When the answer is no, that is not a failed engagement; it is the product. When the answer is yes, what follows is a monthly engineering partnership that builds one workflow into production and keeps it running.
References
- RAND Corporation. The Root Causes of Failure for Artificial Intelligence Projects. RAND Corporation, 2023.
- Sculley, D., et al. Hidden Technical Debt in Machine Learning Systems. NeurIPS, 2015.



