You already know which processes are broken. There is the one that runs on a spreadsheet nobody can open, the one where a person spends every Tuesday chasing the numbers, and the one that has a filed date attached to it and no owner who can say how it will be met.
The question is not whether to build software. It is which of these becomes software first, and that decision is where the money is usually lost. Research on AI initiative failure (RAND, 2024) locates the most common root cause not in the technology but in misunderstandings about what the project was for: a selection failure rather than an engineering one.
ML LABS has taken more than ten systems to production over fifteen-plus years, and will not start one without a written target in the contract: the number the system has to hit for the work to count as done. That single term does most of the selection work below.

Four Tests Before Any Build
Score each broken workflow one to five on four factors, and multiply the scores rather than averaging them. Averaging lets a strong business case hide a data problem. Multiplying makes the weakest factor dominate the result, as it will in production anyway. The lowest factor is what will end the project, and the only one worth arguing about. When it cannot be raised, the workflow is not the one to start with, whatever the total says.
Business Impact
The workflows that pay are those where a decision recurs faster than people can make it. On the roaming optimization platform ML LABS built for a top 10 global telecom company, an engagement run through Gigster, by the time one corridor was modeled its demand pattern was stale. That is a workflow worth funding, and the roaming optimization case study records the return: optimization opportunities across 128% more corridors than the scope targeted, and over 12x return in the first year against engagement cost.
- How often does this decision recur, and what does one instance cost today?
- Can you name the number that has to move, and the date you would check it?
Data Readiness
This is where most comparisons fall apart, and self-assessment is the reason. On a hedge-fund engagement, ML LABS found the client storing large volumes of unnecessary and polluted data without knowing it: the aggregation step was discarding information their models needed, and the storage structure was duplicating records. Correcting it cut storage costs by more than 60% and their models performed 2% better, as the full account records.
- Is the data in one place, or is "the data" a claim about four systems that disagree?
- Has anyone opened a sample, or is quality still a self-report?
Technical Feasibility
A first build should not require a research result. Classification, regression, and extraction are proven patterns with known failure modes and known costs. Anything whose feasibility rests on a paper published this year is a bet, and a first project is the wrong place to make one.
The model is not always the harder half. On the property valuation engine ML LABS built for a PropTech platform, producing the point estimate was the straightforward part. Calibrating the confidence band around it was harder, and the band decided whether the system could run without a human appraiser behind it; the valuation engine case study has the detail.
- What happens when the system is wrong, and what does that cost?
- Who finds out it was wrong, and how long does that take?
Organizational Fit
The best first workflow changes how people work as little as possible. The failure this factor prevents is the result that arrives nowhere. The cloud ECG backend ML LABS built for HeartSciences produced accurate diagnostic findings, and the engineering that made them count was the messaging layer that delivered those findings back into the hospital systems that ordered them, in a format each hospital's system would accept without rejecting the message. A finding that never reaches the clinician who ordered it changes nothing.
- Does output land in a tool people already open, or a new one they must remember?
- Who is accountable when it is wrong, and do they know that yet?
Sometimes The Answer Is None Of Them
Most selection processes never allow the outcome that pays best: none of these, not yet.
A major US TV network had received a quote to build a software system for a workflow that did not need one. The ML LABS scoping session did not pick a better candidate from a list. It established that the build should not happen, and the network kept the money.
"Omar delivered in two weeks what our team estimated would take six months. The scoping session alone saved us from a $200K mistake." — AI Program Manager, a major US TV network
A conversation that stops a $200K build is the highest-return line item in the program. As software gets cheaper to produce, the binding constraint moves to deciding what is worth producing at all, the same question as whether to build the capability or buy it.
First Steps
- List the workflows the people doing the work would name. Eight to twelve, weighted toward operations rather than the people who present to the board.
- Score independently, before anyone talks. The disagreements are what matters. A workflow two groups score far apart hides a constraint only one of them can see.
- Open the data on the leading workflow before funding it. Not the data dictionary, the data. Pull a sample and check each of the fields the build depends on.
Fund The Workflow You Can Write A Target For
Score the four factors, multiply, and let the weakest one speak first. Then write the target. A workflow has survived selection when someone can state, in one sentence, what the system must achieve to count as working, and someone else is willing to be measured against that sentence. The score tells you which workflow is worth writing it for.
If your leading workflow keeps failing on data readiness, the data foundation work comes first. If you are holding a shortlist nobody can rank, or a workflow everyone agrees is broken and nobody can write a target for, a fifteen-minute call maps the candidates against these four factors, names the one worth funding, and produces the recommendation, including, where the evidence points there, the recommendation not to build at all.
References
- RAND Corporation. The Root Causes of Failure for Artificial Intelligence Projects. RAND Corporation, 2024.



