← Back to Intel

How to Find Your First AI Use Case

Jul 7, 2026Omar Trejo9 min read

You have a list. It came out of the meeting where someone asked what the company is doing about AI, and it holds eight ideas, or twelve, each a sentence long and each defensible in a room. None of them are costed. The pressure now is to pick one and start, because starting looks like progress.

That is the moment the first AI budget gets committed to the wrong problem. Research on organizational AI adoption (McKinsey, 2024) puts adoption at 72% of organizations, while research on AI value realization (BCG, 2024) finds only a minority reporting meaningful value from their first deployments. Research on AI initiative failure (RAND, 2024) locates the most common root cause not in the technology but in misunderstandings about what the project was for — a selection failure, not an engineering one. The companion piece on why AI projects fail follows that failure downstream, after the money is spent.

ML LABS has taken more than ten heavy-workload systems to production over fifteen-plus years, across healthcare, telecom, PropTech, and finance. The practice will not start one without a written target in the contract — the number the system has to hit for the work to count as done — because until the client accepts against that target, the fee is refundable. That single term does most of the selection work below: a use case whose success nobody can state in advance is a use case ML LABS cannot sell, however well it presents in a deck.

The Three Selection Traps

The wrong problem does not look wrong. It looks like the most interesting item on the list, and it fails for one of three reasons, all visible before a line of code is written.

  • The shiny object. Chosen because it demonstrates well, not because it changes a decision anyone makes. Demos have no failure modes and production has nothing else — the filter for which pilots deserve production exists for this one.
  • Boiling the ocean. Automates a whole process when a narrow slice carries most of the cost. Scope inflation is why a first project ends up with no date on which anyone can say whether it worked, and why building only the part nobody else owns matters most on a first attempt.
  • The data fairy tale. Assumes data that is described accurately in a data dictionary and exists in that form nowhere. This is the trap the other two hide behind.

None of the three is a technology failure. Each is a choice made before the technology was involved — which is why a structured evaluation catches all three for the price of an afternoon.

Four Tests Before Any Build

Score each candidate one to five on four factors, and multiply — never average. Averaging lets a strong business case hide a data problem. Multiplying makes the weakest factor dominate the composite, which is what that factor will do in production anyway. Research on structured evaluation of AI initiatives (MDPI, 2025) finds that evaluating candidates before implementing them measurably improves time and cost outcomes.

The composite is not a grade to hit. It is a diagnostic pointing at one thing: the lowest factor is what will kill the project, and it is the only factor worth arguing about in the room.

graph TD
    A["8-12 Candidates"] --> B["Score Four Factors 1-5"]
    B --> C["Multiply, Never Average"]
    C --> D{"Weakest Factor"}
    D -->|"Low"| E["Fix It Or Drop It"]
    D -->|"Strong"| F["Write A Checkable Target"]
    F --> G{"Target Defensible?"}
    G -->|"No"| E
    G -->|"Yes"| H["Fund The Build"]

    style A fill:#1a1a2e,stroke:#0f3460,color:#fff
    style B fill:#1a1a2e,stroke:#0f3460,color:#fff
    style C fill:#1a1a2e,stroke:#0f3460,color:#fff
    style D fill:#1a1a2e,stroke:#ffd700,color:#fff
    style E fill:#1a1a2e,stroke:#e94560,color:#fff
    style F fill:#1a1a2e,stroke:#0f3460,color:#fff
    style G fill:#1a1a2e,stroke:#ffd700,color:#fff
    style H fill:#1a1a2e,stroke:#16c79a,color:#fff

Business Impact

The use cases that pay are the ones where a decision recurs faster than people can make it. On the roaming optimization platform ML LABS built for a top 10 global telecom company — an engagement run through Gigster — the decision space was structurally larger than the analysts could traverse: by the time one corridor had been modeled in a spreadsheet, the demand pattern behind it was stale. That is the shape of a fundable use case, and the roaming optimization case study records the return — optimization opportunities across 128% more corridors than the original scope targeted, and over 12x return in the first year against engagement cost.

  • How often does this decision recur, and what does one instance cost today?
  • What would a step-change improvement be worth over a year?
  • Can you name the number that has to move, and the date you would check it?

Data Readiness

This is where most evaluations fall apart, and self-assessment is why. A study on AI-ready data (Gartner, 2025) projects that organizations will abandon 60% of AI projects that are not supported by AI-ready data. On a hedge-fund engagement, ML LABS found the client storing large volumes of unnecessary and polluted data without knowing it: the aggregation step was discarding information their models needed, and the storage structure was duplicating records. Correcting it cut storage costs by more than 60% and their models performed 2% better — and this was a fund whose models are the product, with the full account sitting alongside the ownership argument it anchors.

  • Is the data in one place, or is "the data" a claim about four systems that disagree?
  • Is there enough history for the method, and is any of it labeled?
  • Has anyone opened a sample, or is quality still a self-report?

Technical Feasibility

A first use case should not require a research result. Classification, regression, and extraction are proven patterns with known failure modes and known costs; anything whose feasibility rests on a paper published this year is a bet, and a first project is the wrong place to make one.

The model is not always the harder half. On the property valuation engine ML LABS built for a PropTech platform, producing the point estimate was the straightforward part — calibrating the confidence band around it was harder, and the band was what decided whether the system could run without an appraiser standing behind it. Research on hidden technical debt (NeurIPS, 2015) makes the general case that the model is the small part of a production system; the valuation engine case study makes the specific one, in staged listing photos and stale satellite imagery over active construction.

  • Is there a proven pattern for this, or does it need a result nobody has published?
  • What happens when the model is wrong, and what does that cost?
  • Who finds out it was wrong, and how long does that take?

Organizational Fit

The best first use case changes how people work as little as possible. Research on AI adoption approaches (Deloitte, 2024) finds augmentation-first approaches reaching higher adoption and faster time to value than approaches that automate a role outright.

The failure this factor prevents is not resistance — it is the result that arrives nowhere. The cloud ECG backend ML LABS built for HeartSciences produced accurate diagnostic findings, and the engineering that made them count was the messaging layer that delivered those findings back into the hospital systems which had ordered them, in a form each hospital would accept without rejecting the message. A finding that never reaches the clinician who ordered it changes nothing, however accurate it is. Where AI replaces the interface rather than sitting inside it, adoption becomes a pricing question too — the companion piece on pricing when agents replace UIs takes that up.

  • Does the output arrive inside the tool people already open, or in a new one they must remember?
  • Does it augment a decision someone already makes, or take that decision away?
  • Who is accountable when it is wrong, and do they know that yet?

Sometimes The Answer Is None

The outcome most selection processes never allow themselves is the one that pays best: none of these, not yet.

A major US TV network had already received a quote to build a full software system for a workflow that did not require one. The scoping session ML LABS ran did not pick a better candidate from a list of candidates. It established that the build should not happen — and the network kept the money.

"Omar delivered in two weeks what our team estimated would take six months. The scoping session alone saved us from a $200K mistake."

— AI Program Manager, a major US TV network

A $750 conversation that stops a $200K build is the highest-return line item in the program, and the one nobody puts on the roadmap. Cheaper software makes that discipline more valuable, not less: as software gets cheaper to produce, the binding constraint moves to deciding what is worth producing at all — the same question as whether to build the capability or buy it. A session that decides is a different instrument from a design engagement that proves it by running, and knowing which you need is itself a selection decision.

Where The Framework Breaks

The framework assumes the score decides, and it does not survive a sponsor who has already chosen. When the decision is made before the scoring starts, scores get adjusted to fit it, a low data-readiness factor becomes something to figure out later, and the workshop turns into theater with a spreadsheet attached.

Research on organizational barriers to AI adoption (HBR, 2025) puts a committed sponsor who defends the framework's output ahead of methodological quality — and that cuts both ways. Scoring is a tool for a decision that is genuinely open. If the decision is closed, no rubric reopens it, and the work is upstream, on who owns the call.

First Steps

  1. Collect candidates from the people who do the work. Ask for eight to twelve, one sentence each, weighted toward operations and front-line staff rather than the people who present to the board.
  2. Score independently, before anyone talks. The divergences are the output that matters — a candidate two groups score far apart is hiding a constraint one of them can see and the other cannot.
  3. Open the data on the top candidate before funding it. Not the data dictionary — the data. Pull a sample and check the fields the use case depends on.

Fund The Defensible Candidate

Score the four factors, multiply, and let the weakest one speak first. Then do what the score cannot do for you: write the target. A candidate has survived selection when someone can state, in one sentence, what the system must achieve to count as working — and someone else is willing to be measured against that sentence. The score tells you which candidate is worth writing it for; the sentence is what turns a use case into a project.

If you have a list and no scores, the work is internal and the First Steps above are enough. If your strongest candidate keeps failing on data readiness, selection is not the problem you have, and the data foundation work comes first. And if you are holding a shortlist nobody can rank, a build quote nobody can evaluate, or a candidate everyone likes and nobody can write a target for, an AI scoping session maps the candidates against these four factors, names the one worth funding, and produces the written recommendation the funding decision needs — including, where the evidence points there, the recommendation not to build at all. That answer is worth as much as the other one, and it costs far less to act on.

References

  1. McKinsey & Company. The State of AI. McKinsey Global Survey, 2024.
  2. Boston Consulting Group. From Potential to Profit With GenAI. BCG, 2024.
  3. RAND Corporation. The Root Causes of Failure for Artificial Intelligence Projects. RAND Corporation, 2024.
  4. Oesterreich, T. D., et al. Artificial Intelligence in Project Success: A Systematic Literature Review. MDPI Information, 2025.
  5. Gartner. Lack of AI-Ready Data Puts AI Projects at Risk. Gartner Newsroom, 2025.
  6. Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J., & Dennison, D. Hidden Technical Debt in Machine Learning Systems. NeurIPS, 2015.
  7. Deloitte. State of AI in the Enterprise. Deloitte Insights, 2024.
  8. Harvard Business Review. Overcoming the Organizational Barriers to AI Adoption. Harvard Business Review, 2025.
NEXTTO PRODUCTION

Check your position.

Two minutes. Your main blocker and first move.

Fixed scope · written plan · Design and Build: full refund until you accept