← Back to Intel

Why Your AI Tools Are Losing Money

Jun 27, 2026Omar Trejo8 min read

You can produce the number your AI tools cost you. Marketing has one for content, engineering adopted a coding assistant, support deployed a chatbot, finance is experimenting with forecasting, and each of those lines has an invoice behind it that finance can pull this afternoon. Now produce the other number — what they returned. If that second number takes weeks to assemble, or arrives as a set of anecdotes about how much people like the tools, the portfolio is not being managed. It is being renewed.

The macro picture says this is the normal condition rather than an embarrassing exception. A survey on AI adoption (McKinsey, 2024) put adoption at 72% of organizations, while research on generative AI scaling (BCG, 2024) found only 26% getting past proof-of-concept into anything that produces value. An analysis of AI investment returns (Deloitte, 2024) names the paradox directly: 91% of organizations plan to increase AI investment while most take extended periods to reach a return they would call satisfactory. And a survey on generative AI deployment (Gartner, 2024) found 49% of executives naming difficulty demonstrating value as their top concern. Spend is rising, evaluation is not, and the two curves are not connected by anything.

Unowned spend does not announce itself, and the sharpest version of that ML LABS has met was not in a tool portfolio at all — it was a hedge fund's storage layer, where correcting a data foundation nobody had read cut storage costs by more than 60% and made the fund's models 2% better. Nothing had failed. No alert had fired. The bill had been paid, correctly, every month, and the full account of what a live system costs when nobody owns it is where that engagement is told. The mechanism transfers exactly to a tool stack: a line item nobody reads is a line item that renews.

No AI tool in your stack is failing an evaluation. That is the finding — there is no evaluation to fail, and the renewal happens anyway.

Adoption And Evaluation Never Meet

The structural cause is a split in ownership, and it survives every good intention aimed at it. The person who adopts a tool is choosing at the moment of maximum optimism, against a demo, with a budget that clears without a business case. The person who would evaluate it — if anyone were assigned — works on an annual cycle, sees a line item rather than a workflow, and has no standing to remove something a department says it needs. Those are different people, different clocks, and different information, and nothing in the process forces them into the same room.

By the time a renewal comes around, the switching-cost argument has arrived to protect the tool regardless of what it returns. Migration effort is concrete and immediate; forgone value is abstract and deferred. A tool that has produced nothing measurable and a tool that has produced a great deal look identical at renewal, because the organization never built the instrument that would tell them apart. That is not a technology failure. It is the absence of an owner for one specific question: is this working, and how would we know?

The Audit That Binds Decisions

An audit that ends in a report changes nothing. This one is designed to end in binding keep, cut, and consolidate decisions, which means it needs an owner with the authority to execute them before it starts.

quadrantChart
    title AI Tool ROI Assessment Matrix
    x-axis Low Usage --> High Usage
    y-axis Low Business Impact --> High Business Impact
    quadrant-1 Scale and invest
    quadrant-2 Investigate barriers
    quadrant-3 Eliminate
    quadrant-4 Consolidate or retrain

Register Every Tool, API, And Model

For each one: monthly cost, active users in the last thirty days rather than seats purchased, usage frequency, the business function it serves, and a named owner. The register is where duplicate spend surfaces on its own — two teams paying separately for capabilities that overlap is a consolidation decision that needs no evaluation at all, only visibility. Shadow purchases on personal cards belong in the register too; the ones that never appear in procurement are the ones nobody has ever assessed.

Define Impact Per Category, Not Per Tool

Measuring a coding assistant and a forecasting model against a single metric produces a number that means nothing. Each category earns its own definition of impact, and each needs a threshold below which the tool enters the elimination conversation regardless of how much its users enjoy it. The discipline is the same one that separates an output metric from an outcome metric — measuring AI impact by decisions rather than by accuracy — applied to a purchase instead of a system.

  • Productivity tools: hours returned per user per week. Experimental evidence on generative AI and productivity (Science, 2023) measured roughly a 40% reduction in task completion time — but only for tasks inside the model's capability range, which is why the metric has to be attached to a specific workflow rather than to a seat.
  • Customer-facing tools: resolution rate, satisfaction delta, conversion impact, and escalation frequency, each measured before and after deployment rather than in a vendor's benchmark.
  • Process automation: throughput, error rate, and time-to-completion on the target workflow — the workflow the tool was bought for, not the department it sits in.

Price The True Cost, Not The License

The license is the visible fraction. Research on AI productivity effects (NBER, 2023) found a 14% productivity gain among customer support agents using an AI assistant, and the number worth attending to is not the 14% — it is that the gain was realized inside a workflow that had been instrumented, monitored, and adjusted around the tool. That surrounding work is a cost, and it belongs in the total.

  • Integration and maintenance: the engineering hours that keep the tool connected to the systems it reads and writes.
  • Adoption: onboarding, training, and the ongoing support burden that lands on whoever became the internal expert.
  • Risk and exit: data exposure, vendor lock-in, and the migration effort you are implicitly agreeing to pay later.

Decide, Then Consolidate

Place every tool in the matrix and act on the quadrant. High impact and high usage earns investment and a pricing negotiation. High impact and low usage is a distribution or training problem, not a product problem. Low impact and high usage is popular and not moving anything — consolidate it. Low impact and low usage gets deactivated, and if nobody protests within a cycle, cancelled.

Consolidation is where the largest structural saving sits, and it has a specific rationale. Research on AI's uneven impact across tasks (HBS, 2023) found a "jagged technological frontier" — AI performance varies sharply between tasks that look equally difficult from the outside. A fragmented stack forces every workflow to cross a tool boundary at exactly the points where that jaggedness bites, and each boundary is an integration nobody owns. Fewer, better-integrated tools beat more specialized ones not because integration is elegant, but because a workflow that stays inside one system is a workflow somebody can actually measure.

No Owner, No Audit

This framework has one precondition and it is unforgiving: centralized visibility into spend, and a named owner per tool. Without the first, the audit stalls at the register — half the stack is invisible, and the half you can see is the half that went through procurement, which is not where the waste concentrates. Without the second, the audit produces findings that nobody has authority to act on, and a finding without authority is a slide.

Assign the owners before the audit rather than as a conclusion of it. An owner is the person who will either defend the tool's value with evidence or consent to its removal without a fight — and the assignment itself does useful work immediately, because a tool nobody will agree to own has already been evaluated.

First Steps

  1. Give one person the mandate and the billing access. They need to pull spend and usage across every department, including the purchases that never touched procurement. Without that access, the register is fiction.
  2. Ask the users five questions. Which tools, how often, for what task, what would break if it vanished, and what they use instead when it fails. The last two questions are the ones that separate a habit from a dependency.
  3. Put a decision date on the calendar before you start. An audit with no deadline becomes a permanent workstream, and a permanent workstream is another unowned cost.

Put One Owner On The Spend

Replace tool-led adoption with outcome-led ownership. Every AI tool maps to a measurable business movement, a named owner, and a decision date — and the spend, the usage, and the impact live in one place that somebody reads on a schedule rather than at renewal. That is the whole intervention, and it is unglamorous by design: the reason the waste accumulated was never that the tools were bad, it was that the question of whether they were working belonged to nobody.

Ownership is also the thing that is genuinely hard to staff, because it is a standing job rather than a project. Managed AI operations is that job bought as a service: one accountable owner across a defined scope of live AI systems and spend, with cost and drift review as standing agenda items rather than annual surprises, a monthly written brief covering what ran, what changed, and what is at risk, and cancellation on thirty days' notice — so the arrangement has to keep earning the month it is in. It runs $12,000 a month, which is the number to hold against the spend the register surfaces. If the audit shows the tools were never the problem and the pipeline underneath them was, the data foundation a working AI build actually needs is the more useful place to look next.

References

  1. McKinsey & Company. The State of AI. McKinsey Global Survey, 2024.
  2. BCG. From Potential to Profit With GenAI. Boston Consulting Group, 2024.
  3. Deloitte. AI ROI: The Paradox of Rising Investment and Elusive Returns. Deloitte Insights, 2024.
  4. Gartner. Gartner Survey Finds Generative AI Is Now the Most Frequently Deployed AI Solution in Organizations. Gartner, 2024.
  5. Noy, S., & Zhang, W. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence. Science, 2023.
  6. Brynjolfsson, E., Li, D., & Raymond, L. Generative AI at Work. National Bureau of Economic Research, 2023.
  7. Dell'Acqua, F., Mollick, E., et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School, 2023.
NEXTTO PRODUCTION

Audit your AI spend.

Two minutes. What to keep, cut, or fix.

Fixed scope · written plan · Design and Build: full refund until you accept