← Back to Articles

Your System Is Getting Worse And Nothing Warns You

The system went live, it works, and it has been running ever since. That is usually the moment when a company stops paying attention to it, and it is also the moment the value starts leaving. A production AI system is not an asset that sits still once it ships. It is a position, and a position has to be held — every gain in it is held in place by somebody, and remove that somebody and the same machinery starts to run in reverse.

The models themselves do not hold still. Research on temporal degradation in deployed models (Scientific Reports, 2022) tested 128 model-dataset pairs across healthcare, weather, finance, and transportation, and found 91% of them degrading over time — many with no drift in the underlying data at all. They got old while they were still running.

The expensive version of this is not a model that visibly gets worse. It is a system that gets worse in a direction nobody is watching, while every surface anyone looks at stays green. ML LABS met that system inside a hedge fund: unnecessary and polluted data accumulating at scale without being noticed, while everything that was supposed to run ran.

A data engineer on a headset works at her laptop in a home office above the city
The cost grew while every screen stayed green.

The Data Nobody Was Reading

The fund was storing large volumes of unnecessary and polluted data without noticing. Three things were wrong at once, and not one of them was an error. The aggregation feeding the models was discarding information that mattered for training. The storage structure duplicated data, so the fund was paying to keep several copies of records it did not need. Nothing in the system objected: a pipeline that duplicates still returns rows, and an aggregation that throws away variance still produces a readable table.

Correcting the foundation cut the fund's storage costs by more than 60%, improved the detail and accuracy retained for training, and made their models perform 2% better. For a fund whose edge is the model, two percent is not a rounding error — it is the thing the modeling effort exists to produce, and it came from the storage layer, not from the models.

"We came to Omar with messy data and a deadline we'd already missed once. He turned it into a working system in days, kept us updated throughout, and it held up in production, reliable in a way this field rarely is." — Head of Data, hedge fund

The engagement began as a deadline, not a data-quality project. A loss with no error runs until someone looks, because an alert ends a loss and there was no alert.

Decay Is The Default Direction

Left alone, a live AI system moves in three directions at once, and all three are quiet. Spend creeps wherever usage grows and deletion is nobody's budget line. Quality ages whether or not the world changes. And context leaves, because the knowledge of why the pipeline drops that field lives inside a person, and people move on. The third is the compounding loss, and no dashboard shows it. Nothing pages anyone: alerting fires on symptoms someone named in advance (Google, 2016), and a decay with no error has no symptom to name.

Flowchart: with no owner, a live AI system drifts three ways: spend creeps, quality ages and context leaves. All three lead to nothing alerting. When someone looks, the result is cost cut and the model improved.

One mechanism sits under every live AI surface: the system still runs, so nothing alerts.

What An Owner Actually Does

ML LABS runs an ongoing engineering retainer for HeartSciences, the medical-device company whose cloud ECG backend the practice designed and built. What the ownership has produced is not a maintenance log. It is a set of systems. Two of them are published.

The clinical billing automation built under that retainer ran in shadow mode and was not allowed to cut over until it held 98%+ agreement with expert-adjudicated determinations. The multi-site clinical operations program reaches further out, automating across facilities whose EMR behavior and data paths differ. Neither system is part of the ECG platform build. Both were delivered under the retainer that followed it: the engineer who already holds the context is the one positioned to find the next system worth building.

The contract is ML LABS' own, so it can be stated in full. Three terms carry it:

  • Independent engineering partner — not staff augmentation, not a body against a seat.
  • AI-assisted delivery, acknowledged in writing — one operator leveraged by agents.
  • Month to month after a three-month initial term, on 30 days' written notice

That last clause is the load-bearing one. An arrangement the client can end on short notice must earn each month, and one nobody can end becomes the drift it was hired to prevent.

First Steps

  1. Price what nobody reads. Take one live system and total what it costs to store data that no model, no report, and no auditor has touched since it landed. Storage is where unowned cost accumulates, because nothing in a system ever asks to be deleted.
  2. Trace one aggregation end to end. For a single production model, follow the path from raw capture through to training input, and name every field that the pipeline collapses, drops, or writes twice. The 2% at the hedge fund was in that path.
  3. Name the owner. Write down who answers for each live AI system's cost and quality.

One Owner, Written Targets, Short Notice

Concentrate ownership: one principal engineer accountable for every live AI system, targets written before the work starts, the authority to retire what should not exist, and a contract the buyer can end on short notice. The short-notice clause is what keeps the other three honest, which is why the bar to hold an outside owner to can be set so high.

The hedge fund's storage bill did not fall because of a new model. It fell because someone with authority over the data path read it, and the models improved as a byproduct. That is the position, and it pays only while it is held. Advisory puts an accountable owner on a system we built: monitoring, a monthly written brief, and the engineering that built them a call away, endable on 30 days' written notice if it stops paying for itself.

References

  1. Vela, D., Sharp, A., Zhang, R., Nguyen, T., Hoang, A., and Pianykh, O. S. Temporal Quality Degradation in AI Models. Scientific Reports, 2022.
  2. Google. Monitoring Distributed Systems. Site Reliability Engineering, 2016.

NEXT · TO PRODUCTION

Audit your AI spend.

Two minutes. What to keep, cut, or fix.

15 minutes · no charge · with Omar