A hedge fund cut its storage costs by more than 60% and made its models perform 2% better. Neither result came from a new model. Both came from correcting the data foundation underneath the models: what the fund stored, how it stored it, and what reached training. Nothing there had raised an error, so nobody had ever looked.

Storage Cut Over 60%, Models 2% Better
The fund was storing large volumes of unnecessary and polluted data. Correcting that cut its storage costs by more than 60%, improved the detail and accuracy retained for training, and made its models perform 2% better. For a fund whose edge is the model, two percent is the result the whole modeling effort exists to produce, and it came from the storage layer. The fund's Head of Data described the engagement in their own words.
"We came to Omar with messy data and a deadline we'd already missed once. He turned it into a working system in days, kept us updated throughout, and it held up in production, reliable in a way this field rarely is."
The engagement began as a missed deadline, not as a data-quality project. The savings were found because the person building the system read the path the data took.
Silent Losses Found Before Any Alarm
Three things were wrong at once, and not one of them was an error. The aggregation feeding the models was discarding information that mattered for training. The storage structure duplicated data, so the fund paid to keep several copies of records it did not need. And unneeded, polluted data kept piling up because deleting it was nobody's job.
None of these failures trips a check built to look for errors. A pipeline that duplicates still returns rows. A roll-up that throws away variance still produces a table every dashboard reads happily. A duplicated entity is counted twice, and the evaluation inherits the leak, so a test score can end up measuring the wrong thing with great precision.

Training Data That Keeps Its Detail
The aggregation loss is the one that decides the model, because it cannot be undone downstream. Once the raw detail is gone from the retained record, the ceiling is set upstream, and no amount of tuning will raise it. Correcting the path, rather than the model, is what returned the lost detail to training, which is where the 2% came from.
The cheapest data to store is the data you never collected. The most expensive data is the detail you deleted before you ever knew you needed it.
The method runs from the model backwards. Trace every field a model reads back to the source that produces it, and treat every step that aggregates, deduplicates, rounds or drops a column as a decision to defend or reverse on purpose. What survives that review is what is worth paying to store: one copy, kept at the full detail the models actually need.
Gains Held Because Someone Owns Them
Models do not hold still once they ship. Research on temporal degradation in deployed models (Scientific Reports, 2022) tested 128 model-dataset pairs across healthcare, weather, finance and transportation, and found 91% of them degrading over time, many with no drift in the underlying data at all. They got old while they were still running.
Left alone, a live system drifts three ways at once, and all three are quiet. Spend creeps wherever usage grows and deletion is nobody's budget line. Quality ages whether or not the world changes. Context leaves when the person who knew why the pipeline drops a field moves on. Nothing pages anyone, because alerting fires on symptoms someone named in advance (Google, 2016), and a decay with no error has no symptom to name at all.
That is the loop the fund broke. A loss with no error runs until someone looks, because an alert is what ends a loss, and at the fund there was no alert to end it.
Every Stored Dataset Earns Its Cost
The same test sorts any store. Analysis of AI investment returns (Deloitte, 2024) finds spend on AI rising while returns stay hard to show, and storage is where unowned spend hides best: nothing in a system ever asks to be deleted. Placing each dataset by how often it is read and how much it moves the model turns the bill into decisions.
A dataset nobody reads and no model needs is deleted. Copies that are read but add nothing are consolidated into one. Detail that matters but never reaches training goes back into the path, so the models can learn from all the detail they actually need.
Three Steps To Find The Same Leak
- Price what nobody reads. Total what one live system pays each month to store data that no model, report or auditor has touched since the day it landed.
- Trace one aggregation end to end. Follow a production model from raw capture to training input, and name every field the path collapses, drops or writes twice.
- Name the owner. Write down who answers for each live system's running cost and model quality, with the authority to retire any data that should not exist.
One Owner Keeps The Savings Coming
The fund's storage bill did not fall because of a new model. It fell because someone with authority over the data path read it, and the fund's models improved as a direct result. That gain holds only for as long as someone keeps reading the data path.
Advisory puts one accountable owner on a system we built: cost and drift reviewed as standing items, a monthly written brief, and the engineering that built it a call away, ending on 30 days' notice if it stops paying for itself. If your data path is the open question, the first call is where to prove it on your own sources, before any build is budgeted.
References
- Vela, D., Sharp, A., Zhang, R., Nguyen, T., Hoang, A., and Pianykh, O. S. Temporal Quality Degradation in AI Models. Scientific Reports, 2022.
- Google. Monitoring Distributed Systems. Site Reliability Engineering, 2016.
- Deloitte. AI ROI: The Paradox of Rising Investment and Elusive Returns. Deloitte Insights, 2024.



