A top 10 global telecom company got back more than 12 times the cost of its roaming optimization in the first year. The custom ML platform behind that return surfaced optimization opportunities across 128% more corridors than the original scope, running on roughly 1TB of new data every day: live network telemetry, agreement terms, and historical traffic patterns. Omar designed and built it to optimize cost and routing together. The operator is not named here.
The margin at stake had been leaking for structural reasons. International roaming still settles through bilateral agreements, negotiated one counterparty at a time, and the structural complexity of those agreements (OECD, 2013) is documented: each pair of operators sets its own wholesale terms, on its own renewal cadence, in its own currency of concessions. Every agreement is an input to a routing decision, and the decisions multiply with corridors, hours, and traffic classes until the review cycle is slower than the thing it reviews.

Over 12x Returned In Year One
Measured against the engagement's cost, the optimization returned over 12x in the first year. The platform surfaced optimization opportunities across 128% more corridors than originally scoped, inefficiencies the manual analysis had never reached, and cost per session came down on the targeted corridors against the pre-platform baseline.
The same architecture is a substrate that adjacent problems can be built on rather than a point solution: capacity planning, fraud detection on roaming traffic, and quality-of-service forecasting all consume the same data feeds and the same trained models.
Margin Recovered Where Review Lagged
Before the platform, each routing decision sat at the intersection of carrier agreement terms, real-time congestion, regional regulatory constraints, and per-corridor demand forecasts. The decision space was structurally larger than the team could traverse. By the time analysts had modeled one corridor in a spreadsheet, the demand pattern they had modeled against was already a week old. The portfolio of bilateral agreements grew as the operator expanded, so the surface outgrew the review capacity.
The bilateral terms changed on different cadences, so the pricing baseline was never stable long enough to reason about. Traffic was becoming more volatile; the wholesale settlement structure that governs those flows (GSMA, 2024) assumes commercial terms move slowly while traffic moves fast. Every answer had a shrinking shelf life, and margin leaked in the gap, which is where the operational margin in telecom actually sits (McKinsey, 2024): in the thousands of small routing decisions that nobody has the time to revisit.
Three signals tell an operator that this line has already been crossed. Margin per session drifts downward on corridors that were profitable a quarter earlier, hidden at the aggregate level. Quality complaints rise on a small set of paths where the cheapest carrier is still being chosen after congestion has degraded its throughput. Agreement renewals get signed against pricing baselines built from stale traffic. When two of those are visible at once, adding analysts flattens the curve without changing its shape: the point at which network optimization stops being a staffing problem and becomes an AI problem.
Routing That Keeps Pace With Traffic
What ML LABS delivered was a production ML platform with three capabilities, not a research prototype:
- Time-series models. Trained on global roaming traffic to forecast demand and cost across markets, following the hourly, weekly and seasonal cycles in each corridor.
- Cost optimization engine. Identified the lowest-cost routing paths that still held the quality thresholds, instead of collapsing a carrier under the load it created.
- Routing recommendation system. Adapted to shifting traffic patterns and agreement terms as they moved, so its recommendations held over the operating horizon.
What made it operational were the unglamorous parts. A streaming pipeline absorbing terabytes of daily traffic without a backlog. Models retrained on live data, not a frozen snapshot. And a dashboard where network ops could monitor, override, and audit every decision, because an optimizer that cannot be overridden is one nobody will switch on.
Recommendations Stable Over The Horizon
A static optimizer would have been the wrong tool, and the reason is a property of the cost function rather than a preference about models. Demand for a corridor follows hourly, weekly, and seasonal cycles, so the cheapest path at one moment is the congested path at another. Time-series models capture the autocorrelation in traffic and pricing dynamics, a property the survey of deep learning for network traffic prediction (ACM Computing Surveys, 2024) treats as the load-bearing assumption for telecom forecasting, so the routing engine plans against the demand curve instead of the demand snapshot.
A naive shortest-cost solver routes everything into the cheapest carrier until its quality collapses under the load it just created, then oscillates away. The time-series formulation keeps the optimizer aware that its own decisions shape near-future congestion, which is what makes the recommendations stable over an operating horizon.
Narrow Proof Bought Room To Extend
This was never one workflow, and buying it as one workflow is how it fails. Corridor pricing, agreement modeling, telemetry ingestion, and the network-ops decision surface are four workflows that share an architecture. Scope one of them alone and you get a model that is right about cost and wrong about quality, which is worse than the spreadsheet it replaced.
The narrow start makes program scale survivable. Covering the full agreement portfolio in the first deployment would have made every model error look like a system error and stalled adoption with the team whose sign-off it needed. Proving it where the cost signal was unambiguous bought the credibility to extend. That extension only stays cheap if the architecture grows by configuration, not rebuild, a decision made at design time or not at all.
Three Steps To A Measured Baseline
- Time one corridor by hand. Price it end to end and measure how long it took. If the demand pattern shifts faster than that, adding analysts will not close the gap.
- Build against the costliest corridor first. Deploy to a market subset, and measure cost per session against the pre-platform baseline before anyone discusses expansion.
- Name the owner of the shared substrate before it carries a second workflow.
Later Workflows Reuse The Same Base
Several workflows on one architecture, with targets written per workflow before work starts, is what a sequence of builds on it exists to deliver, each kept running by the person who built it. If the decision surface already outgrows the team reviewing it, and the corridors nobody has time to price are where margin leaks, that is the problem this work solves.
References
- Aouedi, O., Le, V. A., Piamrat, K., & Ji, Y. Deep Learning on Network Traffic Prediction: Recent Advances, Analysis, and Future Directions. ACM Computing Surveys, 2024.
- OECD. International Mobile Roaming Agreements. OECD Digital Economy Papers, 2013.
- GSMA. Roaming. GSMA, 2024.
- McKinsey & Company. How Generative AI Could Revitalize Profitability for Telcos. McKinsey & Company, 2024.



