Skip to content
All writing

Network outage prediction for a telecom operator

Forecasting network failures was the straightforward part. Getting a network operations centre to act on a probability at two in the morning took longer and mattered more.

Intelligent systems3 min read

A national telecommunications operator we worked with already held everything needed to see its own outages coming. Alarms, performance counters, temperature readings, configuration change history, years of ticket records. None of it was being used to predict anything. Outages were diagnosed after the fact, by which point the customers had already found out.

Building models that forecast degradation on that data was the straightforward part of the programme. Getting a network operations centre to act on a forecast took considerably longer and determined whether any of it was worth doing.

A probability is not an instruction

An engineer at two in the morning, with a change freeze in force and a shift handover in three hours, cannot do anything with the sentence there is a 68% chance this element degrades within 48 hours. It is true and it is useless, because it does not say what to do, who does it, or what happens if the engineer is wrong.

So we worked backwards from the intervention. For each failure mode the model could see, we asked what the fix is, what it costs, how long it takes and what the outage would have cost instead. The alerting threshold then falls out of that arithmetic rather than out of an F1 score. On some failure modes the intervention is cheap enough that a weak signal justifies acting. On others, the fix is a truck and two engineers, and the model has to be very sure before anyone is dispatched.

Precision buys trust, recall spends it

An operations team will forgive a missed prediction, because missed failures are what they already live with. Three false alarms in a week is different. Once the alerts are widely regarded as noise, they get filtered, and no subsequent improvement to the model recovers the attention.

The first release was deliberately conservative: a high threshold, a small number of alerts, most of them correct. Coverage widened later, once the team had reason to believe the system. Running that order in reverse, launching with broad recall and promising the false positives will settle down, is a reliable way to kill a good model in its first month.

Route it into the workflow that already exists

The obvious deliverable is a dashboard, and nobody watches a new dashboard after week three. Predictions went into the existing ticket queue instead, with the same severity language, the same routing rules and the same closure codes as every other ticket, distinguishable only by a flag showing they were raised by a forecast. The engineer's day did not change shape. One more class of ticket appeared in a queue they were already working.

We had learned that on the solar generation work for iConcept PowerTech in Indonesia, where an auto-ticketing layer raises and routes critical plant issues without a human involved. A prediction that has to be noticed by somebody is a prediction that will sometimes not be.

The first weeks were ignored

Engineers closed predictive tickets without acting on them. This was not obstruction. Nobody had told them what they were permitted to do on the strength of a forecast, and under a change freeze the safe action is to close the ticket and record that the element was still working.

Two things fixed it. A standing authority note, agreed with operations management, setting out what an engineer may do on a predictive ticket without raising a change request. And a weekly review where correct predictions were shown next to the ones engineers had marked wrong, with the reason attached. The reasons were the valuable output. Several described conditions the model had no visibility of at all, and became features in the next iteration.

40%

Reduction in network outages at a national telecom operator

The rollout was phased, with adoption training at each stage. That line reads as filler in a proposal and it is the difference between a model in production and a model in a report. The reduction belongs at least as much to an operations team that agreed to change how it worked as it does to anything we trained.

Talk to our engineering team

Tell us what you need built, modernised or maintained. We will tell you whether we are the right firm for it and what it costs.