03/03/2020 16:50
Dear Customer,
We are currently experiencing a network outage with one of our SIP carriers.
This is currently been investigated by the network team and we will update you as soon as we have an update from the network team.
Kind Regards
Network Services Team
———————————————————————————————–
03/03/2020 16:58
Dear Customer,
We have been made aware from our SIP supplier that they have had an outage in one of there Data Centre’s.
All traffic has now been re routed over the 2 backup Data Centre’s and we are now showing calls have returned to normal.
As soon as we have a reason for outage we will distribute this via email.
Apologies for the inconvenience this has caused.
The Network services Team
———————————————————————————————
04/03/2020 16:45
Reason for Outage 3rd March 2020
Incident Summary and Background
On Tuesday the 3rd of March at 16:18 Our Supplier experienced a failure of one of its database servers that
caused instability and in some cases a failure of services which were transiting through the THN
datacentre. This would have resulted in the inability for customers serviced directly from THN to register
and make and receive calls, other customers homed to different data centres may have experienced a
delay in making or receiving calls. After diagnosis of the fault and implementation of a fix services
started to recover at 16:40.
Timeline
Time Action
16:18 Alerts received that connectivity to one of the master databases is experiencing
problems.
16:19 Core engineers being investigation to the cause, suspect a networking failure.
16:25 Network engineers begin manual checks no alarms raised
16:30 Core network equipment and links given all clear.
16:33 Further testing indicates a local network stack failure on host database host.
16:35 Configuration updated to point at alternate master.
16:40 Services confirmed recovering.
16:50 Call and registration volumes back to normal levels
Root Cause
The route cause for this outage was the failure of one of the master databases to process incoming
requests. Our Supplier has built in resiliency into its database architecture which in this case limited the
scale of the outage however some customers were still critically affected. The failure was only apparent
on one network interface which the database was served across with other interfaces correctly
connecting and continuing to function normally. Once the affected interface was restarted normal
operation resumed. This indicates that the problem was likely within the IP stack on the affected
interface.
Risk Mitigation
In the 15 years that Our Supplier have been running there infrastructure we have not seen a database failure
happen in this way. While they test the failover and resiliency mechanisms there is always a possibility
that the automated monitoring and recovery systems might miss a fault. In this case the engineers had
to manually instigate the failover to the alternate master for services to recover. This process took 22
minutes from the initial alert of a problem to the systems starting to recover. They will be analysing logs
and updating their detection scripts and thresholds to try to automatically detect this type of failure in
future. While they have not found a specific bug related to the failure they saw, though their BAU
maintenance are already in the process of migrating to a newer version of the databases software and
operating systems on which they run.
Kind Regard
The Network Services Team