Telstra outage caused by failure to address well-known network issues
An independent investigation into the Telstra outage in July found the incident was caused by a faulty power supply being replaced, then reset to the wrong date.
Alarms that would normally alert staff to an error were not monitored outside business hours and there was a lack of knowledgeable staff on duty.
Telstra failed to prioritise fixing known vulnerabilities in its systems despite being warned about them months before they triggered a major nationwide outage.
An independent report investigating the July 8 outage, which affected about 45 per cent of all calls and data sessions on Telstra's mobile network and saw hundreds of Triple Zero (000) calls unable to connect, was released by the telco yesterday.
The review by Technology Audit Partners (TAP) found that a lack of oversight and ownership of key elements of the network contributed to the outage, including that two of the telco's key engineers were on mandatory leave at the time it occurred.
It found that the outage began at 2:50am — nearly an hour earlier than Telstra had previously disclosed — and was triggered after a faulty power supply was replaced that then reset a GPS card date to 2006, which then flowed through the rest of the mobile network.
The review said the alarms that would usually alert staff to the error were only monitored during business hours by a "limited number" of people, and a lack of "knowledgeable staff" contributed to the telco taking "several hours to determine the root cause of the outage".
"There were several reasons for this, primarily due to insufficient clarity of responsibility regarding NTP (Network Time Protocol) ownership and support, lack of visibility to the planned event to replace the Melbourne chassis that night, as well as a lack of understanding of the alarms that were the earliest indicators of a problem in the network," it said.
The report also noted that the telco did not treat its timing system as a priority, nor did it treat it as a "high risk function" of the network.
Telstra chief executive Vicki Brady said the company accepts the findings of the report and is working through its recommendations.
"The cause … was consistent with what we said at the time, so it was not due to infrastructure failing, it was due to an undocumented design change in October, and it was due to a software update having not been done on a GPS card," she told a media briefing.
"The overarching finding is that we did not prioritise our timing system inside our networks at the highest level as a critical capability, which is exactly what it should have been.
The company has already reallocated resourcing to the team to improve its capability, Ms Brady added.
The report's findings come after the ABC revealed in the days after the July outage that Telstra had been warned for months by federal government agencies and academics that it was vulnerable to the type of error it experienced.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.abc.net.au — the content belongs to ABC News Australia - Politics.