On November 9, 1965, one incorrectly set protective relay helped plunge much of the Northeastern United States and parts of Canada into darkness. Within seconds, overloaded transmission lines began tripping one after another. The resulting blackout affected 30 million people and lasted up to 13 hours in some places. The relay was small, but the system surrounding it was enormous. That distinction feels particularly relevant this week.
Last Tuesday, a technical fault at Britain’s air traffic control provider NATS disrupted one of the world’s busiest aviation networks. About 1,300 flights to and from U.K. airports were canceled that day, with the mess spilling into Wednesday as airlines tried to get aircraft, crews, and passengers back where they belonged. NATS CEO Martin Rolfe ruled out a cyberattack. This was, apparently, just a technical problem.
“Just” is doing some heavy lifting there.
The software problem may have lived inside an air traffic control system, but its consequences certainly did not. Travelers missed flights, airlines lost capacity, airports backed up, and aircraft ended up in the wrong places. A computer problem became a transportation problem, a staffing problem, a customer service problem, and an operational problem.
And aviation wasn’t alone.
Last Thursday, September 3, ChatGPT, Claude, and Grok experienced overlapping outages. The incidents were not confirmed to share a common cause, which is an important distinction. But the overlap still exposed something worth paying attention to: AI is no longer just a tool people experiment with when they have a spare minute. Organizations are building it into research, customer support, writing, analysis, software development, and internal workflows.
When those services go down, the impact is not always dramatic. Sometimes it’s just a stalled task or a frustrated employee. But across a large organization, small interruptions can cascade quickly. When three different doors temporarily refuse to open, having three doors suddenly feels less reassuring.
Then came AT&T. Internet disruptions were reported across Texas last Monday, with significant problems around Houston and Dallas. For many customers, connectivity is so embedded in the workday that it barely registers until it disappears. Employees cannot reliably access systems, contact customers, join meetings, or receive updates. Field teams may struggle to coordinate, while managers may not even know whether someone is offline because of the outage or unreachable for another reason.
Once again, something largely invisible until it stopped working suddenly became very visible. And once a communications system fails, the organization has to solve two problems at once: the original disruption and the challenge of telling people what is happening.
Why you should care: Organizations have gotten remarkably good at building operations around systems they rarely have to think about—and dependencies they may not fully appreciate.
Until something fails.
Business continuity plans tend to focus on the thing that failed. The harder question is what that failure takes down with it. If internet access disappears, can employees still receive critical information? If flights stop moving, do you know which travelers are affected? If a service your teams use every day goes dark, who needs to know and what happens next?
Resilience isn’t keeping everything from breaking. That’s a pretty ambitious Tuesday.
It’s knowing what you’ll do when something inevitably does. |