New insights by RSS

Disasters That Shaped Business Continuity Planning: From Mainframes to the Cloud

Reference briefing by Maureen Callahan · Last reviewed · 10 min read
Disasters That Shaped Business Continuity Planning: From Mainframes to the Cloud

Most continuity plans are a record of the disasters their authors feared. A plan written in 1990 assumed the computer room would burn. A plan written in 2002 assumed the building would be lost. Few plans written in 2019 assumed every office would close at once, for months, while the systems ran fine. Each major event exposed an assumption, and the discipline changed in response. That history is a quick way to find the assumptions hiding in your own plan.

The timeline at a glance

Year Event What it taught planners
1970s–80s Growth of commercial hot sites Recovery became a service, but only for the data center
1988 Hinsdale, Illinois telephone central office fire One telecom building can fail a whole region
1992 Chicago tunnel flood; Hurricane Andrew Losing access equals losing equipment; staff homes matter
1993 World Trade Center bombing Plans need evacuation, staff accounting and workspace
1995 Oklahoma City bombing Security, rescue and family support belong in planning
1996 TWA Flight 800 Family and press communication is a discipline
1999–2000 Y2K Dependency mapping works, then goes stale
2001 September 11 attacks Backup sites need separate regions, staff and infrastructure
2003 Northeast blackout Untested generators, fuel and phones fail
2005 Hurricane Katrina Disruption can last months and scatter staff
2009 H1N1 pandemic Absenteeism and telework assumptions need testing
2012 Superstorm Sandy Keep critical equipment out of flood-zone basements
2017 NotPetya A cyberattack can be a full continuity event
2020 COVID-19 Plan for losing everything normal, everywhere, for years
2024 CrowdStrike Falcon outage One trusted vendor's update can halt the world

Before "business continuity": the data-center years

In the 1970s and 1980s the fear was simple: lose the mainframe and you lose the business. Disaster recovery meant backup tapes, an off-site vault and a contract for somewhere to run them. Commercial hot sites, equipped computer rooms that subscribers could occupy after declaring a disaster, grew into an industry led by providers such as SunGard and Comdisco. The vocabulary of hot, warm and cold sites dates from this period.

The weakness was scope: plans covered the computer, rarely the people, phones or buildings around it.

Hinsdale, 1988

On May 8, 1988, a fire at Illinois Bell's central office in Hinsdale, a western suburb of Chicago, knocked out telephone and data service across a large part of the region, and full restoration took weeks. Businesses with untouched computers still could not reach customers, authorize card transactions or talk to branches. The lesson was dependency: one telecom building can be a single point of failure, and "diverse" circuits from different carriers may pass through the same office. Route audits entered the planner's checklist, alongside priority-restoration programs such as Telecommunications Service Priority and carrier mutual-aid arrangements.

1992 to 1996: the building becomes the problem

A run of events moved attention from computers to places and people.

April 1992, Chicago. River water broke into the old freight tunnels under the Loop and flooded basements across downtown. Power was cut and buildings stayed closed for days, so equipment on dry upper floors was useless.

August 1992, Hurricane Andrew. Andrew struck south Miami-Dade County as a Category 5 storm. Businesses learned that a regional disaster damages employees' homes as well as offices, and staff who have lost their houses do not come to work. Insurers took losses heavy enough that several failed, and criticism of the federal response fed into FEMA's reorganization over the following years.

February 1993, World Trade Center. A truck bomb in the parking garage beneath the North Tower killed six people and injured more than a thousand. Power and building systems failed, evacuation down smoke-filled stairwells took hours, and tenants were shut out for weeks. Many firms declared disasters with their recovery vendors, then found their plans said nothing about evacuation, accounting for staff, or where people would sit. Terrorism became a planning assumption.

April 1995, Oklahoma City. The bombing of the Alfred P. Murrah Federal Building killed 168 people. FEMA's urban search and rescue task forces, then a young system, carried out a long and dangerous collapse operation. The aftermath produced federal building security standards through the newly created Interagency Security Committee, and it showed that family assistance, employee counseling and continuity of government operations are part of recovery. Responses like this depend on contractors, engineers and utilities working beside agencies, the subject of our briefing on public–private partnerships in emergency management.

July 1996, TWA Flight 800. The crash off Long Island killed all 230 people aboard, and the handling of victims' families drew sharp criticism. Congress passed the Aviation Disaster Family Assistance Act that year. For corporate planners, it made communication with families and the press a discipline rather than an improvisation.

Y2K: the project nobody thanks

The millennium date problem forced organizations to inventory every system, interface and supplier, test fixes, and staff command centers through New Year's Eve 1999. Little failed, so some called it an overreaction. Practitioners concluded instead that inventory and dependency mapping work, and decay quickly once the deadline passes.

September 11 and the end of the nearby backup site

The attacks of September 11, 2001 destroyed the World Trade Center complex and damaged telecommunications across lower Manhattan, including the heavily used Verizon switching building on West Street. Some firms found their backup sites nearby and just as unreachable; others had working systems but had lost the people who ran them. US stock markets stayed closed until Monday, September 17.

The response reshaped the discipline. In 2003 the Federal Reserve, the OCC and the SEC issued an interagency paper on sound practices to strengthen the resilience of the US financial system. It expected core clearing and settlement organizations to recover within the business day, with a two-hour goal, and pushed firms toward backup sites that did not share the primary site's labor pool, transportation, telecommunications or power. Geographic dispersion, split operations and employee accounting became standard plan elements. If a plan's structure still reflects the pre-2001 data-center view, compare it with an annotated table of contents for a modern continuity plan.

2003: the Northeast blackout

On August 14, 2003, a cascading grid failure that began in Ohio cut power to about 50 million people across the Northeast, the Midwest and Ontario. Generators that had never run under full load failed, fuel contracts proved thin, cell networks congested, and in New York people walked home across the bridges. Cleveland lost water pressure when its pumps stopped. It led to mandatory grid reliability standards under the Energy Policy Act of 2005, and taught planners to test generators under real load and treat fuel as a supply chain.

2005: Hurricane Katrina

Katrina made landfall on August 29, 2005, and levee failures flooded most of New Orleans. More than 1,800 people died. For businesses the defining features were duration and dispersal: employees evacuated to dozens of states, payroll had to reach people with no fixed address, paper records were destroyed, and recovery took months or years. Congress restructured FEMA through the Post-Katrina Emergency Management Reform Act of 2006. Planners added long-duration scenarios, out-of-region alternates and ways to reach scattered staff.

2009: H1N1

In the mid-2000s, organizations wrote pandemic plans around H5N1 avian influenza. When the World Health Organization declared an H1N1 pandemic in June 2009, those plans met a milder virus and a mismatch: triggers keyed to WHO phases, untested telework capacity, and absenteeism estimates that ignored school closures. How those plans fared a decade later is covered in our briefing on pandemic and infectious-disease continuity planning.

2012: Superstorm Sandy

Sandy came ashore on October 29, 2012. Its surge flooded lower Manhattan, subway and road tunnels, and much of the New Jersey shore. The New York Stock Exchange closed for two days for weather, the first such closure since the blizzard of 1888. Buildings with generators on the roof and fuel pumps or switchgear in the basement lost both. Hospitals evacuated patients when basement systems flooded, and one downtown data center survived on diesel carried upstairs by hand. Sandy made flood elevation a continuity question rather than a facilities footnote.

2017: NotPetya

On June 27, 2017, malware disguised as ransomware but built to destroy data spread from a compromised update to Ukrainian accounting software into global corporate networks. Maersk, Merck, FedEx's TNT unit and Mondelez were among those hit. Maersk rebuilt its network around a surviving copy of its domain controller data found in its Ghana office, which had been offline because of a local power outage. The US and UK later attributed it to Russia's military. NotPetya taught that a cyberattack can be a full continuity event, that identity systems need their own recovery plan, and that insurance war exclusions deserve a read before any claim.

2020: COVID-19

The WHO characterized COVID-19 as a pandemic on March 11, 2020. Most plans assumed the loss of one site while the world carried on. Instead every site closed at once, supply chains broke in unexpected places, and the disruption lasted years. The lasting change was a move from scenario-based plans toward impact-based ones: what happens if we lose people, premises, technology or suppliers, whatever the cause.

2024: the CrowdStrike outage

On July 19, 2024, a faulty content update to CrowdStrike's Falcon sensor crashed Windows machines around the world; Microsoft estimated that about 8.5 million devices were affected. Airlines, hospitals, banks and broadcasters stalled. The fix was known within hours, but many machines needed hands-on repair, and organizations that could not find their BitLocker recovery keys waited longer. Delta needed days and thousands of cancellations to recover. The lesson was concentration: a trusted vendor with deep access to every endpoint is a shared point of failure.

Ten questions history would ask of your plan

  1. If our main telecom carrier or central office fails, what still works? (Hinsdale)
  2. If we cannot enter the building for a week, where do people go? (Chicago 1992, WTC 1993)
  3. If employees' homes are damaged, who can actually work? (Andrew, Katrina)
  4. Can we account for every employee within an hour? (1993, 2001)
  5. Does our alternate site share our power grid, transit or labor pool? (2001)
  6. Have the generators run under full load this year, and who delivers fuel? (2003, Sandy)
  7. Is anything critical sitting below flood level? (Sandy)
  8. Can we rebuild identity and directory services from a clean copy? (NotPetya)
  9. Could we operate with most staff remote for six months? (H1N1, COVID-19)
  10. Which single vendor update could stop every laptop we own? (CrowdStrike)

Frequently asked questions

When did "business continuity" replace "disaster recovery" as the name of the discipline?

Gradually, through the 1990s. Events like the 1993 World Trade Center bombing showed that a recovered data center is useless without people, workspace and communications, so planning broadened to the whole organization. "Disaster recovery" survives as the name for technology recovery, which is now one part of a continuity program rather than all of it.

Which single event changed continuity planning the most?

For financial services, September 11 had the deepest effect, producing regulatory expectations for geographic dispersion that still shape where firms put backup operations. For organizations generally, COVID-19 changed more assumptions at once: duration, scale, remote work and supply chains.

Why do so many plans still focus on IT?

Because the discipline started there, and technology recovery is easier to measure and fund than people or supplier recovery. The history above argues for balance: most events that hurt businesses badly involved access, people, utilities or suppliers.

How can historical events be used in exercises?

They make good scenario seeds because they are real and well documented. Adapt the event to your own locations and dependencies rather than replaying it literally, and ask what would differ today. A plain blackout often exposes more gaps than a dramatic scenario.

Disasters That Shaped Business Continuity Planning: From Mainframes to the Cloud | CPE World