New insights by RSS

Cloud Region Outages: Designing Disaster Recovery Across Providers

By Daniel Hirsch · · 9 min read · Filed under Disaster Recovery, Technology
Cloud Region Outages: Designing Disaster Recovery Across Providers

A cloud provider's service level agreement is a refund policy, not a recovery plan. If a region goes down for six hours, the SLA may return a slice of that month's bill for the affected services. It will not process your payments, answer your customers or explain the gap to your regulator.

That is the practical meaning of shared responsibility for resilience. The provider is responsible for the resilience of its infrastructure: data centers, power, hardware, the network between availability zones, and the managed services it runs. You are responsible for how your workloads use that infrastructure. Put everything in one zone, and the provider has kept its side of the bargain when that zone fails. Put everything in one region, and the same applies to the region. AWS, Microsoft Azure and Google Cloud all spell this out in their architecture guidance. Plenty of organizations still discover it mid-outage.

Know which failure you are designing for

"The cloud went down" covers several very different events, each with its own scope, likelihood and fix.

A single instance or node. Routine. Handled by auto-scaling, managed instance groups, health checks and stateless design. If this kind of failure causes you an outage, the problem is architecture, not DR.

An availability zone. A zone is one or more data centers with independent power, cooling and networking, close enough to its sibling zones for low-latency synchronous replication. Zone failures do happen: power events, cooling trouble, network faults. Spreading workloads across at least two zones, ideally three, with managed databases set for multi-zone failover, absorbs most of them without anyone declaring a disaster. That is high availability, and it should be the default for anything above your lowest tier.

A region. A regional event, where several zones or a region-wide service degrade together, is rarer but has a much larger blast radius. Zones within a region share some regional services and control planes, so multi-zone design will not save you. Surviving it requires a second region, which is where disaster recovery in the classic sense begins.

A global service or control plane. Some provider services are global by design (identity and access management, DNS, content delivery, some management APIs), and some keep their control plane in a single region even when the data plane runs everywhere. A problem there can touch many regions at once. Provider documentation says which services work this way; read it with that question in mind.

Something that is not the cloud at all. Your own deployment pipeline, a bad configuration change, an expired certificate, a third-party agent installed on every server. These are more common than region failures, and multi-region design does little against them, because the faulty change propagates everywhere you deploy.

Control-plane dependencies are the quiet killer

Most cloud DR plans assume you will be able to do things during the outage: launch instances in the recovery region, scale out, change DNS records, log in to the console. Each of those is a control-plane action. Control planes tend to struggle first in a regional event, either because they are directly affected or because thousands of other customers are attempting the same failover at the same moment.

The design principle, which AWS calls static stability, is to make recovery depend on the data plane wherever possible. In practice:

  • Pre-provision what you can. Capacity already running in the recovery region does not depend on a launch API that may be throttled or failing. Capacity reservations exist for exactly this, at a price.
  • Pre-configure DNS failover. Use health-check-driven failover records set up in advance instead of planning to edit records during the event. Keep TTLs short enough to matter.
  • Do not depend on the console. Keep tested command-line and infrastructure-as-code paths, with credentials that work if the console or the primary single sign-on route is unavailable.
  • Keep identity out of the failed region. If workforce identity, break-glass accounts or the secrets store live only in the primary region, failover stalls at the login prompt.
  • Check where your tooling runs. CI/CD, monitoring, the runbook wiki and the incident chat channel are all dependencies. A recovery plan stored in a tool hosted in the failed region is not much of a plan.

19 July 2024: a reminder from outside the cloud

The CrowdStrike Falcon outage belongs in every cloud DR discussion precisely because it was not a cloud provider failure. On 19 July 2024, a faulty content update for the Falcon sensor on Windows caused affected machines to crash and fail to restart. Microsoft estimated that roughly 8.5 million Windows devices were affected. Airlines, hospitals, banks and broadcasters were disrupted, some of them for days.

Three lessons carry straight over to cloud DR design:

  1. Correlated failure beats geographic redundancy. Windows servers in a second region, or a second cloud, running the same agent received the same update. Distance does not help when the failure ships with your own software stack.
  2. Remediation needed hands, not APIs. Many machines needed someone to boot into a recovery environment and remove the faulty file. For cloud virtual machines that often meant detaching disks and repairing them from another instance, or restoring from snapshots. Teams with no automation for this did it server by server.
  3. Recovery dependencies sat inside the blast radius. Machines protected by BitLocker disk encryption needed recovery keys. An organization whose key management ran on servers that had also crashed had a problem nested inside its problem.

So the "region outage" scenario should sit alongside scenarios for a bad agent update, a bad internal change and ransomware, which ignores region boundaries entirely. Our briefing on why ransomware downtime costs more than the ransom covers that last case.

Multi-region patterns compared

The standard patterns differ mainly in how much runs in the recovery region before anything goes wrong. The ranges below are typical, not promises; your achieved numbers come from testing.

Pattern What runs in the recovery region Relative cost Typical RTO Typical RPO Main risk
Backup and restore Nothing; backups and images copied cross-region Low (storage and transfer) Hours to a day or more Hours (backup interval) Untested infrastructure code; scarce capacity when everyone fails over
Pilot light Data stores replicated live; compute defined in code but off or scaled to zero Low to moderate Tens of minutes to hours Seconds to minutes Scale-up relies on control-plane APIs during the event
Warm standby Full stack running at reduced capacity Moderate to high Minutes Seconds to minutes Configuration drift; scaling under real load
Active-active Full capacity serving live traffic in two or more regions Highest (often more than double, plus engineering) Near zero to minutes Near zero to seconds, depending on data design Consistency and conflict handling; complexity causing its own incidents

Two practical notes. First, mix patterns by tier rather than picking one for the whole estate; our guide to setting RTO and RPO without guesswork shows how to set the tiers. Second, active-active is a data problem far more than a compute problem. Running stateless web tiers in two regions is easy. Deciding which region wins when the same customer record changes in both is hard, and that decision belongs in the application design, not with the infrastructure team.

Multi-cloud: the honest trade-offs

Multi-cloud DR comes up in every board conversation about concentration risk, and EU financial firms face the question directly under DORA, which has applied since 17 January 2025 and expects firms to assess ICT third-party concentration. It deserves a straight answer.

What a second provider protects you against: a provider-wide failure of a global service, a control-plane problem spanning regions, a commercial or legal dispute that cuts off access, and the supervisory question "what happens if this provider is unavailable?"

What it costs: two sets of skills, tooling and security controls; designs reduced to the features both providers share; egress charges for keeping the second copy current; identity federation across clouds; and a second environment that, unless exercised regularly, drifts into something nobody trusts on the day it is needed.

The middle path most organizations settle on is targeted multi-cloud. Keep backups or a copy of critical data with a second provider. Run DNS with a secondary provider. Keep out-of-band communications on a different platform. Maintain a minimal "lifeboat" for a small set of tier-0 functions rather than a full mirror. That addresses the worst concentration risks for a fraction of the cost.

SaaS dependencies you did not choose

Your recovery design almost certainly depends on software you do not run: workforce identity, email and collaboration, ticketing, payroll, the mass notification system, the status page. Each runs on someone's cloud, often a single provider, sometimes a single region. Ask vendors for their DR architecture, their RTO and RPO commitments, evidence of a recent failover test, and how you would export your data. Put the answers in the plan, and note where two vendors you treat as independent actually share an underlying provider.

Game days

A failover you have never run is a theory. A game day is a scheduled, controlled exercise that breaks something on purpose and measures what happens. AWS and Azure both offer managed fault-injection services, and a lot can be done with network rules and access policies alone. Useful scenarios:

  • Block traffic to the primary region for one application and fail over by the runbook
  • Withdraw console access and recover using only command-line and infrastructure-as-code paths
  • Make the primary identity provider unavailable and test break-glass access
  • Fail the primary database and measure actual data loss against the RPO
  • Push a bad configuration change and practice rolling it back across regions

Set abort criteria and a blast radius before you start, have business owners confirm usability, and record achieved RTO and RPO. That record is what examiners ask for; our guide to surviving a BCM audit or regulatory examination covers how to present it, and the annotated plan table of contents shows where cloud failover procedures belong in the plan.

The first 30 days

Week 1. List tier-0 and tier-1 workloads with their current region and zone layout. Flag anything running in a single zone.

Week 2. Map control-plane and global-service dependencies for each, plus the SaaS tools the recovery process itself relies on.

Week 3. Choose a pattern per tier from the table above and price the gap between what exists and what the pattern needs.

Week 4. Run one game day on one tier-1 application. Record achieved RTO and RPO, write down every manual step, and book the next one before the room empties.

Cloud Region Outages: Designing Disaster Recovery Across Providers | CPE World