New insights by RSS

Immutable Backups and Clean Rooms: Recovering From Ransomware

By Daniel Hirsch · · 9 min read · Filed under Technology, Disaster Recovery
Immutable Backups and Clean Rooms: Recovering From Ransomware

Synchronous replication is very good at one thing: making the second copy look exactly like the first, within milliseconds. When ransomware encrypts a file server at 02:14, the replica at the DR site is encrypted at 02:14 too. The failover runbook that took two years to build will faithfully bring up a synchronized copy of a destroyed environment.

That is the gap between classic disaster recovery and ransomware recovery. DR was designed for losing a place: a flooded data center, a cut fiber run, a building nobody can enter. Ransomware destroys trust in the data and in the systems themselves, including the credentials used to manage them. It usually arrives after the attacker has spent days or weeks inside, often going after the backups first. Recovering from that takes copies the attacker could not touch, a place to restore them that the attacker cannot reach, and an order of operations that starts with identity rather than with the most important application.

Why the DR site does not save you

Three properties of conventional DR work against you in a ransomware event.

Replication is indiscriminate. Storage and hypervisor replication copy blocks, not intentions. Encrypted or deleted data travels as fast as good data. Some replication products keep journals that allow rollback to a point before the damage, which helps, but only if the journal reaches back far enough and the attacker has not tampered with the management plane.

The DR site shares the blast radius. Most DR environments are joined to the same Active Directory, managed from the same backup console, reachable over the same administrative network and patched by the same tools. An attacker holding domain admin rights in production usually holds them at the DR site as well.

Backups are a known target. Attackers commonly hunt for backup servers, delete or encrypt catalogs and repositories, and disable snapshot schedules before launching encryption, because restorable backups weaken their position. If backup administrators log in with domain accounts, a compromised domain means compromised backups.

None of this makes replication useless. It still handles the scenarios it was built for. It means ransomware needs a second, separately designed recovery path, and the plan should state which path applies to which scenario. If you are still deciding how much to spend on that second path, our briefing on why ransomware downtime costs more than the ransom is a useful reality check: the ransom demand is rarely the largest number.

What makes a copy recoverable

Vendors use "immutable," "air-gapped" and "isolated" loosely, sometimes interchangeably. They are different properties, and you want all three.

Immutable

An immutable copy cannot be changed or deleted until its retention period ends, by anyone, including an administrator with full rights. The usual mechanism is write-once-read-many (WORM) storage. In cloud object storage it takes the form of a retention lock: S3 Object Lock in compliance mode, Azure immutable blob storage with a locked time-based retention policy, or a locked retention policy on a Google Cloud Storage bucket. Many backup appliances offer an equivalent.

The distinction that matters is between locks an administrator can override and locks nobody can. S3 Object Lock, for example, has a governance mode, where users with a specific permission can shorten or remove retention, and a compliance mode, where nobody can until the retention date passes. If the attacker becomes that privileged user, a governance-mode lock is a speed bump. Check which mode you are actually running, and who can change the policy.

Isolated

Isolation means the copy is not reachable from production with production credentials. Physical air gaps (tape in a vault, disks rotated offsite) are the classic version and still work, at the cost of slow restores. Logical isolation is the modern compromise: a separate cloud account or tenant, separate identity with no trust relationship to production, separate MFA, management access only from dedicated workstations, and data pulled into the vault rather than pushed from production.

The test is blunt. List every credential that could delete these copies, then ask whether any of them lives in, or can be reached from, the production domain.

Verified

A copy you have never restored is a hope. Verification means routine automated restore tests, integrity checks, and malware scanning of backup content, because backups taken during the attacker's dwell time can contain their tools and persistence mechanisms. You need to know which restore points are clean, not just which ones exist.

Retention: how far back you need to go

Ransomware recovery often means restoring to a point before the attacker got in, not merely before the encryption started. Attackers frequently wait a long time before acting, and the initial access date may stay unknown until forensic work is well along. Short retention, say seven days of dailies, can leave you choosing between a backup that contains the attacker's foothold and no backup at all.

A practical pattern is layered retention: frequent points kept briefly, dailies for 30 days or more, and weeklies or monthlies locked for several months. Locking everything in compliance mode raises storage costs, because nothing can be pruned early. Lock the layers you would actually restore from in a bad week, and budget for them.

Identity comes first

Every recovery runbook ordered purely by business priority runs into the same wall. The payments application will not start because it cannot authenticate. The database is unreachable because DNS is down. Nobody can log in to the hypervisor because the admin accounts live in the compromised directory.

For on-premises Active Directory, a full forest recovery is the usual path after a serious compromise. That means restoring one domain controller per domain from a backup that predates the compromise, in isolation; resetting privileged account passwords; resetting the krbtgt account password twice to invalidate forged Kerberos tickets; removing whatever persistence the attacker added; and only then building out additional domain controllers. Microsoft publishes forest recovery guidance. Follow it and rehearse it, because nobody performs this from memory under pressure.

For Entra ID and other cloud identity providers, you are not restoring the service itself; the provider runs it. You are recovering from what the attacker changed: new credentials on app registrations, added federation trusts, weakened conditional access policies, rogue administrators. Keep an exported, versioned record of that configuration so you can tell what changed, and keep break-glass accounts that do not depend on the compromised path.

NotPetya in 2017 is the standard cautionary tale. Maersk's widely reported recovery depended on a domain controller in an office that happened to be offline when the malware spread. Your plan should not depend on that kind of luck.

The clean room

A clean room, or isolated recovery environment, is where you restore and validate systems before they rejoin anything. It needs its own network with no routes to production, its own identity, its own admin workstations, and tooling to scan what you restore. It can be reserved hardware, a separate cloud account built from code on demand, or a service from your backup vendor. Whatever its form, it has to exist and be tested before the incident. Designing one during an incident costs days you will not have.

Systems leave the clean room for a rebuilt production network only after they pass a gate: scanned, patched, credentials rotated, known indicators of compromise checked, and a business owner confirming the data is usable.

Forensic hold versus speed

The business wants systems back. The incident response team, your insurer and possibly law enforcement want evidence preserved. Wiping and rebuilding destroys the artifacts that show how the attacker got in, and if you never find that out, you may rebuild straight into the same hole.

Settle this trade-off in advance. Agree with legal, the insurer and your retained incident response firm what must be preserved (memory captures, disk images or snapshots of key systems, logs) and who can authorize rebuilding before forensics are complete. Many cyber policies set conditions on which response firms you use and how quickly you must notify the carrier, so read yours now. Snapshotting affected virtual machines before rebuilding them is usually a reasonable middle ground, provided you have set aside the storage.

Rebuild order

A sequence that holds up in most environments:

  1. Contain. Stop the spread, disconnect affected segments, preserve evidence.
  2. Stand up the clean room and an out-of-band communications channel that does not depend on the compromised email or chat platform.
  3. Recover identity: domain controllers or cloud identity configuration, privileged accounts, MFA.
  4. Recover core network services: DNS, DHCP, certificate services, time.
  5. Recover backup and management tooling inside the clean environment.
  6. Restore tier-1 applications and their data stores, validated in the clean room.
  7. Restore tier 2 and below in priority order, with business validation at each step.
  8. Rebuild endpoints, which in a widespread event is often the longest tail.

The tiers come from your recovery targets; our piece on setting RTO and RPO without guesswork covers how to derive them.

Testing restores at scale

Restoring one server proves a backup is readable. Ransomware recovery means restoring hundreds at once, often from a vault with less throughput than the production backup target, through a clean room with limited capacity. Test that scenario. At least once a year, restore a meaningful share of a tier into the clean room in parallel, time it end to end, and compare the result with your RTOs. The bottleneck is frequently storage throughput or a single scanning step that cannot keep pace, not the backup software.

Keep the evidence. Auditors and examiners want proof that backups are protected and restorable, and our guide to surviving a BCM audit or regulatory examination covers how to present it. For an outside reference, CISA's #StopRansomware Guide covers prevention and includes a response checklist that lines up well with the outline below.

A recovery runbook outline

Use this as the skeleton for the ransomware section of your plan; the annotated sample plan table of contents shows where it sits among the other recovery procedures.

Before (review quarterly):

  • Inventory of immutable and isolated copies, with retention, lock mode and every credential that can alter them
  • Clean-room design, capacity and date of last test
  • Identity recovery procedure, break-glass accounts, exported cloud identity configuration
  • Rebuild order by tier, with named owners
  • Agreed forensic preservation rules and who can authorize rebuilding
  • Contacts for the incident response firm, insurer, outside counsel and law enforcement, stored offline

First hours:

  • Contain, preserve, declare
  • Switch to out-of-band communications
  • Notify the insurer and counsel; engage incident response
  • Identify candidate clean restore points

First days:

  • Activate the clean room
  • Recover identity, then core services
  • Restore and validate tier 1, tracking achieved against target RTO
  • Start regulatory notification and disclosure assessments

Afterward:

  • Close the initial access path before reconnecting anything
  • Rotate every credential, service accounts included
  • Write down what failed, what was slow, and what the next test will cover
Immutable Backups and Clean Rooms: Recovering From Ransomware | CPE World