New insights by RSS

RTO vs. RPO: Setting Recovery Targets Without Guesswork

By Daniel Hirsch · · 10 min read · Filed under Disaster Recovery, Business Continuity
RTO vs. RPO: Setting Recovery Targets Without Guesswork

Ask an infrastructure team where its recovery time objectives came from and the honest answer is usually a spreadsheet column someone filled in during a storage refresh. Four hours for the important systems, 24 for everything else, a 15-minute RPO for anything with a database behind it. The numbers look reasonable. Nobody in the business signed them, nobody priced them, and nobody has measured whether the environment can actually hit them.

That is guesswork with a decimal point. The fix is not a better spreadsheet. It is a short chain of reasoning: start with what the business can tolerate, turn that into targets per application and per data store, price the gap, then prove the result in a test.

Four terms, defined properly

The vocabulary gets mangled often enough that it is worth pinning down. The definitions below are broadly consistent with ISO 22301 and with NIST SP 800-34, which says "maximum tolerable downtime" where ISO says "maximum tolerable period of disruption."

Maximum tolerable period of disruption (MTPD), also called maximum tolerable downtime (MTD), is a business number. It is how long a process can be unavailable before the damage becomes unacceptable: a regulatory breach, a contractual default, customers leaving in volume, or harm to people. It belongs to the process owner, not to IT.

Recovery time objective (RTO) is the target time to resume a process or system at a minimum acceptable level after a disruption. For an application, the RTO is the clock from failure (or from declaration, so pick one and write it down) until the system is usable by the people who depend on it.

Work recovery time (WRT) is the time after the system is technically back during which people verify data, re-enter lost transactions, clear queues and reconcile. Most plans forget it. A platform that comes back at hour three but needs two hours of reconciliation before anyone trusts the numbers has not recovered at hour three.

Recovery point objective (RPO) is the maximum acceptable data loss, expressed as time. An RPO of 15 minutes means that after recovery you can tolerate losing, at most, the last 15 minutes of changes. RPO looks backward from the moment of failure; RTO looks forward.

The relationship that matters: RTO plus WRT must be less than MTPD, with margin. If the business can survive eight hours without a process, a six-hour system RTO followed by three hours of reconciliation is already a failure on paper.

Start from the BIA, not the server list

Targets derived from the infrastructure describe what the infrastructure happens to do today. Targets derived from the business impact analysis describe what the business needs.

A usable BIA gives you, per process: impact over time, peak periods, the MTPD, and the applications and data the process depends on. If yours does not, fix that first. Our guide to running a business impact analysis people actually finish covers how to get those answers out of busy process owners.

With the BIA in hand, translate in this order:

  1. Set the process MTPD. Use the impact-over-time curve. The MTPD is where the curve crosses a threshold the executive sponsor agrees is intolerable, not where it first starts to rise.
  2. Subtract WRT and a margin. Ask the process owner how long it takes to verify and catch up once a system returns. Hold back time for detection, declaration and surprises. What remains is the ceiling for every application on that process's critical path.
  3. Give each application the tightest ceiling it serves. An application that supports five processes inherits the most demanding one. This is where shared platforms such as identity, core databases, messaging and the network land in the top tier, usually to IT's surprise.
  4. Ask the data-loss question separately. RPO comes from a different interview question: "If we restored this system to how it looked at 9:00 and it is now 9:40, what would you have to do to recover those 40 minutes?" If the answer is "re-key it from email," a longer RPO may be fine. If the answer is "we would not know which payments went out," the RPO is close to zero, whatever that costs.
  5. Record the dependency chain. A tier-1 application that authenticates against a tier-3 directory service is, in practice, a tier-3 application.

A worked example

Take a hypothetical regional insurer's claims operation. The firm and figures are illustrative.

The BIA finds that claims intake can stop for about a day before complaint-handling timelines and customer harm become intolerable. MTPD: 24 hours. Adjusters say that after an outage they need roughly four hours to re-key claims taken on paper and check open files. Leadership wants four hours of margin for declaration and surprises.

Twenty-four hours, minus four of WRT, minus four of margin, leaves a 16-hour ceiling. The claims platform gets a 12-hour RTO, inside the ceiling with room to spare.

For RPO, the adjusters explain that claims arrive by phone, web form and broker email. Phone and email claims can be reconstructed from call recordings and mailboxes. Web submissions cannot, because customers will not resubmit a form they think they already sent. So the web-intake database gets a 15-minute RPO, while the document store holding photos and repair estimates can tolerate 24 hours.

One "application" ended up with two RPOs because its data stores have different reconstruction costs. That is normal, and a good reason to set RPO per data store.

Tiering: a table that keeps the conversation short

Nobody wants to negotiate targets for 400 applications one at a time. Tiers give you a handful of standard targets, each tied to a standard architecture and price. Treat the numbers below as a starting point.

Tier Typical RTO Typical RPO Example workloads Usual DR pattern Test cadence
0: Foundation Under 1 hr Near zero Identity, DNS, core network, key management, privileged access Redundant across sites; rebuilt first in any scenario Component tests quarterly, full test annually
1: Critical 1–4 hrs Under 15 min Payments, trading, order capture, patient-facing systems Synchronous or near-synchronous replication, warm standby Twice a year
2: Important 4–24 hrs 1–4 hrs Claims, CRM, reporting that feeds daily obligations Asynchronous replication or frequent snapshots, pilot light Annually
3: Deferrable 1–3 days 24 hrs HR systems, intranet, most analytics Restore from backup onto rebuilt infrastructure Annual sample restore
4: Rebuild Over 3 days 24 hrs or more Archives, dev and test, legacy reference systems Restore when capacity allows Spot-check restores

Tier 0 exists because foundation services are what everything else waits on, and they rarely appear in a BIA because no process owner names "the directory" as a dependency. The "usual DR pattern" column is what makes tiering useful: once an application lands in a tier, its architecture and budget follow.

The cost curve nobody draws

Every recovery target has a price, and the price is not linear. Moving an RPO from 24 hours to four is usually cheap: more frequent backups or snapshots. Moving it from four hours to 15 minutes means continuous asynchronous replication and a second set of storage. Moving it from 15 minutes to zero means synchronous replication, which limits the distance between sites (every write waits for the far side to acknowledge) and often adds database licensing.

RTO behaves the same way. Restore from backup is cheap until the restore takes 30 hours because you are pulling 80 TB over a link sized for nightly increments. Warm standby costs a running copy of the environment. Active-active costs that plus the engineering to make the application tolerate it.

Put two curves on one chart: the cost of downtime rising with outage length (from the BIA), and the cost of recovery capability falling as targets lengthen. The sensible target sits near the point where total cost is lowest, adjusted for things that are not purely financial, such as regulatory deadlines and safety. The downtime curve is usually steeper than people expect; our briefing on why ransomware downtime costs more than the ransom breaks down the hourly costs that pile up beyond lost revenue.

It is also the chart most likely to get a CFO's attention, because it prices each hour of target against each hour of outage.

How backup and replication choices set your real RPO

The RPO you get is set by the technology as configured, whatever the plan says. A quick mapping:

  • Nightly backup: RPO of up to about 24 hours, plus the backup window. If the 22:00 job fails quietly two nights running, your effective RPO is three days.
  • Snapshots every N hours: RPO of N hours, provided the snapshots are at least crash-consistent and, for databases, application-consistent.
  • Transaction log backups or log shipping: RPO equal to the log interval, commonly 5 to 15 minutes.
  • Asynchronous replication: seconds to minutes under normal load, but lag grows with heavy write volume or a degraded link. Alert on the lag itself, not only on whether the link is up.
  • Synchronous replication: near-zero RPO for the replicated system, with latency costs and distance limits.

One warning: replication protects against losing a site, not against corrupted data. A bad batch job, a malicious administrator or ransomware damages the primary, and replication delivers the damage to the secondary within seconds. Tier-1 systems therefore need both replication, for availability, and point-in-time copies that cannot be altered, for integrity. We cover that design in immutable backups and clean-room recovery.

Proving the RTO you actually have

An RTO is a hypothesis until a test measures it, and the measurement has to be honest about the clock.

  • Start the clock at the event, not when engineers begin work. Detection, escalation and the decision to declare routinely eat an hour or more. If your RTO runs from declaration, measure that interval anyway.
  • Stop the clock when the business confirms the system is usable, not when a server answers a ping. Have a process owner complete a real transaction.
  • Test at realistic scale. Restoring one virtual machine proves the procedure. Restoring 300 shows whether the backup platform, the network and the people can do it in parallel.
  • Record achieved RTO and RPO per application, next to the target. A test that "passed" without producing a number taught you very little.

Achieved figures belong in the recovery section of the plan itself; the annotated sample plan table of contents shows where they sit alongside procedures and contact lists. Auditors and examiners tend to ask for achieved-versus-target evidence, not just the targets.

Mismatches that show up in almost every review

  1. Tier-1 application, tier-3 dependency. The app replicates; the directory, certificate authority or license server it needs at startup does not.
  2. RPO on paper, nightly backup in practice. The plan says 15 minutes. The only protection is the 22:00 job.
  3. RTO assumes a local copy. The usable copy sits in cloud storage with egress throughput nobody has tested.
  4. No WRT. The system is up at hour four; the business is functioning at hour nine.
  5. Targets set for an ordinary Tuesday. Month-end, open enrollment and holiday peaks change both the impact curve and the data volume.
  6. Nobody owns the gap. The target is four hours, the test showed eleven, and nobody is accountable for closing it or formally accepting the risk.

A decision rule for your next review

For each application, ask four questions in order and stop at the first "no":

  1. Does RTO plus WRT fit inside the MTPD of every process this application supports, with margin?
  2. Does the protection technology, as configured today, deliver the stated RPO for each data store?
  3. Are all upstream dependencies at the same tier or better?
  4. Has a test in the last twelve months measured an achieved RTO and RPO at or under target?

A "no" at question one is a business conversation: change the target, change the process, or accept the risk in writing. A "no" at two or three is an engineering backlog item with a cost estimate attached. A "no" at four means the target is still a guess, however carefully derived.

RTO vs. RPO: Setting Recovery Targets Without Guesswork | CPE World