Path 04Lesson 7 / 10

Set and test RTO and RPO

Define acceptable interruption and data loss. Compare recovery strategies and measure a complete recovery exercise against business requirements.

Practitioner14 minReviewed

Published by How we write

What you will learn

  • Distinguish RTO from RPO and availability.
  • Calculate elapsed recovery time and the data recovery gap.
  • Specify a recovery exercise with evidence and a service owner.

Define two separate objectives

Recovery Time Objective (RTO) sets the maximum acceptable interruption before useful service must return. Recovery Point Objective (RPO) sets the maximum acceptable data loss measured as time. Agree these objectives with the business owner for a defined service and failure scenario.

An availability target describes service performance over a period. RTO and RPO describe recovery expectations. They answer different questions.

For a fictional ordering service, the owner sets RTO to 60 minutes and RPO to 15 minutes. These are example values, not general recommendations. Another service may need different limits because missing orders and delayed reports have different consequences.

Measure the complete recovery

The service stops at 10:00. The team records this exercise:

StageDurationClock time
Detect the interruption8 minutes10:08
Assess and authorize recovery12 minutes10:20
Restore the service and data25 minutes10:45
Validate useful operation10 minutes10:55

Elapsed recovery time is 55 minutes. The exercise meets the 60-minute RTO. Counting only the 25-minute restore operation would hide most of the interruption.

The latest usable recovery point is 09:40. The gap to the 10:00 interruption is 20 minutes. This misses the 15-minute RPO by 5 minutes. Restoring the same data faster would not close that gap.

Inspect actual missing or inconsistent records. A time gap describes exposure; it does not count the affected orders. Reconcile external payment and fulfillment records before resuming normal processing. Try different assumptions in the recovery exercise.

Choose a recovery strategy

A strategy must cover the required service, data, and dependencies. Compare these patterns against measured objectives:

PatternPrepared before the event
Backup and restoreRecoverable data plus a way to recreate the environment
Pilot lightEssential data services; other components need activation or creation
Warm standbyA functioning environment with reduced capacity
Active/activeMore than one environment already serves traffic

There are no universal recovery times for these patterns. The implementation, data volume, dependencies, and test conditions determine the result. Include operating cost and team capability in the decision.

Protect against more than an outage

A replica can copy an unwanted deletion or corrupted record. Keep recoverable versions or point-in-time recovery where required. Verify retention, restore permissions, and access to encryption keys. Match backup isolation to the scenario, including loss of access to the primary account.

For regional recovery, check the permitted data location and the complete dependency chain. Include identity, DNS, certificates, secrets, deployment artifacts, quotas, and network access. A recovery environment that lacks one required key can be unusable.

Define who can declare the event, who performs recovery, and who accepts the restored service. Plan failback or continued operation in the recovery environment. Prevent conflicting writers and reconcile data before switching again.

Turn the plan into evidence

Write a runbook and exercise it under controlled conditions. Record the scenario, dataset size, start and end times, recovered data point, failed steps, and owners. Verify a real business operation with safe test records.

Repeat the exercise after relevant changes and on the agreed schedule. A schema change, new external dependency, or different data volume can invalidate earlier results. Link the exercise evidence to the release and operational responsibilities.

Do the exercise

A fictional service stops at 10:00. Detection takes 8 minutes, a decision takes 12, restoration takes 25, and validation takes 10. The latest usable data is from 09:40. Compare the result with RTO 60 minutes and RPO 15 minutes. Propose one improvement for each objective.

Download worksheet (Markdown)

Check your understanding

A recovery exercise restores useful service in 55 minutes. It recovers data from 20 minutes before the interruption. Targets are RTO 60 minutes and RPO 15 minutes. What is the result?

Sources & further reading

Related reading from Taiga