Set and test RTO and RPO
Define acceptable interruption and data loss. Compare recovery strategies and measure a complete recovery exercise against business requirements.
Published by TaigaHow we write
What you will learn
- Distinguish RTO from RPO and availability.
- Calculate elapsed recovery time and the data recovery gap.
- Specify a recovery exercise with evidence and a service owner.
Define two separate objectives
Recovery Time Objective (RTO) sets the maximum acceptable interruption before useful service must return. Recovery Point Objective (RPO) sets the maximum acceptable data loss measured as time. Agree these objectives with the business owner for a defined service and failure scenario.
An availability target describes service performance over a period. RTO and RPO describe recovery expectations. They answer different questions.
For a fictional ordering service, the owner sets RTO to 60 minutes and RPO to 15 minutes. These are example values, not general recommendations. Another service may need different limits because missing orders and delayed reports have different consequences.
Measure the complete recovery
The service stops at 10:00. The team records this exercise:
| Stage | Duration | Clock time |
|---|---|---|
| Detect the interruption | 8 minutes | 10:08 |
| Assess and authorize recovery | 12 minutes | 10:20 |
| Restore the service and data | 25 minutes | 10:45 |
| Validate useful operation | 10 minutes | 10:55 |
Elapsed recovery time is 55 minutes. The exercise meets the 60-minute RTO. Counting only the 25-minute restore operation would hide most of the interruption.
The latest usable recovery point is 09:40. The gap to the 10:00 interruption is 20 minutes. This misses the 15-minute RPO by 5 minutes. Restoring the same data faster would not close that gap.
Inspect actual missing or inconsistent records. A time gap describes exposure; it does not count the affected orders. Reconcile external payment and fulfillment records before resuming normal processing. Try different assumptions in the recovery exercise.
Choose a recovery strategy
A strategy must cover the required service, data, and dependencies. Compare these patterns against measured objectives:
| Pattern | Prepared before the event |
|---|---|
| Backup and restore | Recoverable data plus a way to recreate the environment |
| Pilot light | Essential data services; other components need activation or creation |
| Warm standby | A functioning environment with reduced capacity |
| Active/active | More than one environment already serves traffic |
There are no universal recovery times for these patterns. The implementation, data volume, dependencies, and test conditions determine the result. Include operating cost and team capability in the decision.
Protect against more than an outage
A replica can copy an unwanted deletion or corrupted record. Keep recoverable versions or point-in-time recovery where required. Verify retention, restore permissions, and access to encryption keys. Match backup isolation to the scenario, including loss of access to the primary account.
For regional recovery, check the permitted data location and the complete dependency chain. Include identity, DNS, certificates, secrets, deployment artifacts, quotas, and network access. A recovery environment that lacks one required key can be unusable.
Define who can declare the event, who performs recovery, and who accepts the restored service. Plan failback or continued operation in the recovery environment. Prevent conflicting writers and reconcile data before switching again.
Turn the plan into evidence
Write a runbook and exercise it under controlled conditions. Record the scenario, dataset size, start and end times, recovered data point, failed steps, and owners. Verify a real business operation with safe test records.
Repeat the exercise after relevant changes and on the agreed schedule. A schema change, new external dependency, or different data volume can invalidate earlier results. Link the exercise evidence to the release and operational responsibilities.
Do the exercise
A fictional service stops at 10:00. Detection takes 8 minutes, a decision takes 12, restoration takes 25, and validation takes 10. The latest usable data is from 09:40. Compare the result with RTO 60 minutes and RPO 15 minutes. Propose one improvement for each objective.
Download worksheet (Markdown)Check your understanding
Sources & further reading
- AWS: Define recovery objectives for downtime and data loss ↗
- AWS: Use defined recovery strategies ↗
- AWS: Testing disaster recovery ↗
Related reading from Taiga
Clearing this selection deletes all progress saved in this browser.
Progress stays in this browser. No account, no tracking.