Choose availability across zones and regions
Compare high availability, Multi-AZ, and multi-region designs. Trace the complete request path and test the failure each design must withstand.
Published by TaigaHow we write
What you will learn
- Explain the difference between an Availability Zone and a Region.
- Find shared dependencies that defeat an availability design.
- Compare the business value and operating cost of multi-region deployment.
Start with the user operation
High availability (HA) aims to keep a service usable despite component failures. Define usability before choosing the architecture. A booking page that loads while all booking requests fail is not an available booking service.
Set a service level objective (SLO) for the important operation. Define which requests count, what success means, and the measurement period. A cloud service SLA describes that provider’s commitment. It does not establish your application’s measured availability.
For illustration, 99.9% time-based availability permits 43.2 minutes of unavailability in a 30-day month. A request-based SLO has a different denominator. Neither measure says how much data you can lose or guarantees a maximum individual outage.
Understand the failure boundaries
An AWS Availability Zone (AZ) is an isolated infrastructure location within a Region. A Region contains multiple AZs. A multi-region design distributes workload components across Regions. Other providers have their own boundaries and service behavior; inspect the selected service.
| Design | Failure it can help address | What still needs a design |
|---|---|---|
| Multiple processes in one AZ | A process or host failure | AZ loss and shared dependencies |
| Multi-AZ in one Region | Loss of an AZ | Regional failure, data corruption, and recovery |
| Multiple Regions | Loss of a Region | Routing, data consistency, capacity, and shared services |
These are design possibilities, not availability guarantees. A label does not prove that every required component uses the intended boundary.
Trace the complete request path
Consider a fictional booking service. Web replicas run in two AZs. Both use one database and one outbound gateway in AZ A. The gateway is required to call the payment provider.
If AZ A fails, the web replica in AZ B may remain healthy while booking still fails. The team must assess the database, network path, identity provider, payment dependency, and routing. Inspect the actual managed database mode: replication, failover, and reader behavior vary by product and configuration.
Also check capacity. The surviving resources must handle the required load. A design that relies on creating capacity during an incident depends on quotas, available resources, and control-plane operations.
Run a controlled exercise with a defined boundary, stop conditions, and an owner. Verify a complete booking, including payment reconciliation. Record failed requests and the time to recover useful operation.
Decide whether another Region solves the problem
Multi-region operation adds data transfer, duplicate resources, deployment coordination, and operational work. Active/passive keeps one environment ready to take traffic. Active/active serves traffic in more than one environment. The required readiness and data behavior differ.
For the booking service, concurrent writes introduce a question: can two Regions sell the same seat? Define the authority for a booking and the behavior during interrupted replication. “Replicate the database” is not a complete answer.
Check permitted data locations, encryption keys, certificates, DNS, secrets, and external services. A common identity outage or a bad release can affect multiple Regions. More locations do not remove every common cause.
Connect availability to recovery
HA handles specified failures during operation. Disaster recovery restores a usable service and its data after a disruptive event. A multi-region service still needs a recovery plan for deletion or corrupted data.
Document the chosen failure scenarios and the ones the business accepts. Keep tests and infrastructure definitions aligned as the application changes. Continue with RTO, RPO, and disaster recovery.
Do the exercise
A fictional booking service runs web replicas in two AZs. Its database and outbound gateway occupy one AZ. Draw the request path. Remove that AZ on paper. Identify what still works, what fails, and which test would verify your conclusion.
Download worksheet (Markdown)Check your understanding
Sources & further reading
- AWS: Deploy the workload to multiple locations ↗
- AWS: Shared responsibility model for resiliency ↗
- Google SRE: Implementing SLOs ↗
Related reading from Taiga
Clearing this selection deletes all progress saved in this browser.
Progress stays in this browser. No account, no tracking.