Path 04Lesson 6 / 10

Choose availability across zones and regions

Compare high availability, Multi-AZ, and multi-region designs. Trace the complete request path and test the failure each design must withstand.

Practitioner12 minReviewed

Published by How we write

What you will learn

  • Explain the difference between an Availability Zone and a Region.
  • Find shared dependencies that defeat an availability design.
  • Compare the business value and operating cost of multi-region deployment.

Start with the user operation

High availability (HA) aims to keep a service usable despite component failures. Define usability before choosing the architecture. A booking page that loads while all booking requests fail is not an available booking service.

Set a service level objective (SLO) for the important operation. Define which requests count, what success means, and the measurement period. A cloud service SLA describes that provider’s commitment. It does not establish your application’s measured availability.

For illustration, 99.9% time-based availability permits 43.2 minutes of unavailability in a 30-day month. A request-based SLO has a different denominator. Neither measure says how much data you can lose or guarantees a maximum individual outage.

Understand the failure boundaries

An AWS Availability Zone (AZ) is an isolated infrastructure location within a Region. A Region contains multiple AZs. A multi-region design distributes workload components across Regions. Other providers have their own boundaries and service behavior; inspect the selected service.

DesignFailure it can help addressWhat still needs a design
Multiple processes in one AZA process or host failureAZ loss and shared dependencies
Multi-AZ in one RegionLoss of an AZRegional failure, data corruption, and recovery
Multiple RegionsLoss of a RegionRouting, data consistency, capacity, and shared services

These are design possibilities, not availability guarantees. A label does not prove that every required component uses the intended boundary.

Trace the complete request path

Consider a fictional booking service. Web replicas run in two AZs. Both use one database and one outbound gateway in AZ A. The gateway is required to call the payment provider.

If AZ A fails, the web replica in AZ B may remain healthy while booking still fails. The team must assess the database, network path, identity provider, payment dependency, and routing. Inspect the actual managed database mode: replication, failover, and reader behavior vary by product and configuration.

Also check capacity. The surviving resources must handle the required load. A design that relies on creating capacity during an incident depends on quotas, available resources, and control-plane operations.

Run a controlled exercise with a defined boundary, stop conditions, and an owner. Verify a complete booking, including payment reconciliation. Record failed requests and the time to recover useful operation.

Decide whether another Region solves the problem

Multi-region operation adds data transfer, duplicate resources, deployment coordination, and operational work. Active/passive keeps one environment ready to take traffic. Active/active serves traffic in more than one environment. The required readiness and data behavior differ.

For the booking service, concurrent writes introduce a question: can two Regions sell the same seat? Define the authority for a booking and the behavior during interrupted replication. “Replicate the database” is not a complete answer.

Check permitted data locations, encryption keys, certificates, DNS, secrets, and external services. A common identity outage or a bad release can affect multiple Regions. More locations do not remove every common cause.

Connect availability to recovery

HA handles specified failures during operation. Disaster recovery restores a usable service and its data after a disruptive event. A multi-region service still needs a recovery plan for deletion or corrupted data.

Document the chosen failure scenarios and the ones the business accepts. Keep tests and infrastructure definitions aligned as the application changes. Continue with RTO, RPO, and disaster recovery.

Do the exercise

A fictional booking service runs web replicas in two AZs. Its database and outbound gateway occupy one AZ. Draw the request path. Remove that AZ on paper. Identify what still works, what fails, and which test would verify your conclusion.

Download worksheet (Markdown)

Check your understanding

Two web replicas run in different AZs. Both require the same database in one AZ. What does this prove?

Sources & further reading

Related reading from Taiga