System designby Learnastra

Concept lesson · Foundations

Multi-region architecture and disaster recovery

By Anup Rai

Start here

Definition

Multi-region architecture deploys a service across geographically separate regions. Disaster recovery is the planned restoration of usable service and data after a major disruption; the recovery point objective (RPO) specifies the targeted data-loss window and the recovery time objective (RTO) specifies the targeted restoration time.

Why it matters: A regional outage, accidental deletion, or failed dependency can affect every local replica. Recovery requires knowing which saved changes survived, ensuring only the designated replacement can accept writes, and providing enough capacity to serve users.

The visual modelRecovery point objective (RPO) and recovery time objective (RTO)

Compare two intervals: how far the recovered data lags behind the disruption, and how long users wait for service to return.

Recovery point objective (RPO) and recovery time objective (RTO)Compare two intervals: how far the recovered data lags behind the disruption, and how long users wait for service to return. West has O16 at 12:00:00. East acknowledges O17 at 12:00:04 but it has not replicated. Connectivity fails at 12:00:05. The safe replica is five seconds behind disruption, and O17 may be lost despite its acknowledgement. Service is validated at 12:07:05: measured recovery takes seven minutes. RPO and RTO are objectives against these separate kinds of loss. Fence the old writer before promotion, and reconcile a possibly successful external payment before retrying it.Two clocks: missing data versus unavailable servicetime (not to scale)12:00:0012:00:0512:07:05West durable O16at 12:00:00Link failsO17 not in WestTraffic validatedservice recoveredData recovery gap: 5 secondsRecovery: 7 minutesO17 was ACKed at 12:00:04. A second region does not automatically give zero data loss.
Read the diagram step by step
  1. West has O16 at 12:00:00. East acknowledges O17 at 12:00:04 but it has not replicated. Connectivity fails at 12:00:05.
  2. The safe replica is five seconds behind disruption, and O17 may be lost despite its acknowledgement.
  3. Service is validated at 12:07:05: measured recovery takes seven minutes. RPO and RTO are objectives against these separate kinds of loss.
  4. Fence the old writer before promotion, and reconcile a possibly successful external payment before retrying it.

Worked example

East acknowledges O17 at 12:00:04, fails at 12:00:05, and West has only data through 12:00:00. Restoring West can miss O17; restoring service at 12:07:05 takes 7 minutes.

Key takeaways

You will learn to

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Replication and durability · Quorums, consensus, leases, and fencing

Workload and timing examples are interview assumptions.

01Multi-region architecture, high availability, and disaster recovery

Multi-region architecture runs a service across geographically separate deployment regions. Disaster recovery is the planned restoration of usable service and data after a major disruption. A region is a geographical deployment area whose infrastructure can share risks such as a regional network failure; an availability zone is a separate failure domain within a region under the provider’s isolation model. Putting servers in two locations does not provide regional recovery if both still depend on the same regional database, credential service, or network.

High availability keeps the service operating through expected component failures. Disaster recovery restores a useful service after a larger disruption. Backups preserve earlier recoverable states. These capabilities overlap, but a replica that immediately copies an accidental deletion is not a substitute for a backup that can restore yesterday's data.

Specify allowed data loss and recovery time before choosing a regional topology. An East-primary/West-asynchronous-replica example illustrates the tradeoff: an acknowledged order O17 can be absent from West when East fails. Whether that loss is acceptable, and whether writes may pause during recovery, determines the required coordination and cost.

02Recovery point objective (RPO) and recovery time objective (RTO)

The recovery point objective, RPO, is the target maximum amount of data loss measured as a time window. An RPO of 30 seconds means the recovery plan targets a recoverable state no more than 30 seconds behind the disruption. The recovery time objective, RTO, is the target time to restore the agreed service after disruption. Neither is a guarantee merely because it appears in a diagram.

Concept in focusRPO looks at lost history; RTO looks at downtime

The timestamps are concrete; distances on this timeline are not to scale. The gaps shown are actual outcomes, to compare with the objectives.

RPO looks at lost history; RTO looks at downtimeThe timestamps are concrete; distances on this timeline are not to scale. The gaps shown are actual outcomes, to compare with the objectives. Measure the history gap before the disruption and the recovery duration after it. Recoverable data stops at 11:59:40; disruption is at 12:00:00: a 20-second gap. Service is usable at 12:05:00: five minutes of recovery.Recovery-point gap and recovery duration11:59:40last recoverable12:00:00disruption12:05:00service usable20 s of history5 min recoveryCompare the first gap with RPO, the second with RTO. Not to scale.

Remember: Look backward for the recovery point; forward for service recovery.

Read the diagram
  1. Measure the history gap before the disruption and the recovery duration after it.
  2. Recoverable data stops at 11:59:40; disruption is at 12:00:00: a 20-second gap.
  3. Service is usable at 12:05:00: five minutes of recovery.
Try from memoryWhich gap would a 10-second RPO fail to meet?

The 20-second history gap from 11:59:40 to 12:00:00. The five-minute service recovery is compared with RTO instead.

Objective Example What must support it
RPO At most 30 seconds of accepted changes lost Replication or recoverable logs within that bound, plus measurement
RTO Ordering restored within 10 minutes Detection, safe promotion, routing, capacity, and validation within that budget
Restore correctness Existing payments reconciled Durable external IDs and recovery procedures

03Active-passive, active-active, and write ownership

A regional topology defines where the service runs and which regions may serve each operation. Compare write ownership separately from replication timing: a region may serve reads while another owns writes, and a write may wait for remote durability before success. These choices determine both normal latency and what remains possible after a region is lost.

Topology Write behavior Benefit Cost or limit
Primary with asynchronous standby East writes; West catches up later Simple normal ownership and lower write coordination cost Acknowledged changes may be missing after regional loss
Cross-region synchronous commit Success waits for the required remote durable state Can protect acknowledged writes against the named regional failure Network latency and possible refusal during partitions
Multiple serving regions, one home writer per key Each tenant/key has a defined write owner Geographic service without arbitrary concurrent conflict Remote writes and ownership-transfer work remain
Concurrent writers with defined merge semantics Regions independently accept mergeable operations More local write availability for suitable data Not safe for arbitrary inventory, money, or ownership changes

Start with a primary region and a standby. East owns writes. West receives the ordered change stream. Reads may use West only under a stated staleness policy. This is easier to reason about than allowing both regions to update the same inventory row independently.

Active-active means more than two copies of a web server. If both regions accept writes, specify ownership or conflict handling. Assigning each tenant a home region gives one authority per tenant. Globally coordinating a row can preserve stricter guarantees but adds cross-region latency. Accepting concurrent updates and merging them requires business-compatible semantics; “last timestamp wins” can silently erase an order or inventory reservation.

Read replicas, immutable assets, and regional caches can reduce geographic read latency without making all writes multi-primary. Choose the narrowest distributed-write requirement the product actually needs.

A different design puts one voting, data-bearing replica in each of three regions and commits through a proven majority protocol. Every acknowledged write is durable in two regions. After any one region is lost, the two survivors can elect according to the protocol and recover the committed history; a lagging survivor cannot simply ignore the protocol's election restrictions. This is a constructed quorum example, not a claim that every three-region product uses this layout. It costs cross-region commit latency and still depends on surviving network and service capacity.

04Regional failover: detection, fencing, promotion, and routing

Assume East acknowledged O16 at 12:00:00 and West durably applied it. East acknowledged O17 at 12:00:04, but its log entry has not reached West. Connectivity fails at 12:00:05.

  1. At 12:00:10 monitoring detects failure. It cannot infer whether East is dead or merely unreachable from West.
  2. A promotion procedure establishes that the old writer cannot continue accepted writes under the ownership protocol. A fencing epoch is an increasing ownership-generation number. Resources that check the current epoch can reject an old writer’s operations; changing a number without an enforcing resource does not stop the old process.
  3. West is promoted from its last safe durable position. In this example O17 may be absent, despite its prior acknowledgement. The observed loss window is five seconds; the missing record was accepted one second before disruption.
  4. Routing moves eligible traffic. DNS caches, connection pools, and clients may keep using old endpoints, so routing changes alone do not fence the old writer.
  5. The team validates order creation and payment reconciliation before declaring recovery complete. If that happens at 12:07:05, service recovery took seven minutes.

The client retries O17 using its original operation identity. A payment might have succeeded outside the lost database state. The recovery path queries the payment attempt or reconciles provider events rather than charging blindly. The write and payment contracts must survive the disaster plan together.

The payment recovery identity must also survive. Store the original client operation ID and provider attempt/resource reference in recoverable state, or ensure the provider can recover the mapping from a durable business reference. If both the mapping and the acknowledged order are lost, the client retry alone does not prove whether a charge exists. Hold new charging attempts while reconciliation reconstructs that fact.

Worked example diagramIn this asynchronous example, recent acknowledged writes may be lost if they never reached the standby. Promotion requires a verified stop of the old writer; merely changing West’s local epoch or DNS cannot enforce that stop. Payment operation identities must remain recoverable.
Multi-region architecture and disaster recovery: architecture diagram1. East: writer epoch 7 to 2. West: asynchronous standby: replicate durable log; 3. O17 acknowledged in East to 1. East: writer epoch 7: may not yet exist in West; 2. West: asynchronous standby to 4. Verify East cannot write; then promote: last safe recovery position; 4. Verify East cannot write; then promote to 5. West: writer epoch 8: only after old writer is stopped; 5. West: writer epoch 8 to 6. Validate and reconcile payments: restore useful service1 → 2: replicate durable log3 → 1: may not yet exist in West2 → 4: last safe recovery position4 → 5: only after old writer is stopped5 → 6: restore useful service01East: writer epoch 702West: asynchronousstandby03O17 acknowledged inEast04Verify East cannotwrite; then promote05West: writer epoch 806Validate andreconcile payments
  1. 1 → 2replicate durable logEast: writer epoch 7 → West: asynchronous standby
  2. 3 → 1may not yet exist in WestO17 acknowledged in East → East: writer epoch 7
  3. 2 → 4last safe recovery positionWest: asynchronous standby → Verify East cannot write; then promote
  4. 4 → 5only after old writer is stoppedVerify East cannot write; then promote → West: writer epoch 8
  5. 5 → 6restore useful serviceWest: writer epoch 8 → Validate and reconcile payments

05Standby capacity and restore-time estimates

A warm standby has some running resources and scales up during recovery. A hot standby keeps more capacity ready. Backup-and-restore starts from stored snapshots/logs and generally has more work on the recovery path. These are cost and recovery-time choices, not universal time guarantees.

Suppose peak traffic is 10,000 requests/s and West is provisioned for 2,000. Promotion without a capacity plan creates a second outage. Reserve or validate capacity, warm critical caches carefully, and use admission control while recovering. Include database connections, queue throughput, key management, identity providers, configuration, and secrets distribution in the dependency inventory.

For backup transfer alone, restoring 6 TB over a sustained 1 GB/s path takes approximately 6,000 seconds, or 100 minutes, before replay, indexing, startup, and validation. That cannot support a ten-minute RTO without another recovery mechanism. Use measured restore throughput, not a network-interface headline rate.

06Backups, point-in-time recovery, and restore validation

Point-in-time recovery restores a backup and replays retained changes only up to a selected moment. Choosing a point before a destructive update can recover data that live replicas have already deleted. The backup, required log history and decryption keys must all be available for that selected point.

Replication can faithfully copy corruption, deletion, or an application bug. Preserve point-in-time recovery logs and backups under access and retention policies that reduce correlated loss. Test restoration into an isolated environment, validate application-level invariants, and measure the entire process.

Retention has a business and security cost. Keep enough history to detect and recover from plausible mistakes while applying deletion and regulatory obligations deliberately. A disaster-recovery copy remains sensitive production data.

07Failback and disaster-recovery exercises

When East returns, it may have different data from West. Keep West in charge of new writes. Rebuild or reconcile East from West, verify replication, then plan the transfer back. Choose a clear switch point and prevent the former writer from continuing afterward. Old clients and running jobs must be rejected if they use an obsolete ownership version.

Run exercises that fail a database, sever regional connectivity, remove a dependency, and restore a backup. Record detection time, last recoverable write, promotion time, routing convergence, and usable capacity. The interview answer becomes credible when it identifies which promise the exercise validates and what would prevent declaring success.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

What are RPO and RTO in disaster recovery?

Reveal a model answer

RPO, recovery point objective, is the targeted maximum data-loss window. RTO, recovery time objective, is the targeted time to restore the agreed service. If the last recoverable state is 5 seconds before a disruption and service returns 7 minutes later, those are separate data-loss and restoration measurements to compare with the objectives.

What the answer must demonstrate: Keep the two objectives separate and distinguish targets from measured guarantees.

Applied · Question 2

Can asynchronous regional replication promise zero loss of acknowledged writes?

Reveal a model answer

“Not by itself. East can acknowledge a write and fail before it reaches West. To survive that regional loss without losing acknowledged writes, acknowledgement must require a durable copy or quorum outside East, within the stated failure model.”

What the answer must demonstrate: Place the acknowledgement boundary.

Applied · Question 3

The East primary stops responding to West. Why is that alone insufficient to promote West safely?

Reveal a model answer

A failed health check cannot prove that East stopped writing. In the asynchronous two-region design, I require a verified stop or removal of its write capability before promotion; if that is impossible, writes remain paused. Alternatively, a proven quorum protocol prevents the isolated minority from committing. A new epoch stored only in West is not sufficient fencing.

What the answer must demonstrate: Separate routing from write authority.

Applied · Question 4

Would active-active remove all regional outages?

Reveal a model answer

“It can improve continuity for some operations, but shared dependencies and write conflicts remain. I would state whether each key has one home writer, uses global coordination, or permits a defined merge. Inventory cannot simply merge arbitrary decrements without a rule.”

What the answer must demonstrate: Describe per-record semantics.

Applied · Question 5

Can a nightly backup meet a 30-second RPO?

Reveal a model answer

“A snapshot alone cannot: it can leave almost a day of changes absent. Continuous recoverable logs or another replication mechanism may narrow that gap. I also need to test restore and replay time against the RTO.”

What the answer must demonstrate: Check both freshness and duration.

Applied · Question 6

The standby has one fifth of peak capacity. Is failover ready?

Reveal a model answer

“Only if the recovery contract allows bounded degradation and the remaining capacity or scaling is verified. I would test databases and dependencies too, prioritize essential operations, and limit admission rather than overload the new primary.”

What the answer must demonstrate: Capacity is part of recovery.

Applied · Question 7

Why keep backups when there are three replicas?

Reveal a model answer

“Replicas can copy an accidental deletion or corruption. Backups and point-in-time recovery preserve earlier states under a separate protection policy. I would regularly restore and validate business records, not only check the backup job status.”

What the answer must demonstrate: Replication is not historical recovery.

Applied · Question 8

An old primary region recovers after failover. Why should writes not immediately be routed back?

Reveal a model answer

“West has accepted new writes, so keep it in charge. Bring East up to date or rebuild it from West, verify the data, then switch writers through a controlled handover. Prevent the former writer from continuing. If the regions have conflicting histories, resolve them before switching.”

What the answer must demonstrate: Failback is a controlled state transition.

Blank-page exercise · 20 minutes

Build the answer yourself

Design recovery for an order service with a 30-second RPO and ten-minute RTO. Then change the requirement to no loss of acknowledged orders.

  • Place each acknowledgement and durable copy.
  • Show the isolated old writer and its fencing mechanism.
  • Budget detection, promotion, routing, and validation time.
  • Include capacity, payments, backups, and failback.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Multi-region architecture and disaster recoveryRPO vs RTORecall first, then reveal

RPO: how far back data may go. RTO: how long recovery may take.

Point = data; time = service.

Return to lesson
Multi-region architecture and disaster recoveryWest is ready to take over. Why not redirect traffic immediately?Recall first, then reveal

East may still accept writes. First prevent the old writer from committing, then promote West, route traffic and validate recovery.

Stop old writes → enable new writer → verify.

Return to lesson
Multi-region architecture and disaster recoveryBackup vs replicaRecall first, then reveal

A replica follows changes; a protected backup preserves an earlier recovery point.

Copies need history.

Return to lesson

Final revision

Summary and interview notes

Disaster recovery is a tested procedure for restoring an agreed service from a surviving data point. Choose RPO and RTO first, then align acknowledgment, replica placement, write authority, capacity, external-effect recovery and failback with those objectives.

Remember these points

  • RPO is the target data-loss window; RTO is the target restoration time, and observed lag is neither promise by itself.
  • Asynchronous replication can lose acknowledged writes; zero-loss acknowledgment must depend on state surviving the named failure.
  • Replica and voter placement matter: a majority concentrated in one region does not survive that region’s loss.
  • Changing routes does not stop the old writer. Before promoting another, enforce exclusive write ownership or verify that the old writer has stopped.
  • Backups protect historical recovery points, while replicas can quickly copy corruption and deletion.

Interview tips

  • Mark every acknowledgment and durable copy on the failover trace.
  • Challenge the design with a partition where the old primary remains alive, not only a clean power-off.
  • Budget detection, authority transfer, capacity, routing and validation; calculate restore bytes divided by measured throughput.

Important qualifications

  • Six decimal TB at one GB/s needs about 100 minutes for transfer alone.
  • Payment identity and encryption-key recovery must survive the disaster along with primary business records.
  • Before moving back to the recovered region, rebuild or reconcile its data from the region currently accepting writes.

Technical references

Practice marks stay in this browser.