Replication Topologies and Failure Domains
LESSON
Replication Topologies and Failure Domains
By the end of this lesson, you will be able to...
map a replication layout to the failure domains and shared dependencies it really uses;
distinguish its commit, failover, and repair paths;
choose a topology that matches a stated latency and recovery promise.
Idea in one sentence: Replica count is not resilience; a topology earns a guarantee only through the domains and paths that must still work when something fails.
Core Insight
At 10:02, an IAD availability zone loses power while a trader submits reservation R-91. The API must decide whether to acknowledge the write, and the on-call engineer must decide whether the remaining replicas can keep serving it. Harbor Point needs a replicated reservation shard. Someone proposes “three replicas plus a disaster-recovery copy.” It sounds like a design, but it leaves out the facts that decide whether a confirmation survives: where the voters are, which dependencies they share, and whether a remote copy is part of the acknowledgement rule.
Three processes on one rack, three zones in one region, and three regions joined by a WAN all have different failure behavior. The useful unit is a failure domain: a boundary inside which failures may be correlated. It can be a host, rack, power zone, region, identity service, DNS dependency, backbone link, or backup store.
The initial model is “more copies mean safer data.” It works only when the copies fail independently and the client promise says which copies must receive a write. The stronger model is a topology map with three visible paths: commit, failover, and repair.
The immediate investigation has three questions. Which nodes can acknowledge R-91 after the zone loss? Which nodes can choose the next authority if the leader was in the lost zone? And, if one surviving replica was already behind, can it repair from retained history or does it need a snapshot? A topology is useful when it answers those questions before an incident, not when it merely labels boxes as replicas.
Start From the Promise
Harbor Point has four product requirements:
confirmed reservations: one ordered authority, no double booking
East Coast writes: usually below 30 ms
Dublin dashboard: may be up to five seconds stale
whole primary-region loss: at most 60 seconds of accepted history lost
The last requirement is an RPO, not a claim of zero-loss regional failover. That difference determines whether a remote replica must vote before a success response.
Make the Domains Visible
The following sketch is a teaching model. It is useful because it names dependencies rather than merely counting processes.
shard 184
iad zone A: iad-1, leader, voter
iad zone B: iad-2, voter
iad zone C: iad-3, voter
dub zone A: dub-1, asynchronous recovery replica
shared dependencies: DNS, identity control plane,
the IAD-DUB network path, and backup object storage
This map supports questions a replica count cannot answer. Losing one IAD zone leaves two voters, so normal local quorum commits can continue. Losing IAD entirely leaves Dublin with only whatever it has durably replayed. Losing the object store may not stop normal replication, but it can prevent a replacement from being seeded or a recovery from being verified.
The map also stops a common mistake: calling zones independent while placing every node behind the same authentication or network dependency. Independence is a claim to investigate, not a label inherited from a cloud diagram.
Two Designs, Three Paths
Compare two layouts for the same service.
A. Local voting quorum + remote asynchronous replica
client -> iad leader -> iad voter 1 and iad voter 2
\-> async stream -> Dublin
B. Stretched voting quorum
client -> iad leader -> iad voter and Dublin voter
Each layout contains three paths.
| Path | Question | A: local quorum + async DR | B: stretched quorum |
|---|---|---|---|
| Commit | What must acknowledge before success? | Local majority. | A quorum that includes the remote voter under this design. |
| Failover | Who can establish authority after loss? | Local voters automatically; regional recovery uses the remote durable prefix and a runbook. | Remote voter participates in the voting story. |
| Repair | How does a lagger return? | Stream retained history or seed from snapshot. | Same need, now sharing more critical WAN capacity. |
Neither layout is simply “better.” Layout A keeps ordinary commits close to IAD users. It is a good fit when the business accepts the stated regional RPO and monitors remote lag. Layout B can support a stronger regional durability promise, but it imports WAN delay and failure behavior into the normal write path. That is the trade-off: lower normal latency versus a stronger acknowledgement boundary.
A Worked Decision
At 10:00, an IAD zone fails. In layout A, iad-1 is gone but iad-2 and iad-3 remain. They still form a local majority and can elect or confirm an authority according to the service protocol. Reservation writes continue with their normal acknowledgement rule.
At 10:03, the IAD region is lost. Dublin has replayed through 10:02:20. Promoting it can make the service available again, but it cannot recover a write that was accepted at 10:02:50 unless that write reached Dublin. The design has not failed its stated 60-second RPO if the lag was within budget; it would fail a zero-loss regional contract.
This trace changes the design conversation. “Dublin is a replica” becomes “Dublin is a recovery replica with this durable prefix, this promotion rule, and this measured loss boundary.” If product changes the promise to no acknowledged loss after regional failure, the commit path must include remote durability or a different globally coordinated mechanism.
The repair path matters before disaster too. If Dublin is offline long enough that the leader cannot retain the needed history, it needs a snapshot from a known committed state. A topology that cannot seed, validate, and route that replacement has copies on paper but a weak recovery plan in practice.
Consider a second failure at 10:04: one IAD survivor has only partial history because it was recovering when the zone failed. The remaining authority cannot simply count a process as useful because it is reachable. It needs a current committed prefix before that replica can acknowledge safely or serve a freshness-sensitive read. The repair path therefore consumes the same network, disk, and operational attention as the normal topology. A design that places every copy in independent zones but cannot replenish one after failure has resilience only for its first incident.
That recovery capability must be exercised before production relies on it.
What This Changes
Before this model, a team might add replicas until a diagram looks redundant. After it, the team writes a contract beside each path:
| Claim | Evidence to require |
|---|---|
| Zone failure does not lose acknowledged writes. | Voters span the relevant zones and acknowledgement uses a surviving quorum. |
| Regional loss stays inside the RPO. | Remote durable lag and archive/recovery evidence stay inside the budget. |
| Dashboard may be stale but confirmations may not. | Routing separates the asynchronous read path from authority reads. |
| A replacement can return safely. | Retained history or snapshot, validation, and an epoch-aware promotion path exist. |
Consequences, Trade-offs, and Limits
More distance can improve survivability, but it costs latency, operational complexity, and sensitivity to WAN failure. Keeping voters local gives predictable commits but leaves a recovery window for regional loss. An asynchronous replica can serve a stale dashboard; it cannot honestly serve a confirmation endpoint that promises the latest committed result.
Topology alone does not solve corruption, duplicate client requests, or unsafe promotion. Watch remote durable lag, quorum health, shared-dependency failures, retained-log pressure, and restore validation. Those signals say when the map no longer supports the promise written beside it.
Check Your Understanding
Check: Three voters run in separate zones but all depend on one identity service that is unavailable. Does “three-zone replication” guarantee write availability?
Think first, then reveal.
Answer: No. The identity service is a shared failure domain if writes or node operation depend on it. The topology must include such dependencies, not only machine placement.
Check: A remote asynchronous replica is 45 seconds behind and the published regional RPO is 60 seconds. May the team call this zero-loss failover?
Think first, then reveal.
Answer: No. It may still be inside the stated RPO, but a recent acknowledged write can be absent after regional promotion. Zero-loss requires a stronger acknowledgement boundary.
Practice: Review a Topology
Sketch a three-replica service with one remote copy. Name its failure domains, its commit acknowledgements, its regional-loss promise, and its repair source.
A good answer should mention: at least one shared dependency; whether the remote copy is synchronous or asynchronous; the RPO implied by that choice; which endpoints may use it for reads; and how a lagger receives retained log or a snapshot.
Connections
- Consistency Contracts and API Semantics supplies the promises that topology must actually support.
- Log Shipping and Ordered Apply explains the ordered history that follows the chosen replication paths.
- Replication Flow Control and Backpressure makes the repair path operational when a follower falls behind.
Resources
- [PAPER] Spanner: Google's Globally Distributed Database — Focus: Relate placement and quorum choices to an explicit multi-region transaction promise.
- [DOC] PostgreSQL Synchronous Replication — Focus: See how standby selection changes a commit acknowledgement boundary.
- [PAPER] Dynamo — Focus: Examine replica placement and preference lists as availability design choices.
Key Takeaways
- A topology must name failure domains and shared dependencies, not only replica count.
- Commit, failover, and repair paths can have different participants and different failure behavior.
- Local quorum plus remote asynchronous recovery buys low latency and bounded regional loss; stronger regional durability moves remote cost onto the acknowledgement path.