Replication Models: Primary-Backup, Multi-Leader, and Leaderless

LESSON

Consensus and Coordination

009 30 min intermediate

Replication Models: Primary-Backup, Multi-Leader, and Leaderless

By the end of this lesson, you will be able to...

  • Locate the authority boundary in primary-backup, multi-leader, and leaderless designs.

  • Choose a replication model from a data item's conflict cost, locality needs, and availability requirements.

  • Predict which evidence, repair work, or operational signal each model makes important.

Idea in one sentence: Replication models do not remove coordination; they decide whether the system pays for it before a write, while reconciling several writers, or while reading and repairing replicas.

Core Insight

A travel-booking service runs in Europe and North America. Its designers want one replication policy for every piece of state because one policy sounds easier to operate.

The service stores three very different things:

seat/AB123/17A       one traveler may hold it
trip/481/notes       two travelers may edit it from different regions
user/92/last-viewed  a low-stakes “recently viewed hotel” value

All three values can have replicas in both regions. But copying bytes is not the hard question. The hard question is: when two requests disagree or a region is disconnected, which request is allowed to become authoritative?

A seat can be sold only once. A note can often preserve two independent additions. A recently viewed hotel can tolerate a stale value or a simple reconciliation rule. Treating those objects as if they had the same conflict cost either makes the low-stakes path needlessly slow or makes the seat unsafe.

Primary-backup, multi-leader, and leaderless replication are three ways to put that authority boundary in different places. The names describe a family of designs, not one fixed implementation. Their useful comparison is about where a write is admitted, what evidence makes it acceptable, and who must clean up disagreement later.

The Small Situation: Same Copies, Different Promises

Keep the three objects in view. The product needs a different promise for each one.

Object What a bad conflict costs Useful promise
seat/AB123/17A two people may believe they hold the same seat one controlled decision before confirmation
trip/481/notes a collaborator's edit may disappear preserve or explicitly resolve concurrent edits
user/92/last-viewed a recommendation is briefly less relevant keep serving through a replica fault and reconcile later

This table is a teaching model, not a universal schema rule. A real product may decide that trip notes need stronger control, or that a “last viewed” value carries privacy or billing meaning. The point is to name the object-level promise before selecting a topology.

The Initial Model: More Replicas Mean More Safety

The first reasonable design says: send every write to three replicas. If one replica is slow, return after the other two respond. That improves durability in many situations, but it does not say what concurrent writers mean.

Imagine a partition between Europe and North America. An agent in each region tries to reserve seat/AB123/17A. If both local replica groups can accept the request independently, each can return success without knowing about the other. Replication will eventually reveal the disagreement, but there is no safe merge rule for “two people own one seat.”

The same pattern is not automatically bad for the trip notes. Two independent additions can be represented as two edits and merged. The last-viewed value may even accept a policy such as replacing an older value, if the product explicitly accepts the loss of the older update.

So “replicate to several nodes” is incomplete. It tells us how many copies participate, but not where authority lives or what happens when copies disagree. The stronger model is to move coordination deliberately to the point where its cost is acceptable.

The Better Model: Choose Where Disagreement Is Paid For

Plain meaning: A replication design decides whether one writer prevents conflicts early, several writers accept them and reconcile later, or replicas use quorum and version evidence without a fixed writer.

In this scenario: Use a tightly controlled write path for the seat, allow regional note writes only with an explicit merge rule, and use a leaderless-style replica set only for the low-cost last-viewed value if its reconciliation behavior is acceptable.

Technical names: These choices are commonly called primary-backup, multi-leader, and leaderless replication.

primary-backup: one normal writer sequences updates
multi-leader:   several normal writers accept updates, then exchange them
leaderless:     replicas and configured quorum/version rules decide participation

None of these labels alone guarantees a particular consistency level, latency, or failure behavior. Reply rules, replication mode, read routing, membership, versioning, and application semantics still matter. The label is useful because it reveals the first place to look for authority and reconciliation.

Worked Design: Three Paths Through the Same Product

Path 1: Primary-backup pays before confirming the seat

For seat/AB123/17A, designate one normal write authority for the seat's partition. A client in Europe may reach it directly or through a regional gateway.

traveler request
  -> seat primary checks “is 17A already held?”
  -> primary records and replicates the decision
  -> reply follows the system's configured acknowledgement policy

The key benefit is not that the primary has magic knowledge. It is that normal conflicting requests meet at the same authority boundary before the service promises a hold. The primary can order “hold for Maya” and “hold for Luis” and reject or queue the later one.

The acknowledgement policy still matters. An implementation might wait for a backup, for a quorum, or for a weaker asynchronous condition before replying. Those choices change durability and failover risk. They do not change the central design move: normal writes do not race through two independent authorities.

If the primary becomes unreachable, the system must either fail over through a safe mechanism or temporarily refuse new seat holds. That is an availability cost chosen to avoid inventing a merge for double booking.

Path 2: Multi-leader accepts local note edits, then owns the conflict

For trip/481/notes, Europe and North America can each accept a local write. The regional leaders later replicate accepted edits to each other.

Europe:        add “passport check”
North America: add “airport transfer”
                 \             /
                  \-- merge ---/

If the notes model records independent additions, the product can merge both changes. Local writing is attractive because neither traveler has to wait for an inter-region round trip.

Now change the edit: both travelers replace the same itinerary title with different text while the regions are partitioned. Both writes were locally legitimate. The replication protocol cannot learn the intended product answer from topology alone. The design needs a stated rule: show both versions, pick a defined winner, merge structured fields, or ask a user to resolve it.

This is the real cost of multi-leader. It moves coordination out of the immediate write path and into conflict detection, reconciliation, and user experience. A system that quietly overwrites a conflict has still chosen a merge rule; it has merely hidden it.

Path 3: Leaderless moves evidence and repair into the replica set

For user/92/last-viewed, suppose a leaderless-style store places the key on three replicas. The following values are illustrative:

N = 3 replicas
W = 2 acknowledgements required for this write
R = 2 replicas consulted for this read

With ordinary fixed replica sets, W + R > N means a write quorum and a read quorum intersect. That shared replica can carry useful version evidence. It does not, by itself, settle every real implementation question: a system may use temporary substitutes, asynchronous repair, configurable consistency levels, or concurrent versions.

Here a request writes last-viewed = hotel-7 to two replicas while the third is slow. A later read consults two replicas. If it sees an older value and a newer version, the store or client can use the version information to select or reconcile the result and repair the stale replica later.

write: replica A <- hotel-7, replica B <- hotel-7
read:  replica B -> hotel-7, replica C -> hotel-3
then:  return/reconcile according to version policy; repair C when appropriate

This path has no permanent primary for the key. It has not escaped coordination. It relies on quorum choice, version metadata, repair, and a definition of what to do with concurrent values. Dynamo is a classic example of this design space: it used versioning and application-assisted conflict resolution to favor high availability for selected workloads.

So far: the three models move the same burden rather than deleting it. Primary-backup constrains writers early. Multi-leader permits several writers and makes conflict handling a first-class product feature. Leaderless designs make quorum evidence, versions, and repair part of the data contract.

The Design Review: Start With the Conflict, Not the Product Name

Before this lesson, the team might ask, “Which database gives us the most replicas and the lowest latency?” After it, the more useful review starts with the object:

1. What conflicting writes can occur for this key or aggregate?
2. What is the cost of accepting both while disconnected?
3. Is there a correct merge rule that the product can explain?
4. Must a write wait for one authority, a quorum, or a remote region?
5. What evidence will a reader or repair job need later?

For the seat, a single authority or consensus-backed boundary is usually the better fit because the conflict is not safely mergeable. For notes, multi-leader is a situated preference when local writes matter and the product has a credible merge and audit story. For last-viewed, leaderless replication can be a good fit when availability and locality matter more than an immediately uniform view.

This does not mean a real service should run three unrelated databases by default. It means a broad “one policy everywhere” decision needs a reason. A single system can also offer different consistency and conflict policies, while a service boundary can separate data with fundamentally different promises.

Consequences, Trade-offs, and Limits

The trade-off is where disagreement becomes visible and expensive.

Model What improves What costs more Signal near the boundary
Primary-backup simple normal write authority primary latency, failover safety, and availability during loss primary queueing, replica lag, or repeated failover
Multi-leader local writes and regional autonomy conflict detection, merge design, and reconciliation delay replication backlog, divergent versions, or rising conflict rate
Leaderless no fixed primary and configurable replica participation quorum tuning, version handling, repair, and stale/conflicting reads repair backlog, inconsistent replicas, or quorum timeouts

These are not scorecards where one row always wins. A primary can be geographically close to its writers. A multi-leader system can reject some conflicts. A leaderless system can be configured for stronger reads. The table identifies the questions that remain after the label.

The boundary is especially important across objects. A leaderless last-viewed write and a primary-backed seat hold do not automatically form one transaction just because they live in the same product. Cross-object invariants need their own authority and commit boundary. Replication topology alone does not create it.

Common Confusions

Confusion: “Multi-leader is primary-backup with more primaries, so it is simply faster.”

Why it is tempting: Local writes can avoid a remote round trip.

Better model: It creates several legitimate writers. The saved time moves into conflict detection and reconciliation whenever their writes overlap.

Confusion: “Leaderless means no coordination.”

Why it is tempting: There is no named primary.

Better model: Coordination moves into quorum thresholds, version evidence, anti-entropy, and merge rules.

Confusion: “A leaderless quorum always returns one obvious newest value.”

Why it is tempting: Quorum intersection sounds like a universal latest-value guarantee.

Better model: It provides useful overlap under its assumptions, but concurrent writes, temporary replicas, and chosen consistency settings can still require version-aware reconciliation.

Check Your Understanding

Check: Two regions must accept seat/AB123/17A holds while disconnected. The product cannot ever confirm two holds. Is multi-leader a good default for this object?

Think first, then reveal.

Answer: No. It would allow two locally valid writers unless an additional shared authority boundary prevented that. Because the conflict cannot be merged after confirmation, a controlled writer or consensus-backed decision is the safer fit.

Check: A leaderless store reads two replicas and observes two concurrent versions of a trip-note title. Is “return whichever response arrived first” a valid conflict policy?

Think first, then reveal.

Answer: It is a policy only if the product explicitly accepts arbitrary data loss. Usually the system should expose or apply a defined version-aware merge rule; arrival time does not express user intent.

Practice: Review One Proposed Policy

A team proposes: “Use multi-leader replication for all travel data. It gives each region fast writes, and conflicts are rare.” Their application includes seats, collaborative notes, and recently viewed hotels.

Write a review. Identify one object that should not use this policy by default, one object for which it could be appropriate only with a named merge rule, and one signal that should make the team investigate whether its reconciliation cost is growing.

Model answer: Seats should not use multi-leader by default because independently confirmed holds are not safely mergeable; they need one controlled authority boundary. Collaborative notes can use it when the product defines how independent and conflicting edits are merged or presented. A growing cross-region replication backlog, more divergent versions, or an increasing conflict-resolution rate signals that the assumed reconciliation cost is no longer small. Recently viewed hotels may instead suit a leaderless or eventually reconciled path if the product accepts its freshness boundary.

Connections

The ZAB lesson showed why ZooKeeper pays for a leader-based, totally ordered write path: coordination state decides who may act. This lesson generalizes the choice. The more expensive a conflict is, the more deliberately a system should constrain where normal writes become authoritative.

The next lesson narrows from replication topology to distributed logs. A log is an ordered history, but its ordering scope depends on the authority and partitioning boundary beneath it.

Resources

Key Takeaways

PREVIOUS ZAB and Total Order Broadcast in Practice NEXT Distributed Logs and Ordering Guarantees