Leaderless Replication, Sloppy Quorums, and Hinted Handoff
LESSON
Leaderless Replication, Sloppy Quorums, and Hinted Handoff
By the end of this lesson, you will be able to...
Trace a sloppy-quorum write from an unavailable home replica to a fallback node and back.
Explain why the usual
R + W > Nargument does not by itself promise a fresh read after fallback placement.Decide when a repair-backed availability trade-off fits an API, and name the repair signals it needs.
Idea in one sentence: A sloppy quorum keeps accepting writes by storing them on reachable fallback nodes, then relies on hinted handoff and broader repair to restore the intended replica placement.
Core Insight
Harbor Point stores a trader's watchlist in a leaderless replicated store. The key watchlist:trader-17 normally belongs to three home replicas: A, B, and C. The user is editing preferences during a zone incident, and A and B cannot be reached.
The obvious response is to reject the update: its normal homes are unavailable. That keeps placement tidy, but it turns a local infrastructure fault into a user-facing failure. A leaderless, availability-first store can make a different choice: send temporary copies to reachable nodes farther down the preference list, such as D and E.
That choice is called a sloppy quorum. It is useful only when the API can tolerate temporary placement drift and later convergence. A watchlist can usually do that. A reservation approval or a change to a legal trading limit usually cannot: those actions need one clear authority decision, not a write that may be temporarily out of its normal place.
The Situation
The previous lesson used quorum arithmetic for one key:
N = intended replicas for the key
W = write acknowledgements required
R = read responses required
R + W > N
For the home set A, B, C, let N=3, W=2, and R=2. If both operations choose only from A, B, C, a successful write and a successful read must share at least one home replica.
strict write: A, B
strict read: B, C
^ overlap
This is a useful proof, but it has a condition: both sets come from the same fixed family of replicas. It does not say that every two groups of two nodes in a larger cluster overlap. The phrase “quorum read” is not a freshness spell; the candidate set matters.
The Initial Model
It is tempting to say, “We use W=2 and R=2, so every later quorum read will find the write.” That model works while the write and read use the key's home set.
It becomes insufficient when a failure changes the placement.
home replicas for K: A B C
extended preference list: A B C D E
A and B are unreachable.
If the service refuses to use D or E, it cannot collect enough home acknowledgements. The update fails even though several storage nodes are healthy. If the service accepts fallbacks, it gains availability but must stop claiming that simple home-set overlap proves the next read is fresh.
The Mechanism Step by Step
Plain meaning:
Store the update on healthy temporary holders now, and remember where the update belongs when the normal holder returns.
In this scenario:
DandEtake copies that would normally have gone toAandB. Their metadata records the intended home replicas.
Technical names:
Choosing reachable nodes from an extended preference list is a sloppy quorum. The temporary-delivery record and later replay are hinted handoff.
The following trace uses illustrative node names and versions. It describes the mechanism, not a promise that every Dynamo-family product makes the identical routing choice.
Starting state
A, B, C hold watchlist version v7
home replicas A and B are unreachable
C, D, E are reachable
Step 1: coordinator receives PUT K = v8
It walks the preference list until it finds reachable nodes.
It sends v8 to C, D, and E.
Step 2: fallback nodes record intended placement
C stores v8 as a home replica.
D stores v8 with hint "deliver a copy to A".
E stores v8 with hint "deliver a copy to B".
Step 3: acknowledge the write
Once the configured write condition is met, the API can return success.
The success means v8 is durably accepted under this availability policy.
It does not mean all home replicas already have v8.
Step 4: recovery and handoff
When A and B become reachable, D and E scan their stored hints.
D sends v8 to A; E sends v8 to B.
After confirmed delivery, each fallback may remove its temporary copy.
Why does the fallback wait for confirmation? A timeout after sending v8 does not reveal whether A stored the copy. The handoff request therefore needs an idempotent identity or version evidence: retrying delivery must converge on one version, not create a second edit. Only a confirmed, reconciled delivery lets D safely retire its temporary responsibility. If the fallback dies first, another repair path must detect the missing home copy.
The original Dynamo paper describes this general pattern: operations use healthy nodes from a preference list, and a fallback replica carries metadata naming the intended recipient. The crucial detail is the last part of the trace. Fallback placement is not the final state. It creates delivery work.
A Worked Trace: Where the Overlap Story Breaks
Now apply pressure to the initial model. The write completed while A and B were down. Before hinted handoff finishes, C becomes unreachable during another short network problem. A and B have recovered but still hold v7.
after sloppy write: C=v8, D=v8(hint->A), E=v8(hint->B)
after partial recovery: A=v7, B=v7, C=unreachable
an illustrative home-only read: A, B
returned versions: v7, v7
The read has two responses. The old arithmetic still says 2 + 2 > 3. Yet that arithmetic was about a fixed home set, while the successful write's acknowledgements could have included C and a fallback. The read has no overlap with those copies in this trace.
This does not prove that every sloppy-quorum implementation returns stale data. A system can route reads through its expanded preference list, compare versions, wait, or expose a weaker contract. The lesson is narrower and more useful: after fallback placement, freshness depends on the actual read routing, version reconciliation, and repair progress. It does not follow from R + W > N alone.
So far: sloppy quorum changes an availability failure into a repair obligation. The write may be durable and still be absent from the normal homes for a while.
What This Changes for the API
Harbor Point can make the trade-off explicit instead of giving every key the same behavior.
| Operation | A reasonable policy | Why |
|---|---|---|
| Save a watchlist preference | Sloppy quorum plus repair | User progress matters; a temporary stale read can be repaired or reconciled |
| Save a dashboard layout | Often similar | The UI can show a freshness boundary and recover later |
| Approve a reservation | Refuse when authority is unclear | Two independently accepted approvals can violate the business invariant |
| Change an issuer limit | Use an authoritative, strongly coordinated path | Stale or conflicting placement has a high financial and audit cost |
This is a situated design preference, not a database taxonomy. Sloppy quorum is a good fit when all of these are true: the value can converge later, the product has a conflict rule, the caller can tolerate a repair window, and the operators can observe and pay the repair work. It is a poor fit when success must immediately establish a unique, externally visible authority fact.
Cost, Limits, and Signals
This mechanism improves write availability during temporary node or network failure. It costs extra storage, replay traffic, metadata, and operational attention. It also does not solve concurrent-write conflicts: if two clients make incompatible changes, hinted handoff delivers copies but does not decide their business meaning.
The boundary is visible in the repair backlog.
| Signal | What it reveals |
|---|---|
| Fallback-write rate | How often normal placement is unavailable |
| Oldest undelivered hint | How long temporary placement has lasted |
| Hint queue size and replay failures | Whether handoff can catch up after recovery |
| Replica version mismatch rate | Whether reads or background repair still find drift |
| Repair bytes and latency | The cost that availability shifted into the recovery path |
Hinted handoff works best for temporary failures. A fallback node can fail before delivery, and a long outage can make the hint queue large or unreliable. That is why the next lesson adds read repair and anti-entropy: hinted handoff is targeted delivery, not a complete convergence system.
Common Confusions
Confusion: A sloppy quorum has the same read-freshness proof as a strict quorum.
Why it is tempting: Both may use the same values for N, R, and W.
Better model: Strict overlap is about operations choosing from one fixed replica family. Fallback placement changes the candidate sets, so routing and repair determine what a later read can find.
Confusion: A hint is a copy that can be kept forever.
Why it is tempting: The fallback node already has the new value.
Better model: The hint represents temporary custody and a delivery obligation. Keeping it forever wastes capacity and leaves the intended replica set incomplete.
Confusion: Hinted handoff resolves conflicts.
Why it is tempting: It eventually moves an update home.
Better model: Handoff moves versions. A separate version and conflict policy decides how concurrent updates relate or merge.
Check Your Understanding
Check: A, B, and C are home replicas. A write succeeds on C and fallback D while A and B are down. Later, a read contacts only recovered A and B, before handoff completes. What is missing from the claim “R+W>N, so the read must see the write”?
Think first, then reveal.
Answer: The claim assumes both operations selected replicas from the same fixed home set. The successful write used a fallback, while the read used only stale homes; their actual response sets need not overlap.
Check: Which metric most directly tells you that temporary fallback placement is lasting too long?
Answer: The age of the oldest undelivered hint. It measures the time between accepting temporary custody and restoring intended placement.
Practice: Review a Write Contract
An incident-notes service wants to accept note edits during a zone outage. Notes are visible to the author immediately, but an audit export must not silently omit an acknowledged edit. The team proposes sloppy quorum for the write path.
Write a short contract for the successful write and the audit export. Then name one required repair control and one case where the API should return a clear failure instead of accepting a fallback write.
A good answer should mention:
- What write success means: durable acceptance under fallback placement, not immediate delivery to every home replica.
- How the author gets a token-aware or version-aware follow-up read rather than an unexplained stale response.
- How the audit export detects a pending or missing version, waits, uses an authoritative path, or clearly marks the result incomplete.
- An age or backlog objective for undelivered hints plus repair beyond handoff for missed delivery.
- A refusal for an operation that cannot tolerate ambiguous authority or unresolved conflict.
Connections
- Quorum Reads, Writes, and Tunable Consistency establishes the fixed-home-set overlap argument that this lesson qualifies under failure-time fallback placement.
- Read Repair, Anti-Entropy, and Merkle Divergence Checks explains how hot and cold data converge when hinted handoff is delayed or misses a replica.
Resources
- [PAPER] Dynamo: Amazon's Highly Available Key-value Store — Focus: Read the failure-handling section for sloppy quorum, hinted handoff, and the limits that require anti-entropy.
- [DOC] Riak KV Glossary — Focus: Compare its concise sloppy-quorum and read-repair definitions with the worked trace.
- [BOOK] Designing Data-Intensive Applications — Focus: Relate leaderless replication, fallback writes, repair, and conflict resolution to an application's actual promises.
Key Takeaways
- Sloppy quorum accepts a write on reachable fallback nodes when the intended homes are unavailable.
R + W > Nonly gives its familiar overlap argument when reads and writes choose from the same fixed replica set.- Hinted handoff records temporary custody and tries to return data to its intended homes after recovery.
- Fallback placement buys availability but creates repair debt, version-reconciliation work, and a freshness boundary that the API must state.