Replication Failure Mode Check

LESSON

Consistency and Replication

023 30 min advanced REVIEW

Replication Failure Mode Check

By the end of this lesson, you will be able to...

  • review a replicated service by connecting ownership, read guarantees, and failover promises;

  • classify a stale read, an ambiguous write, and a regional-loss outcome against the right contract;

  • identify the missing control when an attractive multi-region design makes an unsupported promise.

Idea in one sentence: A replicated-data design is coherent only when the writer, every read path, and the failover runbook tell the same story about one business action.

Core Insight

What You Can Now See

Harbor Point wants reservations from Madrid and New York to feel local. The first proposal sounds sensible: let each office write to its nearby replica, replicate between regions, and fail over if a region disappears.

That proposal does not yet define a service. For issuer MUNI-77, every accepted reservation consumes the same remaining exposure. Two regions cannot independently approve the last available amount and later merge the answer without deciding which confirmation to undo.

The question is: can every client-visible result be explained by the same authority, read, and recovery rules?

The simple model, “the closest replica handles the request,” works for data that can be merged later or for read-only copies. It breaks for a decision that must be final at the time a user sees confirmed. The stronger model assigns that decision to one current owner, then makes weaker paths and recovery boundaries explicit.

The Concepts Together

Use this compact contract for the rest of the review. The values are illustrative design choices, not claims about a particular database product.

Service question Authority or serving path Promise
Can MUNI-77 accept another reservation? Current leader for the issuer shard One ordered decision for the exposure record and reservation.
Did token K7 take effect? Same shard's durable token record One token resolves to one outcome, even after a timeout.
What did this caller just create? Leader, or follower caught up to the caller's commit token Read-your-writes.
What does compliance search show? CDC-fed index with a displayed watermark Bounded or eventual freshness, not authority.
What survives loss of the home region? Remote replica's durable replayed prefix Only the stated recovery point objective (RPO).

The rows must agree: the leader orders the write, the token resolves uncertainty, the commit token limits session reads, and remote replay bounds disaster loss. A search index remains useful but is not authority.

An authority boundary is the current component allowed to decide one invariant. Harbor Point can have many shards, but one issuer's exposure decision needs one owner at a time.

Common Confusions

“A write is local if the user is local.”

This is tempting because a Madrid gateway can accept the HTTP request quickly. But receiving a request is not the same as being entitled to order the issuer's next reservation. If MUNI-77 belongs to a New York shard, the Madrid gateway may route the decisive operation there.

The cost is a cross-region hop. The benefit is one ordered exposure decision. This fits work that cannot be repaired after confirmation; mergeable data may use another rule.

“A healthy follower can answer a strong read.”

Health says that a process is running. It does not say that the process has replayed a particular commit. After reserve(K7) completes at log position 842, a follower at 836 may be healthy and still be too old for a read-your-writes request.

The stronger rule is visible: carry the observed commit position with the session. Serve locally only when the follower has reached at least that position; otherwise route to the leader or wait within a defined budget. The signal is the difference between the caller's required position and the follower's replay position.

“A successful local commit means zero data loss after regional loss.”

Not when the remote replica is asynchronous. A same-region acknowledgement can honestly promise local durability and still leave a recent commit outside the remote durable prefix. That is why an RPO must name a time or position budget and a response when that budget is exceeded.

If the product promises that every 201 survives losing the home region, remote durability belongs in the acknowledgement rule. If it promises only an RPO of five seconds, then a very recent locally successful write can be within the disclosed loss window after a regional disaster. The API, runbook, and customer expectation must use the same version of this rule.

Synthesis Example: Three Events, One Contract

Assume the following illustrative topology. Shard 184 owns MUNI-77 in New York. It has a local synchronous follower, an asynchronous replica in Madrid, and a CDC search index. A write response includes the shard commit position.

Madrid gateway -> shard 184 leader in New York -> local sync follower
                                      |
                                      +-> async Madrid replica -> CDC search

Event A: two desks request the last capacity

At 09:00, both desks request a reservation for the final €1m of available exposure. They reach the New York leader in some order. The first transaction decreases the remaining amount to zero. The second observes zero and is rejected.

This is not because the network made the two requests non-concurrent. It is because the authority boundary serialized the exposure decision. If both regional replicas had independently approved a local copy of the same €1m, the system would need a later conflict rule. A later conflict rule cannot make both earlier confirmations true.

Event B: the confirmation is missing from a local screen

The Madrid client receives success for K7 at position 842, then refreshes the issuer dashboard. Its local follower has replayed only through 836; the CDC search index reports a watermark before 842 too.

The reservation did not disappear; both read paths are weaker than the caller needs. Use the caller's token to wait, route to the leader, or label the dashboard stale. Search remains useful for discovery, not as proof of a just-confirmed write.

Event C: the home region is lost

At 09:05, the remote replica has durably replayed through position 850. At 09:05:02, it is promoted after the New York region becomes unavailable. The runbook publishes a new shard generation before accepting traffic, so the old route cannot keep serving writes if it later returns.

The client can ask the promoted shard for K7 and avoid blind replay. But a write at 854 was not in the promoted prefix. That violates a zero-loss regional contract; it may be inside an explicit five-second RPO, which must have been monitored and communicated.

So far: authority prevents two final approvals for the same capacity; commit-aware reads prevent a stale replica from masquerading as current; and the remote replay prefix defines the recovery boundary. These are three parts of one contract, not three independent features.

A Review Method You Can Reuse

For every endpoint or workflow, ask these questions in order:

  1. What state decides the business rule, and who can order it now?
  2. Which read paths are allowed, and what version can each one prove?
  3. What result does a success response mean during a timeout, failover, or regional loss?
  4. How does the runbook fence the old owner and resolve an ambiguous client request?
  5. Which evidence proves the answer today? Examples include commit positions, replay lag, token lookup, leader generation, and a checked client history.

This method does not require every path to be linearizable. It requires every path to be named. The failure is claiming a stronger property than its path can support.

Retrieval Check

Check: A follower is healthy but has replayed position 836; the caller needs position 842. Can it serve the caller's read-your-writes request?

Think first, then reveal.

Answer: No. Health does not prove freshness. The gateway needs to wait until the follower reaches 842, route to the leader, or use a path whose contract admits a stale answer.

Check: After a regional failure, the promoted replica lacks an acknowledged write. Is that automatically a correctness defect?

Think first, then reveal.

Answer: Only if the acknowledgement contract promised survival across that regional failure. With an asynchronous remote copy and a stated RPO, the loss may be within the declared budget. It is still an operational event that needs an honest RPO, lag signal, and token-resolution path.

Transfer Challenge

An online ticket service sells the last ten seats for a concert. It proposes local writes in two regions, an asynchronous cross-region replica, a regional seat map, and a global search index. Review the proposal using the five questions above.

A good answer should mention:

What Comes Next

The capstone asks you to turn this review into a complete service design: choose the authority boundary, state the guarantees, show the data paths, name the operational evidence, and defend the trade-offs. This lesson's job is smaller but essential: reject a design whose pieces are individually plausible yet incompatible together.

Resources

Key Takeaways

  1. A region close to a client is not automatically the authority for that client's business decision.
  2. A healthy replica is not automatically fresh enough for a session-sensitive or strong read.
  3. An RPO is part of the meaning of a successful write during regional failure, not an operations footnote.
  4. A coherent design aligns shard ownership, read routes, token handling, monitoring, and promotion fencing.
PREVIOUS Guarantee Matrix Design Review NEXT Replicated Data Service Capstone