Replicated Data Service Capstone
LESSON
Replicated Data Service Capstone
By the end of this lesson, you will be able to...
design a replicated data service whose write, read, recovery, and operational paths support named business guarantees;
defend the design with a guarantee matrix, two traces, and observable readiness evidence;
identify the unsupported promise when a failure or product request exceeds the proposed architecture.
Idea in one sentence: A replicated data service is ready to defend when every client-visible result can be traced to one authority, one stated freshness or recovery boundary, and evidence that the promise still holds under failure.
Core Insight
Harbor Point can add replicas in Madrid and New York, yet still break its reservation promise if two paths disagree about who decides the last available exposure for MUNI-77. The difficult part is keeping every client-visible promise aligned with the path that can actually support it.
The Scenario
Harbor Point is launching a reservation service for municipal-bond exposure. Desks in Madrid and New York submit reservations. Each issuer has a limited amount of exposure, and an accepted reservation must not over-allocate it. Compliance needs to search reservations across issuers. Risk wants dashboards near each desk. The service must survive loss of one region, and support must resolve a request that timed out without guessing whether it succeeded.
The tempting design is simple: give both regions a writable replica, copy data between them, and call the service “strongly consistent.” That model works for a record that can be independently edited and reconciled later. It breaks at the last available euro of an issuer's exposure. If Madrid and New York each confirm the final reservation from separate copies, a later merge cannot make both earlier confirmations true.
This capstone does not ask for a particular database product or a consensus implementation. It asks for a coherent contract. Use the mechanisms from the track only when they support an observable promise. The design below is a worked exemplar with illustrative numbers; your final challenge asks you to produce and defend the same kind of dossier.
Constraints
Assume these constraints for the worked exemplar:
- An issuer may have many reservations, but its remaining exposure may never fall below zero.
- A confirmation shown to a desk is final for that issuer, unless the request itself is reported as uncertain.
- A client that receives confirmation should be able to read its reservation immediately, even if it is routed to a nearby region.
- Compliance search may be behind the authoritative state for up to five seconds, but its response must make that freshness visible.
- The remote recovery copy may lose at most five seconds of acknowledged local work after loss of the home region. The recovery time objective is ten minutes.
- A delayed, partitioned, or returning former owner must not resume accepting writes after a new owner is active.
The last two targets are assumptions for this exercise, not properties supplied by a replica count. They become real only when the acknowledgement rule, remote durable progress, promotion runbook, and restore drills support them.
Design Goal
Deliver an architecture dossier that answers five questions for each important operation:
- What business state decides the operation?
- Who is the authority for that state now?
- What does success, a read result, or a timeout mean to the client?
- Which weaker path is allowed, and what happens when it cannot prove its promise?
- Which trace, signal, or test could show that the answer is false?
This is stronger than an architecture diagram alone. A diagram can show replicas. The dossier must say which replica may decide, which replica may only observe, and what a client can infer from the response.
Proposed Model
Harbor Point shards the authoritative reservation state by issuer_id. For the exercise, issuer MUNI-77 maps to shard 184, whose current home region is New York. A leader owns the issuer decision at one configuration generation. A same-region follower participates in the normal durable acknowledgement rule. A Madrid replica receives the committed log asynchronously for recovery. CDC consumers build search, dashboard, and reporting views from committed changes.
Madrid or New York gateway
|
v
issuer_id -> shard map -> shard 184, generation 31
|
v
New York leader: issuer exposure + reservations + token outcomes
|
+--> required same-region durable follower
|
+--> asynchronous Madrid recovery replica
|
+--> committed outbox / log -> search, dashboards, reports
The key is not that New York is special. It is that one current owner can serialize the MUNI-77 exposure decision. A later rebalancing or regional promotion must publish a new generation and fence the old one before it serves writes. This consumes the previous lessons on sharding, membership, and safe ownership without reimplementing their consensus machinery.
Guarantee matrix
| Operation | Authority and public promise | Serving path and fallback | Evidence |
|---|---|---|---|
POST /reservations |
Shard 184 leader atomically decides exposure, reservation, and idempotency token. confirmed means this commit met the configured local durability rule. |
Route by current shard generation. If the authority or durability rule cannot be proven, fail or return uncertain, not a weaker confirmation. |
Commit position, leader generation, durable-ack trace, and a checker for duplicate token or lost completed result. |
GET /reservation-tokens/K7 |
The token record has one durable outcome. A timeout is not treated as a failed reservation. | Read the authority, or a replica that has replayed the token's position; otherwise wait or route to the leader. | Token outcome, replay position, and retry history. |
GET /my-reservations after success |
The caller sees at least its observed commit. | A local follower serves only after its replay position reaches the caller's commit token. Otherwise route or wait within a stated budget. | Required position minus follower replay position. |
| Compliance search | Base reservations are authority; the index is a discovery view with freshness at most five seconds. | Return the index watermark. When late, label the response degraded or use an issuer-backed investigation path. | Index lag, watermark age, and an alert on the five-second budget. |
| Regional recovery | Promotion exposes only the remote durable prefix; loss must fit the five-second RPO. | Stop claiming readiness when remote lag, retained-log continuity, or restore proof exceeds the target. Fence the old generation before writes resume. | Remote durable position, RPO lag, restore age, promotion record, and a regional-loss exercise. |
The matrix deliberately does not say “the system is linearizable.” That sentence is too coarse. It says which operation needs a final decision, which reads need a session guarantee, which data may lag, and which success response survives which failure domain. This is a design preference under Harbor Point's constraints. A product that requires every confirmation to survive regional loss would need remote durability in the acknowledgement rule, with the resulting latency and availability cost.
Walkthrough
The following trace uses invented times and positions. It is a teaching model, not an observed production incident.
Normal write and session read
At 09:00, MUNI-77 has €1m of remaining exposure. Madrid submits POST /reservations for €1m with token K7. The current shard map says MUNI-77 -> 184 -> New York, generation 31.
09:00:00 Madrid gateway route K7 to shard 184, generation 31
09:00:01 NY leader checks exposure = €1m; no outcome exists for K7
09:00:01 NY leader commits exposure = €0, reservation R-91, token K7 -> R-91
09:00:02 local follower durably confirms position 842
09:00:02 gateway returns confirmed(R-91, commit=842, generation=31)
09:00:03 Madrid follower replay position = 836
09:00:03 client asks GET /my-reservations, required position 842
09:00:03 gateway routes to leader or waits; follower 836 does not answer as current
The first reasonable but wrong response at the last line is, “the follower is healthy, so it can answer.” Health is a process signal. It does not show that the follower contains this caller's commit. The required position makes the missing state visible. Once the follower reaches 842, it can serve this session read. Before then, the design pays a small latency cost for honesty.
The trace also shows why the token belongs in the same authoritative decision. If the response at 09:00:02 is lost, the client does not submit a fresh business action. It asks about K7 or repeats K7; the leader returns R-91. One token maps to one business outcome across timeout and retry.
Derived read and recovery boundary
The outbox for position 842 reaches compliance search at 09:00:05. Until then, a search response can omit R-91, but it must expose a watermark older than 842. The omission is permitted by that row's bounded-freshness contract; it is not permitted as proof that the reservation failed.
At 09:04, Madrid has durably replayed through position 850. At 09:04:02, New York becomes unreachable. The runbook first prevents the old generation from serving, then promotes Madrid at generation 32 from its durable prefix. A reservation at 854 cannot appear after promotion because it never reached the promoted prefix.
This has two possible interpretations:
- If the public confirmation rule promised survival of every acknowledged write through regional loss, the design is wrong. The acknowledgement path must wait for remote durability or change the promise.
- If the service explicitly promised an RPO of five seconds and the lost writes fall inside a measured five-second window, the loss may match the stated recovery contract. It is still an event that needs token resolution, communication, and evidence.
The meaningful boundary is not “replication happened.” It is the durable prefix that the promoted authority can actually prove.
So far: the write path protects the invariant, the session path prevents a stale follower from impersonating current state, the derived path names its lag, and the promotion path names its loss boundary. These are separate mechanisms serving one contract.
Failure Review
Use this review before calling the design ready.
| Pressure | What must remain true | Failure that reveals an unsupported promise |
|---|---|---|
| Two desks submit the final available exposure. | At most one receives confirmed; the other sees insufficient exposure or an honest uncertain result. |
Two regional owners independently confirm the final amount, then attempt to merge later. |
| A response times out after the leader commits. | Repeating the same token returns the existing outcome, not a second reservation. | The retry creates R-92 because token status was not durable with the decision. |
| A follower lags after a completed write. | A session-sensitive read waits, routes, or returns a contractually stale answer. | A healthy follower returns an older state as if it were read-your-writes. |
| CDC or the index falls behind. | Search exposes watermark age and does not decide issuer exposure. | A stale index is used to approve or reject a reservation. |
| Home region is lost. | New writes use one fenced generation; recovery loss is no larger than the declared RPO. | A former leader accepts writes after promotion, or the measured remote lag exceeds the advertised RPO. |
| A restore is required. | The team can reconstruct authoritative state to the stated point and then rebuild derived views. | A replica exists but no restore proof, retained log, or replay plan supports the recovery target. |
A fault exercise should record client invocation and completion, request token, shard generation, commit/replay positions, and topology events. A green cluster after the exercise is not sufficient evidence. The checker needs to reject a history with an impossible confirmation, duplicate token, forbidden session read, or recovery result outside the stated contract.
Trade-offs
Harbor Point gets one final exposure decision per issuer. It pays cross-region latency when a desk is far from that issuer's leader. It also pays for token state, follower progress checks, remote replication, log retention, monitoring, and rehearsed recovery.
In return, compliance search and dashboards can avoid the decisive write path. They are scalable because they are derived, not because they somehow become authoritative. The cost is that their freshness must be measured and displayed, and their operators need a fallback for investigations.
The design can still fail. A five-second RPO does not make the last five seconds impossible to lose in the specified disaster. A token does not make a downstream email exactly-once. A fenced promotion cannot repair a false shard map before it is detected. These are boundaries, not reasons to abandon the model. They tell the team which promise to narrow, which mechanism to strengthen, or which signal to gate before release.
Evidence and Readiness
Before launch, Harbor Point should have an answer for every cell in this compact dossier:
| Readiness question | Acceptable evidence |
|---|---|
| Does each irreversible action have one authority boundary? | Guarantee matrix and shard-key review showing that the invariant is local to the chosen owner. |
| Does acknowledgement mean what product and operations think it means? | Write trace with the exact durable condition and a documented treatment of timeout. |
| Can clients obtain the promised read after a write? | Commit-token propagation test; follower replay metric; leader fallback behavior. |
| Are derived surfaces honest about staleness? | Watermark in the response, lag alert, and a tested stronger lookup path. |
| Is the recovery promise real? | Remote durable-progress history, retained-log or backup continuity, timed restore, and fenced-promotion drill. |
| Can the team falsify the central claims? | A recorded workload, fault schedule, checker, and retained failure artifacts. |
These are not generic production checkboxes. They correspond directly to the matrix. If a row has no evidence, it is an aspiration rather than a guarantee. If an alert has no contract threshold, it may be useful for debugging but cannot tell the team whether the service is meeting its promise.
Final Challenge
Create an architecture dossier for one service with an invariant that cannot be repaired after success. You may use Harbor Point, an online ticket service, inventory allocation, or another nearby domain. Keep the same level of specificity; do not solve the challenge with “use a strongly consistent database.”
Your dossier must contain:
- a one-sentence user-visible invariant and the authority boundary that decides it;
- a guarantee matrix with at least four rows: decisive write, timeout/status lookup, session-sensitive read, and one derived or recovery path;
- a topology and routing sketch that shows the current owner, durability path, remote recovery copy, and derived consumers;
- a normal-operation trace that includes a write, returned commit token or equivalent evidence, and a safe immediate read;
- a failure trace involving either timeout plus retry, stale read, rebalancing, or regional promotion;
- two explicit trade-offs and the product constraints that justify them;
- a signal, test, or runbook step for every high-value promise; and
- one limitation the design does not solve.
Readiness rubric:
- Correctness: every irreversible decision has one authority at a time; retries and promotion cannot silently create a second effect.
- Clarity: the write acknowledgement, each read guarantee, and the RPO/RTO language can be understood without product slogans.
- Operability: lag, required read position, topology generation, and recovery progress have named evidence and a response when outside budget.
- Trade-offs: weaker reads or recovery loss are explicit, bounded, and connected to the latency or availability they buy.
- Scope: the dossier uses the track's mechanisms to support a contract; it does not wander into database-engine internals, full consensus proofs, or a vendor deployment guide.
Resources
- [PAPER] Spanner: Google's Globally-Distributed Database — Focus: Compare a globally ordered transaction design with the capstone's deliberately narrow authority boundary.
- [DOC] Spanner: TrueTime and external consistency — Focus: Contrast an explicit strong-read contract with bounded stale reads and their different costs.
- [PAPER] In Search of an Understandable Consensus Algorithm — Focus: Revisit leader ownership, committed log prefixes, and safe membership change as concepts consumed by this design.
- [DOC] PostgreSQL logical decoding concepts — Focus: See a concrete log-to-consumer interface, including replay, duplicate-delivery, and retention boundaries.
- [DOC] Jepsen Analyses — Focus: Find examples of client histories and faults that make a claimed guarantee testable.
Key Takeaways
- A capstone design begins with client-visible invariants, then selects authority, acknowledgements, reads, recovery, and evidence to preserve them.
- A healthy replica, a CDC projection, and a promoted node each answer different questions; none is authority by default.
- Every weaker path needs a named freshness or recovery boundary and a fallback when it cannot prove that boundary.
- The final proof of a replicated service is a falsifiable dossier: matrix, traces, signals, runbook, and fault test all describe the same contract.