Guarantee Matrix Design Review
LESSON
Guarantee Matrix Design Review
By the end of this lesson, you will be able to...
turn a service endpoint into a row that names its invariant, authority, read or write path, recovery boundary, and evidence;
find a contradiction between an advertised guarantee and the path that actually serves it;
defend a deliberately weaker path when its freshness, fallback, and failure behavior are explicit.
Idea in one sentence: A guarantee matrix makes a replicated service reviewable by requiring every client-visible operation to name who decides, what a client may see, what happens under failure, and how the team will prove it.
Core Insight
What You Can Now See
Harbor Point is preparing a reservation service for desks in Madrid and New York. The service has shard leaders, followers, a remote replica, a search index, backups, and fault tests. That sounds reassuring, but it is not yet a contract.
The simple model is: “choose a highly consistent database, then place replicas wherever latency is good.” It works for a small system with one kind of request. It breaks as soon as the same service has a final reservation, a retry after a timeout, a dashboard, a cross-issuer search, and a regional-recovery promise. Those operations do not need the same answer, and a slogan cannot tell a caller which one it received.
The stronger model is a guarantee matrix. It starts with operations people can observe, rather than with a topology diagram. Each row joins seven decisions that must agree:
- the business invariant or question;
- the current authority for that decision;
- the acknowledgement or read guarantee;
- the normal serving path and an allowed freshness budget;
- the fallback when that path cannot prove the guarantee;
- the recovery boundary; and
- the signal or test that makes the row falsifiable.
The Concepts Together
For issuer MUNI-77, every accepted reservation consumes the same remaining exposure. That is the narrow invariant. A reservation can be final only if one current authority orders the exposure update and the reservation together. A local gateway may receive the HTTP request, but receipt is not authority.
Other surfaces can be weaker. A search index can be useful before it is current. A dashboard can be fast before it is authoritative. The mistake is not using those paths; the mistake is silently treating their answers as proof of a just-completed reservation.
| Operation | Authority and contract | Normal path | When it cannot prove the contract | Evidence |
|---|---|---|---|---|
POST /reservations |
The issuer-shard leader orders issuer_exposure and the reservation together. A success is one accepted decision. |
Route by issuer_id to the current leader; wait for the configured durable acknowledgement. |
Fail or return an uncertain result that the caller resolves with the same token. | Commit position, leader generation, and a history checker that rejects duplicate tokens or lost acknowledged writes. |
GET /reservation-tokens/K7 |
The same durable token record maps one token to one outcome. | Read the authoritative shard, or a replica known to have replayed the relevant position. | Wait or route to the leader; never create another reservation just because the first call timed out. | Token record, replay position, and retry trace. |
GET /my-reservations after success |
The caller must see at least its observed commit. | A nearby follower may serve only after it reaches the caller's commit position. | Route to the leader or say that the answer is not yet available. | Caller-required position minus follower replay position. |
| Compliance search | Base reservations remain authoritative; search is a derived discovery view with a published freshness budget. | Query the index with its watermark. | Mark degraded freshness or use an issuer-backed investigation path. | Index lag and an alert on the published budget. |
| Regional recovery | The promoted site can expose only its durable recovered prefix. The stated RPO bounds the permitted loss window. | Remote asynchronous replica plus tested backup and log-recovery path. | Do not claim readiness when lag, log continuity, or restore evidence is outside the stated budget. | Durable replay position, backup/restore proof, and promotion exercise. |
Common Confusions
“Every row should be linearizable.”
That would simplify the wording, but it may impose unnecessary latency and coordination on search and reporting. The better rule is narrower: make the final business decision strong enough for its invariant, then name exactly what weaker paths may return.
“A healthy follower can serve a session-sensitive read.”
Health says that the follower process is up. It does not prove that it has replayed commit 842 from this caller. If the follower is at 836, the missing state is visible: it lacks six positions. The gateway needs a required position from the caller, then it can wait, route to the leader, or return a response that honestly permits staleness.
“An RPO is a backup-team detail.”
An RPO changes what a completed local write means after a regional loss. If remote replication is asynchronous, a local acknowledgement may survive a node failure but still be absent from the remote recovered prefix. That may be acceptable under an explicit five-second RPO; it is not compatible with a promise that every success survives losing the whole home region. The boundary becomes operational when remote replay lag, archive continuity, or restore rehearsal leaves the budget.
Synthesis Example
At 09:00, the New York leader for MUNI-77 has 1 million euros of remaining exposure. Madrid sends POST /reservations with token K7 for the last 1 million euros. The leader commits the exposure update and token record at position 842, and returns success. The Madrid follower is still at 836; the search index watermark is also before 842.
The matrix tells each actor what to do:
- The write succeeds only through the issuer leader, so a second desk cannot independently confirm the same last capacity.
- If the client refreshes immediately, the local follower cannot yet serve a read-your-writes response. The gateway waits for
842or routes to the leader. It does not interpret a healthy follower as current. - Compliance search may still omit
K7. That is not evidence that the reservation vanished because this row promised bounded freshness, not immediate authority. The response must show the index watermark or use the stronger lookup path. - If the client saw a timeout instead of success, it repeats
K7or asks for its status. The token converts an ambiguous network outcome into one resolvable business outcome. - If New York later fails, promotion starts from Madrid's durable recovered prefix and publishes a new shard generation before accepting writes. A reservation beyond that prefix is either within the declared RPO or a broken recovery promise; the matrix prevents calling both outcomes “successful failover.”
This is the useful review move: follow one business action across write, read, retry, derived data, and recovery. The topology is not judged by its number of replicas. It is judged by whether every observed result still matches its row.
Retrieval Check
Check: A team says, “Compliance search is linearizable because the base table is strongly replicated.” The index has a 12-second ingestion delay, while the public search contract says results appear within five seconds. Which matrix columns contradict each other?
Think first, then reveal.
Answer: The serving path and freshness budget contradict the claimed guarantee. A strongly protected base record does not automatically make a derived index current. The team must improve the index path, increase the published budget, or provide a stronger fallback; relabeling the base table does not repair the row.
Check: The remote replica is 18 seconds behind, but the service advertises a five-second RPO. Can the on-call team call a failover drill successful because promotion works?
Answer: No. Promotion may be mechanically possible, yet the recovery row is already outside its stated loss budget. The signal says the promise is false until the lag returns within budget or the advertised RPO changes.
Transfer Challenge
A ticket platform sells the last ten seats for a concert. It proposes local writes in two regions, an asynchronous replica, a seat map served from a cache, and a global search index. Create four matrix rows: purchase, order-status retry, seat-map read after purchase, and search.
A good answer should mention:
- the concert inventory or performance as the authority boundary for the purchase, not the buyer's nearest region;
- an idempotency key and a durable status record for an ambiguous payment or order request;
- a required commit position, leader fallback, or explicit stale label for the seat map immediately after purchase;
- a visible freshness budget for search; and
- a recovery promise, fencing step, and operational evidence before claiming that a regional loss preserves paid orders.
What Comes Next
The next review lesson attacks this matrix with partitions, timeouts, stale reads, membership changes, backlog, restore, and regional loss. First make each row coherent. Then ask whether it still holds when several mechanisms are under pressure at once.
Resources
- [DOC] DynamoDB read consistency — Focus: Compare table, index, and multi-Region consistency contracts instead of assigning one label to a whole service.
- [DOC] DynamoDB global secondary indexes — Focus: See why an efficient alternate query path is a separate data surface with its own synchronization behavior.
- [DOC] Spanner: TrueTime and external consistency — Focus: Contrast strong and stale read contracts in a geo-distributed system.
- [PAPER] Spanner: Google's Globally-Distributed Database — Focus: Study the design trade-offs behind global transaction ordering.
- [DOC] Jepsen Analyses — Focus: Connect a matrix row to a workload and fault history that can reject its promise.
Key Takeaways
- A guarantee matrix begins with client-visible operations and business invariants, not a preferred database label.
- Authority, acknowledgement, serving path, freshness, fallback, recovery, and evidence must agree within every row.
- A weak path is a valid design choice only when its boundary is visible and callers have an honest fallback.
- The next step is not adding more components; it is testing whether the matrix remains true during compound failure.