Failure Testing Replication Claims
LESSON
Failure Testing Replication Claims
By the end of this lesson, you will be able to...
turn a replicated-service promise into operations, faults, and a checker;
distinguish an acknowledged result, a definite failure, and an uncertain timeout in a test history;
diagnose what a failure-test anomaly says about fencing, retries, or a stated guarantee.
Idea in one sentence: A failure test earns confidence only when a checker can reject the client-visible histories that the promised contract forbids.
At 10:02, Harbor Point's reservation API returns 201 Created for request token K7. A few seconds later, the current leader is partitioned away from the other replicas and a new leader takes over. The cluster recovers. Its dashboard is green.
That is reassuring, but it does not answer the important question: after the 201, can a strong status read find exactly one reservation for K7? A smooth election proves that some recovery path worked. It does not prove that the acknowledgement, retry path, and reads all obeyed the API contract.
This distinction is the point of Jepsen-style failure testing. The test is not “add chaos while generating load.” It records what clients invoked and observed, injects a fault that threatens a named promise, and checks whether the resulting history could have happened in a system with that promise. Faults provide pressure; the history and checker provide the evidence.
Core Insight
The trade-off is deliberate: a precise workload and checker cost more to design than a generic failover drill, but they can reject a history that violates a business promise instead of merely showing that processes restarted.
Start With a Promise a Client Can Observe
Harbor Point first writes a small contract for its reservation endpoint:
POST /reservations with idempotency key K7 returns 201
-> one reservation for K7 exists durably
-> a later strong GET /reservation-tokens/K7 finds that reservation
-> retrying K7 never creates a second reservation
The exact durability boundary matters. If the product says 201 means only that one process accepted a request, the promise is weak. If it says 201 means the reservation survived a failover, the test may insist that a successful response remains discoverable after that failover. The harness must test the promise actually offered, rather than silently upgrading or weakening it.
The same discipline applies to reads. A dashboard allowed to lag may legitimately miss a fresh reservation. A strong token-status endpoint cannot. Giving the two paths different names prevents the checker from calling an intended stale read a bug, or worse, accepting a stale read on the path users rely on for confirmation.
The resulting operations should be small enough to model:
| Operation | Client result that matters | Invariant to check |
|---|---|---|
reserve(K) |
ok with reservation ID |
An acknowledged reservation is not later absent without a matching release. |
status(K) on the strong path |
the ID or absence | After reserve(K) returns ok, it returns that same ID. |
retry reserve(K) |
same ID or an explicit already-created result | One token has at most one business effect. |
release(K) |
completion | A later absence is valid only when the history includes a valid release. |
These are not performance assertions. “p95 stayed below 40 ms” may be useful context, but it cannot establish that an acknowledged action was not lost. A checker needs a rule about the state transition visible to a client.
Record Uncertainty Instead of Erasing It
Each client operation enters the history with an invocation and later receives one of three useful outcomes:
ok: the client received a successful response.fail: the client received a definite rejection, such as validation failure or a fenced-leader error before the request was accepted.info: the client timed out or lost its connection and cannot tell whether the operation took effect.
info is neither success nor proof of failure. Treating every timeout as a failure can hide a committed reservation. Treating every timeout as success can invent one. The safe test design keeps the uncertainty, then uses the idempotency key to resolve it later.
For example, a harness can issue status(K7) after the topology settles. If it finds one reservation, the earlier timeout may have hidden a commit. If it finds none, the client may safely make a new attempt with the same key. If it finds two IDs, the test has caught a duplicate effect. Generating a fresh key after every timeout would make that last anomaly much harder to see and would not match a safe production retry protocol.
This is also why client instrumentation belongs in the test, not just server logs. Server logs explain a suspected cause after a failure. The client-visible history establishes what the service promised and what it delivered. A correct system can have noisy logs; it cannot make an ok mean two incompatible things.
Work One Adversarial History
Consider three replicas, A, B, and C. A is initially the leader. The test uses two clients and distinct tokens, while the fault injector can partition, pause, and restart nodes.
10:00:00 client 1 invoke reserve(K7) -> A
10:00:01 client 1 info timeout waiting for response
10:00:01 nemesis partition A away from B and C
10:00:03 B and C elect a new leader
10:00:04 client 1 invoke reserve(K7) -> new leader
10:00:04 client 1 ok reservation R-91
10:00:05 client 2 invoke status(K7) on strong path
10:00:05 client 2 ok R-91
This trace can be acceptable. The first request was uncertain. The retry used the same token, and the final status sees one reservation. The checker does not need to decide whether the first request reached A; it only needs to see that the observable outcome has one valid effect for K7.
Now change the final observations:
10:00:04 retry reserve(K7) returns R-91
10:00:05 strong status(K7) returns R-77 and R-91
No healthy-looking dashboard repairs this history. There are two business effects for one key. The likely engineering question is whether the old leader accepted an operation without durable authority, whether a request-record table was not committed with the reservation, or whether the retry handler did not resolve an existing token before creating a row. The checker reports the violation; logs, commit positions, and leader epochs help locate the cause.
There is another, stricter case. Suppose the first reserve(K7) returned ok before the partition, and, after recovery, the strong status path says no reservation exists. If the documented ok is a failover-surviving confirmation, this is an acknowledged-write-loss anomaly. If the system only promised a weaker acknowledgement, the test has instead exposed a contract mismatch: the API response is being read as stronger than its implementation.
The important reasoning move is to separate these cases. A timeout creates uncertainty that must be resolved. A success response creates an obligation according to the published acknowledgement contract.
Choose Faults That Threaten the Contract
A fault is useful when it attacks a mechanism the guarantee relies on. The goal is not an impressive collection of destructive actions.
| Fault schedule | Promise under pressure | Evidence the checker needs |
|---|---|---|
| Partition leader from a quorum during writes | Old authority must not acknowledge unsafe work. | ok writes remain visible after recovery; no duplicate token. |
| Pause or kill a leader during an acknowledgement | Recovery must preserve the stated commit boundary. | The post-failover history is compatible with every completed result. |
| Delay replica application while issuing strong reads | A strong path must not serve an older state. | A read after a completed write sees the required version. |
| Change a node clock when leases determine authority | Lease and fencing assumptions survive the fault model. | No obsolete leader produces accepted results. |
Clock manipulation is not mandatory just because a tool can do it. It belongs in scope when leadership or safe reads depend on time-based leases. Similarly, a stale analytics view is not a failure unless that particular view promised freshness or strength it did not deliver. Name the fault model and its boundary in the test plan, so a result cannot be dismissed vaguely as “unrealistic” after it fails.
The fault injector is often called a nemesis, but it should be boringly deliberate. Schedule the partition near the acknowledgement window; vary timing and concurrent load across runs; retain the seed or schedule that reproduces an anomaly. Randomness can explore more interleavings, yet reproducibility is what turns an alarming trace into a regression test.
What the Checker Can and Cannot Conclude
For a small token-based API, a checker can verify local rules directly: no token maps to two reservation IDs, a completed strong read follows the required completed write, and a release explains a later absence. For a richer concurrent object, the question may be whether all completed operations fit one allowed sequential ordering. Linearizability is the familiar strong form: each operation appears to take effect at one point between its invocation and response, while non-overlapping operations retain their real-time order.
That does not make one passing run a proof of correctness. It means the run found no forbidden history within its workload, duration, topology, and faults. Nondeterministic tests may miss a rare timing window. A model can also omit a business rule. Keep the checker narrow enough to understand, run it repeatedly around releases and configuration changes, and grow it when a real incident reveals a missing invariant.
The cost is real: capturing histories, building a trustworthy model, and operating faults takes time. The benefit is equally concrete. Instead of “failover looked fine,” Harbor Point can say: “under this documented partition schedule, our checker found no acknowledged-write loss, duplicate idempotency token, or forbidden strong read.” That claim is bounded, falsifiable, and useful in a design review.
Check Your Understanding
Check: A fault run causes 200 client timeouts. The test calls every timeout a failure, then reports no duplicate reservations. Why is that conclusion weak?
Answer: Some timed-out calls may have committed. Without recording them as uncertain and resolving their token status, a duplicate created by a retry can be hidden. The history lacks the evidence needed to check at-most-once application.
Check: A partition-and-failover test completes, every node is healthy, and no checker ran. What did the team establish?
Answer: It established that this run recovered operationally. It did not establish that successes survived, strong reads were valid, or retries were idempotent. Those require observable invariants and a history checker.
Practice: Design the Smallest Useful Test
Write a one-page test plan for one endpoint in your system, or use reserve(K).
- State one success-response promise in a sentence that a client can observe.
- Name the operation, idempotency or correlation key, and the result the harness records.
- Pick one fault that threatens the mechanism behind that promise.
- Write the checker rule and how it treats
ok,fail, andinfo. - Name the extra evidence you will keep for diagnosis: request IDs, leader epoch, commit position, topology events, or logs.
Self-review: If your checker can pass when an ok write disappears, your promise and rule do not match. If it cannot say what a timeout means or how a retry reuses identity, add a resolution step before adding more faults.
Connections
- Observability for Replicated Data Systems maps internal signals to operational promises; this lesson checks whether those promises survive adversarial schedules.
- Guarantee Matrix Design Review turns each important promise into an architecture, fallback, measurement, and test obligation.
- Clocks, Leases, and Safe Reads explains why a clock fault becomes relevant when authority depends on lease time.
Resources
- [DOC] Jepsen Analyses
- Focus: See how generated concurrent histories, real clusters, and injected failures expose replica divergence, stale reads, and lost data.
- [PAPER] Linearizability: A Correctness Condition for Concurrent Objects
- Focus: Learn the history-based correctness model behind the strong-read and acknowledged-write examples.
- [PAPER] Elle: Inferring Isolation Anomalies from Experimental Observations
- Focus: Study a checker that infers transactional anomalies from observed histories.
Key Takeaways
- Begin with a client-visible promise, then choose operations, faults, and a checker that can falsify that promise.
- A timeout is uncertainty, not a result; stable tokens and follow-up status checks make retries testable.
- A passing failure test is bounded evidence, while a rejected history points directly to a contract, fencing, retry, or durability decision.