Consistency Contracts and API Semantics
LESSON
Consistency Contracts and API Semantics
By the end of this lesson, you will be able to...
Turn a user-visible surprise into a precise read or write contract for one endpoint.
Distinguish convergence, session guarantees, causal order, sequential consistency, and linearizability by the histories they rule out.
Choose the weakest useful contract and name the routing, version, or coordination mechanism required to earn it.
Idea in one sentence: A consistency model is not a database adjective; it is an API promise about which sequence of observations a client may rely on.
Core Insight
An incident team uses a replicated checklist. Priya is on a phone and marks database failover verified as complete. Leo refreshes the web dashboard. Marta opens the audit timeline to understand which action caused which follow-up.
All three clients read copies of the same data. They do not need the same experience.
- Priya must not receive success and then immediately see the task as incomplete.
- Leo may tolerate a dashboard that is two seconds behind, as long as the page says so and does not jump backward for him.
- Marta must not see “rollback started” before the verified failover action it depends on.
The tempting design is to put all three behind “the database replicates asynchronously.” It is attractive because it is short. It is not a contract. It does not tell Priya, Leo, or Marta which histories the API rules out.
The Promise We Need to Keep
A consistency contract starts with the surprise a user must never see. In this incident tool, the relevant API questions are:
PATCH /tasks/42/complete
What does 200 OK mean?
GET /tasks/42
Must the writer see the completed task on the next read?
GET /incidents/9/dashboard
How stale may the summary be for one viewer?
GET /incidents/9/audit-log
Must a dependent event appear after the event it depends on?
Those are product and API questions. Replication only becomes meaningful after the promise is clear.
The Naive Design
Suppose every request goes to its nearest replica. A write is accepted in Madrid and is copied asynchronously to Virginia. Reads also use whichever replica is nearest.
Priya's phone -> Madrid: PATCH /tasks/42/complete
Madrid: task 42 = complete @18
Priya's next read -> Virginia: GET /tasks/42
Virginia: task 42 = incomplete @17
The system may converge later. That is not enough for Priya. The API has already said the write succeeded, then displayed a state older than the acknowledged write.
The same shape creates different harm for the other endpoints. Leo may be fine with a slightly stale dashboard. Marta is not fine with an audit story that shows an effect before its cause. The lesson is not “make every read linearizable.” It is “state what each endpoint must prevent.”
A Better Model: Histories Clients May Observe
Plain meaning:
A consistency contract says which stories about reads and writes are legal for a caller to observe.
In this scenario:
Priya's successful update should remain visible to Priya. Leo's dashboard may lag within a stated budget. Marta's audit timeline should preserve cause before effect.
Technical names:
Session guarantees constrain one client's successive observations. Causal consistency preserves cause-and-effect order. Linearizability gives each operation one point between request and response, so completed operations fit a real-time single-copy history.
These models are not a universal “weak to strong” ladder. They rule out different bad histories and may require different mechanisms. Choose by the surprise the endpoint cannot tolerate.
| Contract | Rules out | Does not automatically provide |
|---|---|---|
| Eventual convergence | Permanent divergence after writes stop | Read-your-writes, monotonic reads, causal order, or a freshness bound |
| Read-your-writes | A session reading before its own acknowledged write | Agreement among other clients or a global order |
| Monotonic reads | One session moving backward to an older observed version | Seeing its own new write unless read-your-writes is also promised |
| Causal consistency | Observing an effect without an observed cause | One total order for concurrent, unrelated updates |
| Sequential consistency | Histories that cannot fit one order respecting each client's program order | Real-time ordering across clients |
| Linearizability | A completed operation appearing after a later-started operation | Cheap reads, low latency, or availability during every failure |
A Worked Contract Trace
Start with the task-completion endpoint. The product chooses read-your-writes for the writer.
1. Priya -> Madrid: PATCH /tasks/42/complete
2. Madrid durably accepts task 42 = complete at revision 18.
3. Madrid -> Priya: 200 OK, observed_revision=18
4. Priya -> a nearby replica: GET /tasks/42, min_revision=18
Step 4 is where the contract becomes real. A replica at revision 17 cannot return its old value as a successful answer to this request. The system has several possible ways to honor the token:
| Option | What the read path does | Cost |
|---|---|---|
| Route to the write authority | Reads the known current path | May add distance or concentrate load |
| Wait for local catch-up | Delays until the replica has revision 18 | Adds variable read latency |
| Use a quorum or validated read | Asks enough replicas for current evidence | Adds coordination cost |
The client need not know which option the service uses. It needs to know what min_revision=18 means and what happens if the deadline expires. A clear API might return a retryable 409 or 503 rather than silently returning revision 17.
So far: an asynchronous write pipeline can coexist with a session guarantee. The guarantee requires an extra read-path rule, not wishful language about replication.
Three Endpoints, Three Choices
The incident tool can now write a contract table.
| Endpoint | Promise | Mechanism that earns it | Explicit boundary |
|---|---|---|---|
PATCH /tasks/42/complete |
Success includes a durable revision token | Authoritative write path records revision 18 | A timeout may leave the client unsure whether the write committed |
GET /tasks/42?min_revision=18 |
Never returns a version older than 18 as success | Token-aware routing, waiting, or fallback | May wait or return a retryable failure |
GET /dashboard |
May be at most 2 seconds old and never move backward in one session | Freshness metadata plus session version tracking | Not proof that a task write committed |
GET /audit-log |
Dependent events appear after their causes | Causal dependencies or one ordered event stream | Unrelated concurrent events need not have one “true” order |
For an authority-changing endpoint, the team may need linearizability. For example, assigning the only incident commander or approving an irreversible rollback may require every client to understand completed operations in real-time order. Herlihy and Wing describe this as each operation taking effect at some point between its call and response.
That is powerful. It is not free. Across replicas, the service may need leader routing, quorum evidence, leases, or a read-index-like mechanism. During a partition, the endpoint may have to refuse instead of pretending it can preserve a real-time single answer.
The Trade-off
Before this lesson, a team might say, “We use eventual consistency.” After this lesson, it can state an endpoint contract with its price.
This helps: Priya does not see her own acknowledged action disappear.
This costs: Some reads route farther or wait for a replica to catch up.
This can still fail: A client can lose the response to a committed write and need an idempotent retry path.
The signal to watch: Token-wait latency, stale-read fallback rate, and the age of dashboard data.
The weakest useful promise is usually the best one. A dashboard can spend little coordination when the product accepts bounded staleness. An authority-changing write should spend more when a wrong answer would grant conflicting power. The right choice depends on named harm, not on a comforting database label.
Common Confusions
Confusion: Eventual consistency means “slightly stale.”
Why it is tempting: Both phrases describe data that may arrive later.
Better model: Eventual convergence alone does not bound staleness or protect a session from moving backward. A freshness budget and session guarantee are separate promises.
Confusion: A commit token makes a later read fresh by itself.
Why it is tempting: The writer now knows revision 18.
Better model: The read path must honor the token by routing, waiting, or failing clearly. A token that no server checks is only a receipt.
Confusion: Causal order means every event has one global order.
Why it is tempting: The audit log looks like one timeline.
Better model: Causality orders an event after its known causes. Independent concurrent events may remain unordered unless the product pays for a stronger total-order contract.
Check Your Understanding
Check: Priya receives 200 OK, observed_revision=18. Her next GET /tasks/42?min_revision=18 reaches a replica at revision 17. Which response honors read-your-writes?
Think first, then reveal.
Answer: The replica must wait, route elsewhere, or return a retryable outcome. Returning revision 17 as a normal success breaks the promise because the client has already observed a successful write at revision 18.
Check: The dashboard shows revision 18, then revision 17, then revision 18 again to the same viewer. Which guarantee is missing?
Answer: Monotonic reads. A bounded-staleness budget alone does not prevent one session from moving backward between reads.
Practice: Contract an Incident API
Add a POST /rollbacks/approve endpoint. Two operators may press approve during a partial network failure. An approved rollback triggers an external change that cannot safely run twice.
Write the contract for the write and the immediate follow-up read. Name the user-visible failure response and the signal the operations team should monitor.
A good answer should mention:
- One authoritative approval decision, often requiring linearizable or equivalently strong conditional-write semantics for this invariant.
- An idempotency key or approval identifier, because a lost response does not prove the write failed.
- A token-aware status read or authoritative fallback after approval.
- A retryable refusal when the service cannot establish current authority, rather than two independent approvals.
- Signals such as conditional-write conflicts, authority/read latency, and unresolved approval status.
Connections
- Partition-Time Guarantees: CAP and PACELC explains why an authority-changing endpoint may refuse during a partition; this lesson turns that choice into an API promise.
- Replication Topologies and Failure Domains asks where the write authority, readers, and recovery copies should live so these contracts survive real failures.
Resources
- [PAPER] Linearizability: A Correctness Condition for Concurrent Objects — Focus: Compare the real-time rule with weaker contracts before demanding a single-copy history.
- [PAPER] Session Guarantees for Weakly Consistent Replicated Data — Focus: Read the four per-session guarantees and the client observations each one protects.
- [PAPER] Time, Clocks, and the Ordering of Events in a Distributed System — Focus: Use happened-before to separate causal order from a total order.
- [BOOK] Designing Data-Intensive Applications — Focus: Connect models of consistency to replication and API design choices.
Key Takeaways
- A consistency contract describes histories clients may observe; it is not a blanket database label.
- The contract should start from a concrete user-visible surprise, such as seeing an acknowledged write disappear or seeing an effect before its cause.
- Tokens, routing, waiting, quorums, and ordered streams are useful only when they earn a named endpoint promise.
- Stronger guarantees reduce some surprises but add latency, coordination, or failure-time unavailability; use the weakest one that protects the invariant.