Consistency Contracts and API Semantics

LESSON

Consistency and Replication

002 30 min intermediate

Consistency Contracts and API Semantics

By the end of this lesson, you will be able to...

  • Turn a user-visible surprise into a precise read or write contract for one endpoint.

  • Distinguish convergence, session guarantees, causal order, sequential consistency, and linearizability by the histories they rule out.

  • Choose the weakest useful contract and name the routing, version, or coordination mechanism required to earn it.

Idea in one sentence: A consistency model is not a database adjective; it is an API promise about which sequence of observations a client may rely on.

Core Insight

An incident team uses a replicated checklist. Priya is on a phone and marks database failover verified as complete. Leo refreshes the web dashboard. Marta opens the audit timeline to understand which action caused which follow-up.

All three clients read copies of the same data. They do not need the same experience.

The tempting design is to put all three behind “the database replicates asynchronously.” It is attractive because it is short. It is not a contract. It does not tell Priya, Leo, or Marta which histories the API rules out.

The Promise We Need to Keep

A consistency contract starts with the surprise a user must never see. In this incident tool, the relevant API questions are:

PATCH /tasks/42/complete
  What does 200 OK mean?

GET /tasks/42
  Must the writer see the completed task on the next read?

GET /incidents/9/dashboard
  How stale may the summary be for one viewer?

GET /incidents/9/audit-log
  Must a dependent event appear after the event it depends on?

Those are product and API questions. Replication only becomes meaningful after the promise is clear.

The Naive Design

Suppose every request goes to its nearest replica. A write is accepted in Madrid and is copied asynchronously to Virginia. Reads also use whichever replica is nearest.

Priya's phone -> Madrid:   PATCH /tasks/42/complete
Madrid: task 42 = complete @18

Priya's next read -> Virginia: GET /tasks/42
Virginia: task 42 = incomplete @17

The system may converge later. That is not enough for Priya. The API has already said the write succeeded, then displayed a state older than the acknowledged write.

The same shape creates different harm for the other endpoints. Leo may be fine with a slightly stale dashboard. Marta is not fine with an audit story that shows an effect before its cause. The lesson is not “make every read linearizable.” It is “state what each endpoint must prevent.”

A Better Model: Histories Clients May Observe

Plain meaning:

A consistency contract says which stories about reads and writes are legal for a caller to observe.

In this scenario:

Priya's successful update should remain visible to Priya. Leo's dashboard may lag within a stated budget. Marta's audit timeline should preserve cause before effect.

Technical names:

Session guarantees constrain one client's successive observations. Causal consistency preserves cause-and-effect order. Linearizability gives each operation one point between request and response, so completed operations fit a real-time single-copy history.

These models are not a universal “weak to strong” ladder. They rule out different bad histories and may require different mechanisms. Choose by the surprise the endpoint cannot tolerate.

Contract Rules out Does not automatically provide
Eventual convergence Permanent divergence after writes stop Read-your-writes, monotonic reads, causal order, or a freshness bound
Read-your-writes A session reading before its own acknowledged write Agreement among other clients or a global order
Monotonic reads One session moving backward to an older observed version Seeing its own new write unless read-your-writes is also promised
Causal consistency Observing an effect without an observed cause One total order for concurrent, unrelated updates
Sequential consistency Histories that cannot fit one order respecting each client's program order Real-time ordering across clients
Linearizability A completed operation appearing after a later-started operation Cheap reads, low latency, or availability during every failure

A Worked Contract Trace

Start with the task-completion endpoint. The product chooses read-your-writes for the writer.

1. Priya -> Madrid: PATCH /tasks/42/complete
2. Madrid durably accepts task 42 = complete at revision 18.
3. Madrid -> Priya: 200 OK, observed_revision=18
4. Priya -> a nearby replica: GET /tasks/42, min_revision=18

Step 4 is where the contract becomes real. A replica at revision 17 cannot return its old value as a successful answer to this request. The system has several possible ways to honor the token:

Option What the read path does Cost
Route to the write authority Reads the known current path May add distance or concentrate load
Wait for local catch-up Delays until the replica has revision 18 Adds variable read latency
Use a quorum or validated read Asks enough replicas for current evidence Adds coordination cost

The client need not know which option the service uses. It needs to know what min_revision=18 means and what happens if the deadline expires. A clear API might return a retryable 409 or 503 rather than silently returning revision 17.

So far: an asynchronous write pipeline can coexist with a session guarantee. The guarantee requires an extra read-path rule, not wishful language about replication.

Three Endpoints, Three Choices

The incident tool can now write a contract table.

Endpoint Promise Mechanism that earns it Explicit boundary
PATCH /tasks/42/complete Success includes a durable revision token Authoritative write path records revision 18 A timeout may leave the client unsure whether the write committed
GET /tasks/42?min_revision=18 Never returns a version older than 18 as success Token-aware routing, waiting, or fallback May wait or return a retryable failure
GET /dashboard May be at most 2 seconds old and never move backward in one session Freshness metadata plus session version tracking Not proof that a task write committed
GET /audit-log Dependent events appear after their causes Causal dependencies or one ordered event stream Unrelated concurrent events need not have one “true” order

For an authority-changing endpoint, the team may need linearizability. For example, assigning the only incident commander or approving an irreversible rollback may require every client to understand completed operations in real-time order. Herlihy and Wing describe this as each operation taking effect at some point between its call and response.

That is powerful. It is not free. Across replicas, the service may need leader routing, quorum evidence, leases, or a read-index-like mechanism. During a partition, the endpoint may have to refuse instead of pretending it can preserve a real-time single answer.

The Trade-off

Before this lesson, a team might say, “We use eventual consistency.” After this lesson, it can state an endpoint contract with its price.

This helps: Priya does not see her own acknowledged action disappear.
This costs: Some reads route farther or wait for a replica to catch up.
This can still fail: A client can lose the response to a committed write and need an idempotent retry path.
The signal to watch: Token-wait latency, stale-read fallback rate, and the age of dashboard data.

The weakest useful promise is usually the best one. A dashboard can spend little coordination when the product accepts bounded staleness. An authority-changing write should spend more when a wrong answer would grant conflicting power. The right choice depends on named harm, not on a comforting database label.

Common Confusions

Confusion: Eventual consistency means “slightly stale.”

Why it is tempting: Both phrases describe data that may arrive later.

Better model: Eventual convergence alone does not bound staleness or protect a session from moving backward. A freshness budget and session guarantee are separate promises.

Confusion: A commit token makes a later read fresh by itself.

Why it is tempting: The writer now knows revision 18.

Better model: The read path must honor the token by routing, waiting, or failing clearly. A token that no server checks is only a receipt.

Confusion: Causal order means every event has one global order.

Why it is tempting: The audit log looks like one timeline.

Better model: Causality orders an event after its known causes. Independent concurrent events may remain unordered unless the product pays for a stronger total-order contract.

Check Your Understanding

Check: Priya receives 200 OK, observed_revision=18. Her next GET /tasks/42?min_revision=18 reaches a replica at revision 17. Which response honors read-your-writes?

Think first, then reveal.

Answer: The replica must wait, route elsewhere, or return a retryable outcome. Returning revision 17 as a normal success breaks the promise because the client has already observed a successful write at revision 18.

Check: The dashboard shows revision 18, then revision 17, then revision 18 again to the same viewer. Which guarantee is missing?

Answer: Monotonic reads. A bounded-staleness budget alone does not prevent one session from moving backward between reads.

Practice: Contract an Incident API

Add a POST /rollbacks/approve endpoint. Two operators may press approve during a partial network failure. An approved rollback triggers an external change that cannot safely run twice.

Write the contract for the write and the immediate follow-up read. Name the user-visible failure response and the signal the operations team should monitor.

A good answer should mention:

Connections

Resources

Key Takeaways

PREVIOUS Partition-Time Guarantees: CAP and PACELC NEXT Replication Topologies and Failure Domains