Linearizable Reads, Leader Leases, and Fencing
LESSON
Linearizable Reads, Leader Leases, and Fencing
By the end of this lesson, you will be able to...
Trace the evidence a leader needs before returning a current coordination read.
Explain when a lease is a conditional read optimization rather than permanent authority.
Show how fencing makes a paused former controller unable to change a protected resource.
Idea in one sentence: A current read needs evidence that the leader and its applied state are still current; a fencing token makes an old actor harmless when that evidence later becomes stale.
Core Insight
Atlas has two scheduler controllers, A and B, and a separate workload database. Only the current controller may assign a job. Controller A was leader a moment ago, so its local memory says it owns token 41. Then A pauses during a long runtime pause and loses contact with the consensus cluster. The remaining replicas elect B and give it a newer token, 42.
The easy model says: “A was leader, so A can read its own state and continue.” The model works only while A has evidence that it is still leader. After a partition or pause, A cannot learn from its local log alone whether the cluster has advanced without it.
Three mechanisms address different parts of this problem:
- a linearizable read returns a value that can be placed at one instant between the request and response;
- a leader lease can reduce the cost of some current reads, but only under explicit timing assumptions;
- a fencing token is checked by the protected workload database, so delayed actions from
Alose to token42.
They complement rather than replace one another. A fresh read does not prevent a controller from being paused after the read. A lease does not stop a process from sending a delayed request. A fencing token does not tell a controller which desired state it missed.
The Situation: A Read That Grants Authority
Not every read has the same consequence. Atlas can show a slightly stale replica count on a dashboard without harming a job. It cannot safely use a stale value to answer “may I assign this job?”
| Read | Staleness consequence | Suitable evidence |
|---|---|---|
| Dashboard: current queue length | An operator may see an older number | Explicitly stale read can be acceptable |
| Controller: current job owner | Two controllers may both act | Current, ordered authority evidence |
Client: was configuration revision 901 applied? |
A follow-up write may use a wrong premise | Revision-aware or linearizable read |
Workload database: may token 41 write? |
A former controller may overwrite newer work | Target-side comparison with highest accepted token |
A linearizable read does not mean “the value is true forever.” It means the response reflects a single point during this read's execution. If a client will act later, it still needs a conditional update, a new read, or an enforcement token appropriate to that later action.
The Initial Model: Serve from the Leader's Local Memory
Suppose A was leader in term 9. It has applied through index 120 and remembers that it owns scheduler token 41.
if local_role == leader:
return owner=A, token=41
This is fast. It also has no freshness proof. During a partition, A may be isolated with one follower while the other three replicas elect B in term 10. B commits a new owner record with token 42. A's memory is internally consistent but no longer current.
The evidence from the previous lesson explains why the larger side can advance: its quorum intersects the history that later leaders must preserve. The isolated former leader cannot infer that it still holds authority merely because it has not received a revocation message. Silence can mean delay, partition, or replacement.
The stronger model asks two separate questions before A returns an authority read:
- Does
Ahave evidence that it is still the current leader for this read? - Has
Aapplied enough committed state to answer at that evidence boundary?
The Better Model: Establish a Read Boundary
A conservative path puts a read barrier through the replicated log. Once the barrier is committed and applied, the leader can return a state-machine result that includes earlier committed writes. This is easy to reason about, but it adds log work and latency.
A common alternative is a quorum-confirmed read, often described as a read-index style path. The leader contacts a quorum in its current term and receives confirmation that it still leads. That establishes a safe read index R. The leader then waits until its state machine has applied through at least R before answering from local state.
quorum confirms current leadership
-> leader obtains read boundary R
-> leader waits for applied_index >= R
-> leader reads state and replies
The wait matters. Quorum confirmation alone can show that the leader is current, while its local state machine might still lag behind the committed boundary. Reading before application could return an older owner even though the current log already contains the new one.
This is a teaching model, not a claim that every consensus API uses the same message sequence. Some systems expose a built-in linearizable read; others append a no-op or read barrier; details depend on protocol and implementation. The invariant is the same: the reply must be tied to evidence of current leadership and a state machine applied far enough to reflect the relevant committed history.
A Worked Trace: From Current Read to a Fenced Write
The following trace uses illustrative terms, indexes, and tokens.
| Step | Consensus state | Controller or database action | What is established |
|---|---|---|---|
| 1 | A has term 9, token 41, applied index 120 |
A begins an authority read |
Its old memory alone is not enough |
| 2 | A cannot obtain a quorum after a partition |
A must not claim a current read |
Lack of quorum confirmation blocks the fast path |
| 3 | Remaining quorum elects B in term 10 |
B commits owner B, token 42, at index 121 |
A newer authority record has durable evidence |
| 4 | B confirms leadership with a quorum and obtains read boundary 121 |
B waits until applied_index >= 121 |
B can read the current owner result |
| 5 | Workload database has accepted token 42 |
B writes assignment job-7 -> node-3 with token 42 |
The protected target records the high-water token |
| 6 | A resumes with a delayed request carrying token 41 |
Database compares 41 < 42 and rejects it |
Old process may run, but old authority cannot succeed |
Step 6 is why fencing belongs at the protected target. A consensus cluster cannot prevent A from waking up or prevent a delayed network packet from arriving. The workload database can reject authority that is older than the authority it already accepted.
Question: Could A use a linearizable read after it resumes, see token 42, and then safely send the old request with token 41?
Answer: No. The new read corrects A's knowledge, but it does not retroactively make the old request safe. A must abandon or rebuild the action using current state. The database's token check remains the final protection if A races or behaves incorrectly.
So far, the read path makes a present-tense claim about coordination state. Fencing makes a future side effect compare its authority against the target's present record.
Leader Leases: A Fast Path with Preconditions
A leader lease is a bounded claim that no other valid leader can exist during a specified interval. When an implementation can justify that claim, it may serve some linearizable reads without contacting a quorum for every one.
The lease is not a universal shortcut. Its correctness depends on the implementation's named assumptions, such as conservative clock bounds, monotonic time measurement, bounded lease duration, renewal rules, and an election protocol that prevents overlap under those bounds. Long process pauses and clock anomalies are part of the safety analysis, not mere operational inconveniences.
The trade-off is explicit:
| Path | What it buys | What it costs or assumes |
|---|---|---|
| Log/read barrier | Straightforward current-read evidence | Replication round and added latency |
| Quorum-confirmed read | Avoids appending every read | A quorum communication round when fresh evidence is needed |
| Leader lease | Very low per-read coordination cost | Strong, carefully enforced timing assumptions and conservative expiry |
| Explicitly stale read | Low latency and high availability | Cannot grant authority or validate a current precondition |
Choose the path by the claim the read supports. A metrics screen may intentionally choose stale data. A lease grant, lock owner, or update precondition normally needs evidence that is current enough for that authority decision.
Where It Breaks
Leases and linearizable reads do not solve every authority problem.
- A lease can be misconfigured or its timing assumptions can be violated. The signal is clock-skew alarms, long pauses, missed renewal deadlines, or overlapping-leader reports.
- A current read can become old before the client acts. The signal is a failed compare-and-swap or a rejected token; the response is to refresh the decision rather than retrying blindly.
- Fencing works only when every protected write path checks the monotonic token. The signal is a downstream action that succeeds with a token lower than a previously accepted one.
- A token may order authority without carrying the desired workload state. The controller still needs revisions, resync, and idempotent operation identity from the earlier lessons.
The important boundary is this: consensus can make the current owner record authoritative, but protection is incomplete until the resource that can be harmed rejects stale owners.
Trace It Yourself
Check: A leader has just received quorum confirmation for read boundary 200, but its local applied_index is 198. May it return the state-machine result for a linearizable read at boundary 200?
Think first, then reveal.
Answer: No. It has evidence that it is current enough to establish the boundary, but its state machine has not executed entries 199 and 200. It must apply through at least 200 before returning a result that claims to include that boundary.
Practice: Review a Lease API Claim
An API returns owner=A from a locally cached leader lease. It gives callers no revision or token. A caller later uses that result to delete a shared resource.
Review the design.
A good answer should mention:
- whether the read is linearizable or intentionally stale, and what evidence makes that claim true;
- if a lease fast path is used, its bounded timing assumptions and expiry behavior;
- a revision or fencing token that represents the authority the caller obtained;
- target-side rejection of an older token after a newer one succeeds;
- a conditional operation or fresh read if the action depends on state that may change between reading and acting.
Connections
The previous lesson explained why quorum intersection preserves evidence across leaders. Here that evidence becomes a safe read boundary and an external enforcement token.
The next lesson turns these requirements into usable coordination APIs: compare-and-swap, leases, watches, and recovery contracts.
Resources
- [PAPER] The Chubby Lock Service for Loosely-Coupled Distributed Systems — Focus: Locks, leases, sessions, and client-visible coordination semantics.
- [DOC] etcd API Guarantees — Focus: Revisions, linearizable reads, leases, and watch recovery semantics.
- [PAPER] In Search of an Understandable Consensus Algorithm — Focus: Terms, leader completeness, and the conditions behind safe leader reads.
- [PAPER] Spanner: Google's Globally-Distributed Database — Focus: Timing assumptions and the cost of stronger time guarantees.
Key Takeaways
- A linearizable authority read needs current-leader evidence and a state machine applied through the read boundary.
- A current read is not a promise that the state remains current when a client acts later.
- Leader leases can reduce read latency only under explicit, conservatively enforced timing assumptions.
- Fencing moves enforcement to the protected resource, which rejects actions carrying an older token.
- Revisions, conditional operations, and idempotent action identities complete the gap between learning authority and using it safely.