Replication Lag and Read-Your-Writes
LESSON
Replication Lag and Read-Your-Writes
A trader releases reservation R-88421. Harbor Point's primary commits the release and returns success. The trader immediately refreshes the exposure page, but the page still says 9.9M rather than the lower exposure that the release should produce.
The write did not half-work. The page was sent to a read replica that had not replayed the release yet. This is easy to miss because the replica is healthy, connected, and fast. It is simply serving an earlier prefix of history.
The reasonable first model is that a replica is “the same database, only cheaper for reads.” That model works for a dashboard whose contract allows a short delay. It fails for the next read in the same user workflow. Read-your-writes means the session must not observe an older state than its own completed write.
The solution is not to avoid replicas. It is to make freshness part of routing: carry evidence of the write, compare it with the state a replica has made visible, then wait, route elsewhere, or state that staleness is allowed.
Core Insight
A replica can receive and durably store a transaction before it applies that transaction to the tables a query sees. For a read-your-writes request, the relevant question is therefore not “is the replica connected?” or even “is average lag small?” It is:
Has this node made the specific committed write my session depends on visible to reads?
Represent the required write with a commit token. Depending on the database and topology, that token can be a log position, a GTID set, a hybrid timestamp, or another monotonic commit marker. The router may use a replica only when its visible position is at least the session token in a comparison that is valid for the current authority epoch.
This is a session guarantee, not a claim that every client sees the same global order. It protects one session from going backward after its own success response.
The Moving Parts
At 10:04, the primary commits the reservation release at illustrative position 8A/58. Two replicas report different progress:
commit token returned to trader: 8A/58
replica east: received=8A/58, flushed=8A/58, replayed=8A/40
replica west: received=8A/70, flushed=8A/70, replayed=8A/60
The numbers are a teaching trace, not output from a particular engine. Still, they let us make a precise routing decision.
| Node | Has the release durably? | Can a normal query see the release? | Suitable for token 8A/58? |
|---|---|---|---|
| Primary | Yes | Yes | Yes |
| Replica east | Yes | No; replay ends at 8A/40 |
No |
| Replica west | Yes | Yes; replay ends at 8A/60 |
Yes |
Lesson 004 introduced received, flushed, and replayed positions. Here the key distinction is that replay determines ordinary read visibility. A transport-health metric might show that east has already received the bytes, but it cannot prove that the trader's query will see them.
The Mechanism Step by Step
Harbor Point gives the write and its immediate read a shared piece of evidence.
1. Trader sends POST /reservations/R-88421/release with idempotency key K.
2. Primary commits the release and returns 200 plus commit_token=8A/58.
3. Client stores that token for the session or sends it with its next GET.
4. Router inspects each candidate replica's replay position.
5. Router chooses a replica only if its replay position has reached 8A/58.
6. If none has reached it, the router waits briefly, falls back to the primary,
or returns the explicitly stated freshness outcome.
In plain language: the successful write tells the session what it must be able to see next. The replay position tells the router whether a replica has caught up enough. The technical name for that session property is read-your-writes consistency.
The router's choice is a comparison, not a guess from elapsed time:
required session token: 8A/58
east replay position: 8A/40 -> reject for this session
west replay position: 8A/60 -> eligible
For a single, stable primary and its physical replicas, an ordered log position can make this comparison direct. Systems with failover need a token format or accompanying epoch that remains meaningful across leadership changes. A bare node-local sequence from an old primary may not be comparable to a new leader's sequence. This is an implementation boundary, not a reason to drop the guarantee.
A Worked Session Trace
First see the failed route. The product has a fast load balancer that chooses the closest healthy replica.
10:04:00.010 POST release R-88421 -> primary
10:04:00.016 primary commits release at 8A/58
10:04:00.017 client receives 200 OK
10:04:00.020 GET exposure/CA-MUNI -> replica east
10:04:00.021 east has replayed only through 8A/40
10:04:00.022 east returns exposure=9.9M
The incorrect inference is “the release failed.” The evidence says otherwise: the primary committed it, while east is behind in replay. A generic lag graph could be low, but it still cannot answer whether 8A/58 is visible on east.
Now use a token-aware path:
10:04:00.017 client receives 200, commit_token=8A/58
10:04:00.020 GET exposure/CA-MUNI carries token 8A/58
10:04:00.020 router sees east replay=8A/40 and west replay=8A/60
10:04:00.021 router sends request to west
10:04:00.022 west returns the lower exposure
The router did not make west globally freshest. It made a smaller, sufficient claim: west has applied at least the history required by this session. The distinction matters because “freshest node” can be costly or undefined during normal replication, while a supplied token has a concrete comparison.
If neither replica had reached 8A/58, Harbor Point would have a policy choice:
wait up to 40 ms for an eligible replica -> preserve read offload if it catches up
otherwise route to primary -> preserve read-your-writes at primary cost
otherwise return a retry/freshness status -> preserve honesty, but add a client step
The 40 ms is an illustrative budget. The product should choose it from the user-visible latency budget and observed replay distribution, not copy it as a universal setting.
Freshness Classes Are Product Contracts
Not every endpoint needs the same route. Treating them all as primary reads wastes replica capacity; treating them all as replica reads makes session bugs inevitable.
| Endpoint | Required freshness | Route rule | What the caller must be told |
|---|---|---|---|
| Approve a new reservation | Current authoritative state | Primary or a path with an equivalent authority guarantee | A stale replica is not acceptable for this decision. |
| Trader checks their own release | Read-your-writes | Token-aware replica, brief wait, then primary fallback | The response includes the session's committed outcome. |
| Regional exposure dashboard | Bounded staleness | Replica within a stated replay-age limit | The page shows its freshness age. |
| Historical export | Intentional snapshot | Read from a named snapshot | It is not a live position view. |
This classification is a preference under product constraints. A risk admission check may need stronger authority than read-your-writes, because seeing your write does not prevent another session from changing a limit. A dashboard may tolerate stale data but should never be reused for a decision that assumes current state.
The trade-off is direct. Token-aware routing, waiting, and primary fallback improve the session promise, but add routing state, occasional latency, and primary load when replicas cannot catch up. Bounded-stale reads gain cheap capacity and locality, but must expose the allowed delay rather than masquerading as current data.
The Failover Token Boundary
Assume the trader holds token 8A/58, then the old primary fails before the next read. A router must not merely compare that token to whichever number a newly promoted node reports. First it needs evidence that the new authority includes the committed history the token represented.
before failure
old primary: commit token 8A/58
replica west: replayed through 8A/60
promotion
west recovers its durable history and becomes leader in epoch 19
router records that epoch 19 contains the old prefix through 8A/60
next session read
token 8A/58 is recognized as included in epoch 19
router may read from the new leader or an eligible replica of epoch 19
If the service cannot establish that inclusion, it must treat the token as unresolved and use a safe command-status or authority path. A protocol can encode this with a globally meaningful transaction ID, a timeline plus log position, or an epoch-to-prefix mapping. The exact format is product-specific. The requirement is not: “use an LSN everywhere.” It is: “after authority changes, retain a truthful way to decide whether the session's write is visible.”
This check belongs in a failover drill. Record a completed write, force a leadership change, then ask the session to read through every route the router might use. The expected result is not merely that the new leader is alive. It is that the session either observes its recorded write or receives a clear unresolved-outcome response. A fast stale answer reveals a broken token or epoch mapping.
Failure Boundaries and Signals
Read-your-writes does not solve every replication problem.
- It does not make a derived projection correct. The projection must first be updated correctly at the source; the token then proves whether the selected read node has replayed that update.
- It does not make two different keys or shards atomic. A session can observe its own committed write and still need a separate transaction or workflow to protect a cross-key invariant.
- It does not turn a timeout into a failed command. If a client timed out after sending a write, it must resolve the idempotency key or command status before adopting a token for a supposedly completed result.
- It does not automatically survive failover. The system must define how tokens relate to a new leader, its timeline or epoch, and promotable replicas.
Use signals that identify the exact failed inference.
| Signal | Question it answers |
|---|---|
| Replica replay position versus session token | Can this node serve this session's next read? |
| Wait-for-token latency and timeout rate | Is token-aware routing preserving replica use within budget? |
| Fallback-to-primary rate by endpoint | Which workflow is losing its intended read offload? |
| Replay conflicts, I/O pressure, and apply queue | Why is a replica not becoming eligible? |
| Token failures after an epoch change | Does the read contract remain interpretable during failover? |
Do not rely on one sampled “lag seconds” chart as the routing authority. Database metrics can be sampled or delayed, and different products measure transport, flush, and apply at different points. The session token and the replica's comparable visible progress are the evidence for the individual request.
Check Your Understanding
Check: A replica has received and flushed the trader's release but has not replayed it. Can it serve a read-your-writes dashboard request for that release?
Answer: No. Durable receipt supports recovery evidence, not ordinary query visibility. The router needs replay or another mechanism that proves the query sees the token's state.
Check: A primary returns 200 for a release, but the client loses the response and times out. Should the client invent a token and route a later read by it?
Answer: No. The command outcome is ambiguous. The client should use the idempotency key or a status lookup to discover whether a commit occurred and obtain the authoritative token or result.
Practice: Choose the Fallback
For the trader's post-release page, every replica is behind token 8A/58. The primary is healthy, but the product wants to protect it from unnecessary reads. Choose a policy that includes a maximum replica-wait budget, a fallback, and one metric that would tell the team the policy is no longer buying useful read scale.
A strong answer might wait briefly for a nearby replica, then route to the primary to preserve read-your-writes. It should state that the wait budget is a product decision and observe fallback-to-primary rate plus wait latency. If fallback becomes common, the team should investigate replay capacity or reduce the traffic assigned to this freshness class instead of silently serving stale results.
Connections
- Conflict Resolution and Convergence Policies explains why observing one session's write does not decide the semantic winner of concurrent writes.
- Chain Replication and Ordered Failover supplies a contrasting topology where the tail, rather than arbitrary replica routing, establishes the ordinary committed read point.
Resources
- [DOC] PostgreSQL: Monitoring Replication Statistics — Focus: Distinguish sender and standby progress signals before using them for a client-visible routing claim.
- [DOC] MongoDB: Causal Consistency — Focus: Compare read-your-writes with related session guarantees in a database API.
- [DOC] MySQL: Replication with GTIDs — Focus: Examine a portable transaction identity that can support session-aware routing or waiting.
Key Takeaways
- A replica serves a replayed prefix of history; received or flushed data is not necessarily visible to a normal query.
- Read-your-writes needs request-specific evidence: a commit token and a replica whose visible position has reached it.
- Token-aware wait and primary fallback preserve the session contract while bounded-stale endpoints retain inexpensive replica reads.
- Tokens, routing, and metrics must remain meaningful across timeout, replay delay, derived projections, and failover boundaries.