Log Shipping and Ordered Apply

LESSON

Consistency and Replication

004 30 min advanced

Log Shipping and Ordered Apply

By the end of this lesson, you will be able to...

  • Trace a committed change from a primary's recovery log to a replica's query-visible state.

  • Distinguish received, flushed, and replayed log positions by the operational question each answers.

  • Diagnose why a replica can be connected and durable yet still serve a stale read.

Idea in one sentence: Log shipping gives a replica the primary's ordered recovery history, but received bytes, durable bytes, and query-visible state are three different milestones.

Core Insight

At 09:31, a trader receives success for reservation R-88421. The regional dashboard immediately reads a standby and does not show the reservation. The first message in the incident channel says, “Replication is down.”

It is not down. The standby has received the log record and stored it locally. It has not yet replayed that record into the database pages used by the dashboard query.

The incident reveals a common mistake: treating a replica as either “healthy” or “behind.” Log shipping has an ordered pipeline. Each position in that pipeline answers a different question about transport, crash recovery, and read freshness.

The Situation

The primary commits R-88421 at an illustrative log sequence number, 8A/58.

8A/10  update issuer exposure
8A/28  insert reservation R-88421
8A/40  update an index entry
8A/58  commit transaction 88421

The exact storage records vary by database engine. The teaching model is stable: the write-ahead log contains the ordered recovery history that lets a replica rebuild the same state transition. Shipping only an informal list of changed rows would not necessarily preserve all index, visibility, and commit-order details the engine needs.

The Initial Model

The first reasonable model is “if the standby received the reservation, its queries can read the reservation.” That works only after the standby has applied the ordered log prefix containing the commit.

It fails between network receipt and replay:

primary latest position: 8A/90
standby received:        8A/90
standby flushed:         8A/90
standby replayed:        8A/40

The reservation commit at 8A/58 is already on the standby's disk. It is not yet in its query-visible state. A connection check and a dashboard freshness check therefore need different evidence.

The Mechanism Step by Step

Plain meaning:

The replica first receives the primary's ordered log, then persists it, then replays it into the data state that reads can see.

In this scenario:

The standby has enough log bytes to recover R-88421 after a crash, but must replay through 8A/58 before its dashboard query can show that reservation.

Technical names:

In PostgreSQL-style monitoring, these milestones are often expressed as receive, flush, and replay log sequence numbers (LSNs).

1. primary writes and fsyncs the commit record at 8A/58
2. sender streams WAL to the standby
3. standby receives bytes through 8A/90       -> receive position
4. standby fsyncs bytes through 8A/90         -> flush position
5. standby replays records through 8A/40      -> replay position
6. dashboard queries see only the replayed prefix

The primary and standby do not have to be bit-for-bit at the same position for replication to be useful. The important rule is order: the standby cannot skip from 8A/40 to a later query state without applying the intervening history in sequence.

A Worked Diagnostic Trace

Support asks three questions about R-88421. Each one maps to a different watermark.

Question Position to compare Answer at 09:31
Is the network path carrying the log? Receive Yes; receive has reached 8A/90
Could the standby recover this commit after its own crash? Flush Likely yes; flushed through 8A/90
Can the dashboard read this commit now? Replay No; replay is only 8A/40

Why is replay behind? In this example, a long-running dashboard query holds a snapshot that delays apply. Other real pressures include slow standby storage, a burst of write volume, replay conflicts, or expensive index changes.

symptom: dashboard omits R-88421
wrong inference: the transport link is broken

evidence: receive=8A/90, flush=8A/90, replay=8A/40
correction: transport and durable receipt work; ordered apply is behind

The response should target apply: reduce conflicting reads, give replay more I/O, isolate heavy analytics, or route a freshness-sensitive request elsewhere. Restarting the sender would not fix a replica that already has the bytes but cannot apply them.

So far: log shipping is remote recovery plus ordered apply. Replay, not receipt, defines what ordinary standby queries see.

What This Changes

Before this model, a service might send every dashboard query to any connected standby. After it, the service can state a freshness contract:

stale-tolerant dashboard
  may use a replica with a visible replay-age label

read-your-writes workflow
  needs a replica replayed through the session's commit token,
  or must wait, use the primary, or fail clearly

failover planning
  needs durable-log and recovery evidence, not only dashboard freshness

This distinction also explains why log shipping is not a generic event bus. It is strongly tied to the primary engine's recovery format and ordered state transitions. It is a good fit for another copy of the same database. A service that needs a stable, independently evolving business-event contract needs a different interface.

The Trade-off, Limits, and Signals

Log shipping improves recovery and can provide read capacity. The trade-off is that a replica serving heavy reads may delay its own replay, making the read-scale copy less fresh exactly when it is busiest.

Signal What it reveals
Receive lag Transport or sender path is falling behind
Flush lag Remote durable-log window and standby storage pressure
Replay lag User-visible freshness and apply pressure
WAL retention age Whether the primary still has the continuous prefix the standby needs
Replay conflict or long-query count Why apply cannot advance

The continuous-prefix requirement is a hard boundary. If a disconnected standby falls behind the retained log window, later records cannot repair the missing middle. It must retrieve the missing prefix from an archive or be reseeded from a base backup. PostgreSQL documents the same pressure: streaming replication needs retained WAL, an archive, or a replica slot to prevent premature recycling.

From Recovery Copy to New Primary

The three positions also prevent an oversimplified failover rule. Suppose the primary fails while the standby has flushed through 8A/90 but replayed only through 8A/40. A promotion procedure may first replay WAL already durable on the standby before it opens for new writes. The new primary can then make more state visible than the old dashboard had shown.

before failure
  receive=8A/90, flush=8A/90, replay=8A/40

promotion path
  1. stop depending on the old primary
  2. recover from locally durable WAL through 8A/90
  3. establish a new writable timeline
  4. publish the new authority only after recovery completes

This is why “replay lag is six seconds” is not a complete recovery-point statement. It may describe read freshness, while durable WAL gives a different recovery boundary. The exact failover promise still depends on the database, synchronous policy, fencing of the old primary, and whether unreplicated primary-only commits existed. Never infer zero loss merely because a standby was connected.

The operational lesson is to record both views during an incident: the latest position confirmed durable on a promotable standby, and the latest position safely visible to its reads. Those values may be equal on a quiet system and diverge under load. Treating their difference as a defect hides the useful decision they support.

There is also a client boundary. If the old primary may still be reachable, it must be fenced before the promoted standby accepts authoritative writes. Otherwise two machines can each extend a different history after 8A/90. Log recovery preserves the old prefix; a failover protocol still needs leases, a control-plane decision, or another fencing mechanism to ensure that only one new suffix is written. This is separate from shipping speed, but it is essential to the claim that recovery produced one authoritative database. Test this boundary under network partitions, not only clean crashes.

Where This Mechanism Stops

Log shipping replays one engine's physical or recovery-oriented history. It does not automatically give another service a stable business event such as reservation.confirmed. A schema change, engine upgrade, or replication-format mismatch can limit which replicas can replay the stream. PostgreSQL, for example, documents compatibility and setup requirements for primary and standby servers.

That boundary matters when Harbor Point builds downstream workflows. A regional read copy can consume WAL because it needs the same database state. A notification, warehouse, or partner integration should consume an explicit application-level contract, not rely on a storage recovery log whose meaning and format belong to the database engine.

Common Confusions

Confusion: A connected standby is a fresh standby.

Why it is tempting: A healthy network connection is easy to observe.

Better model: Connection and receive progress show transport. Fresh reads require replay progress.

Confusion: Flushed log is already query-visible.

Why it is tempting: The standby has persisted the commit.

Better model: Flush improves crash-recovery evidence. Replay changes the state that queries can read.

Confusion: A later WAL record can fill a missing gap.

Why it is tempting: Logs are append-only, so newer data looks more complete.

Better model: Recovery replays one ordered prefix. If the retained middle is gone, the standby needs an archive or a new base copy.

Check Your Understanding

Check: A standby has receive_lsn=8A/90, flush_lsn=8A/90, and replay_lsn=8A/40. The commit is at 8A/58. Which user-visible claim is false?

Think first, then reveal.

Answer: “A standby dashboard can show the committed reservation now” is false. The commit is received and flushed, but replay has not reached it. The correct concern is apply lag, not a missing transport connection.

Check: A standby was offline long enough that the primary no longer retains the WAL it needs. Why cannot it simply start from the newest WAL record?

Answer: The missing records contain part of the ordered recovery history. It needs the complete prefix from an archive or a new base backup before it can apply later records safely.

Practice: Choose the Correct Watermark

An incident report says: “The remote standby is six seconds behind.” The report contains only the timestamp of the last received log byte. The service has both an analyst dashboard and an emergency failover requirement.

Rewrite the report as two questions, name the watermark each needs, and state one action if replay is the only lagging stage.

A good answer should mention:

Connections

Resources

Key Takeaways

PREVIOUS Replication Topologies and Failure Domains NEXT Synchronous and Asynchronous Replication