Synchronous and Asynchronous Replication

LESSON

Consistency and Replication

005 30 min advanced

Synchronous and Asynchronous Replication

Harbor Point's reservation API has accepted R-88421. Its primary database has made the commit record durable locally. A standby in another zone is connected, but it has not confirmed that record yet. May the API return 201 Created now?

It is tempting to say yes: there is a primary, there is a replica, and replication is running. That model works when the product can tolerate losing a very recent confirmed reservation after a primary failure. It fails when “accepted” means that the reservation must survive losing that one machine.

The useful question is narrower: which event must happen before this client receives success? That event is the acknowledgment boundary. Asynchronous replication puts it at local durability on the primary. Synchronous replication puts some remote milestone on the path too. The choice changes write latency, what can be lost on failover, and whether a slow replica can stop new writes.

This lesson uses one leader and one standby. The same reasoning applies to other replication designs, but it does not by itself explain quorum or consensus commits; the next lesson takes up the quorum case.

Core Insight

An acknowledgment is not merely a response-time choice. It tells the client which failure story is still allowed. Local-only acknowledgment allows a primary-only loss window. Waiting for a remote durable flush trades part of the write path for the claim that the record exists beyond that primary. Waiting for apply makes a stronger read-visibility claim and costs more still.

The Moving Parts

Lesson 004 separated the three positions a standby can report. Reuse them here, because “the replica has it” is too vague to be a contract.

Milestone What happened on the standby What it supports What it does not promise
Receive The standby has received log bytes, perhaps only in memory. Evidence that transport reached another machine. Survival of a standby crash.
Flush The standby has durably stored the log bytes. Recovery of that history after the primary fails. A read on the standby can see the change.
Apply The standby has replayed the record into its queryable state. A standby query can observe the change. A short or cheap commit path.

The primary has a local milestone too: it makes its own commit record durable. That protects against a process restart or clean recovery of that machine. It does not protect against losing the primary host, its storage, or the failure domain that contains both storage and process.

For the reservation, call the commit position 8A/58. The positions below are illustrative, not measurements from a particular database:

primary durable through 8A/58
standby received through 8A/58
standby flushed  through 8A/40
standby replayed  through 8A/22

Only the primary has a durable copy of 8A/58. A connected standby is useful evidence about the pipeline, but it is not yet the evidence Harbor Point needs for a primary-loss durability promise.

The Mechanism Step by Step

Both modes start the same way. The primary validates the reservation, changes its state, and makes an ordered commit record durable locally. The difference is when it replies to the caller.

Asynchronous replication

1. Client sends POST /reservations/R-88421/approve.
2. Primary commits and flushes 8A/58 locally.
3. Primary returns 201 Created to the client.
4. Sender later transmits 8A/58 to the standby.
5. Standby receives, flushes, and eventually applies 8A/58.

The primary does not wait in steps 4–5. This keeps the normal write path short: it needs local storage, not a network round trip plus remote storage. It also lets the service continue accepting writes when the standby is slow or disconnected.

That availability is purchased with a recovery window. Between steps 3 and 5, a successful response may describe a commit that exists durably on only the primary. If that primary fails permanently, promotion of the standby can start from an older durable prefix. The client may have a confirmation for a reservation the new primary cannot recover.

This is not a bug in asynchronous replication. It is the allowed outcome of its acknowledgment rule. An operation can fit that rule when it is derived, replaceable, or explicitly allowed to have a recovery-point objective. A search-index update or analytics projection often fits. A confirmed reservation may not.

Synchronous replication with remote flush

1. Client sends POST /reservations/R-88421/approve.
2. Primary commits and flushes 8A/58 locally.
3. Primary sends 8A/58 to a selected standby.
4. Standby writes and flushes 8A/58 durably.
5. Standby acknowledges that flush to the primary.
6. Primary returns 201 Created to the client.
7. Standby applies 8A/58 later, when replay reaches it.

Now success means that the commit has durable copies on at least the primary and the selected standby. If the primary fails after step 6, a correctly selected and promoted standby has the log record it needs to recover R-88421. The standby might not have shown it to a dashboard at step 6, because apply is later, but the durable recovery evidence exists.

The word synchronous hides an important choice. Some systems wait for remote receipt, some for remote durable logging, and some for remote apply. PostgreSQL exposes remote-write, flush, and apply wait points; their names describe different commitments. MySQL's semisynchronous mode similarly waits for configured replica acknowledgments after events have been written and flushed to a relay log, rather than waiting for application on every replica.

Client waits for Stronger statement the client may infer Remaining gap
Primary local flush The primary can recover the commit if it survives. The primary alone may be lost.
Standby receipt Another machine observed the bytes. Its crash may erase an in-memory record.
Standby flush Another machine has durable recovery history. Its ordinary reads may still be stale.
Standby apply That standby can read the changed state. Replay work and remote delay lengthen commit.

The exact vendor setting, number of required standbys, storage guarantees, and promotion procedure matter. So the table is a teaching model, not a substitute for an engine's documentation or a failure test.

A Worked Failure Trace

Consider two possible histories for R-88421. The network link is healthy until it is not; no timeout can make missing evidence appear.

The asynchronous history

09:30:00.120  client sends the approval
09:30:00.123  primary flushes 8A/58 locally
09:30:00.124  primary returns 201 Created
09:30:00.125  primary host and its storage fail
09:30:00.126  standby is promoted with durable log only through 8A/40

The client did everything reasonably: it received a normal success response. Yet the promoted database lacks R-88421. The missing interval is not “replication lag” in the vague sense. It is the interval that the acknowledgment policy deliberately exposed to the caller.

The right remediation depends on the product contract. Harbor Point might reconstruct the reservation from an independent durable command log, reconcile it manually, or tell the client the outcome is uncertain. Retrying blindly can create a duplicate reservation. A request idempotency key helps make a retry safe, but it does not restore a record that no surviving replica has.

The remote-flush history

09:30:00.120  client sends the approval
09:30:00.123  primary flushes 8A/58 locally
09:30:00.126  standby flushes 8A/58 durably
09:30:00.127  primary returns 201 Created
09:30:00.128  primary host and its storage fail
09:30:00.129  standby recovers and is promoted

Here 201 Created was delayed by the remote path, but a primary-only failure should not erase the confirmed log record. During promotion, the standby may replay durable records before accepting new writes. That is why an earlier dashboard result and a later promoted database can differ: dashboard visibility follows replay, while recovery can use flushed log.

This is a boundary, not a universal safety proof. The standby must be in an independent enough failure domain to make the second copy meaningful. Both replicas can still fail together, a bad promotion can choose a stale node, and an unfenced old primary can create two writable histories after a network partition. Remote flush narrows one loss window; it does not solve authority or disaster recovery by itself.

What This Changes in a Service Contract

“This database uses synchronous replication” is not a usable API guarantee. A service needs to state which operations wait for what, and how it behaves when that wait cannot complete.

Operation A reasonable acknowledgment rule Why User-visible degraded rule
Approve a reservation Primary flush plus one independent standby flush. A confirmed approval should survive primary loss. Reject or queue the approval when the required standby is unavailable.
Update a dashboard projection Primary flush only; replicate asynchronously. It is derived and can be rebuilt. Show freshness age; rebuild after failover if needed.
Write an audit event tied to approval Match the approval's remote-flush rule. The evidence should not disappear while the approval survives. Do not silently weaken the rule.
Read a standby immediately after approval Wait for apply, route using a commit token, or use the primary. Flush does not make the row query-visible. Return a bounded-staleness or retry response.

These are preferences under stated costs, not universal rules. A service may choose availability over remote durability during a regional incident, but then the response contract must change visibly. A silent fallback from synchronous to asynchronous means identical 201 responses represent different failure outcomes.

Timeouts are especially deceptive. If the primary waits for a standby and the client times out, the reservation may still commit and later become durable. The client has an ambiguous outcome, not evidence of failure. The service should expose an idempotency key or a status lookup so the caller can resolve the command without inventing a second one.

The Cost and the Signals

Remote waiting adds at least a network round trip plus work at the replica. Tail latency matters more than the average here: one overloaded standby, full disk, or distant link can become part of every synchronous commit. Strict settings can therefore turn a replica incident into a write outage. That is sometimes the intended protection; it must be chosen and tested as such.

Watch signals that identify the boundary rather than a single generic “replication health” light.

Signal Question it answers
Commit wait time or synchronous-commit latency How much delay is the acknowledgment rule adding to the client path?
Required standby count and connection state Can the system still satisfy the configured promise?
Receive and flush positions Has a candidate standby received and durably stored the relevant prefix?
Replay position Can a standby serve a freshness-sensitive read yet?
Timeout and fallback events Did the system block, reject, or weaken the promised acknowledgment rule?

A low replay-lag chart is not proof that a confirmed write would survive failover. Conversely, replay lag does not prove that remote durability is absent. Inspect the relevant position and the actual acknowledgment policy.

Check Your Understanding

Check: A primary returns success after its own flush. The standby receives the record but has not flushed it. What does the success response prove about a primary failure followed by a standby crash?

Answer: It proves only local primary durability. The standby's in-memory receipt may disappear in its own crash, so the response does not prove that a surviving durable copy exists elsewhere.

Check: Harbor Point waits for a standby to flush 8A/58, then immediately reads a dashboard on that standby. The dashboard omits the reservation. Is the synchronous rule necessarily broken?

Answer: No. A remote-flush rule protects durable recovery history. The dashboard needs replay or another read-routing rule. Check the standby's replay position before calling this data loss.

Practice: Write the Degraded-Mode Rule

Harbor Point uses remote flush for reservation approvals. At 10:02, its only required standby loses its network link for six minutes. Write a two-part policy for the API:

  1. What response does a new approval receive while the remote-flush requirement cannot be met?
  2. How does a client safely resolve a request that timed out while the server was waiting?

A strong answer names the trade-off. For example: reject new approvals with a retryable unavailable response rather than silently acknowledging them asynchronously; require an idempotency key; and provide a status endpoint that returns the recorded outcome for that key. A different policy may queue commands or explicitly enter a lower-durability mode, but it must tell both operators and clients that the old survival promise no longer applies.

Connections

Resources

Key Takeaways

PREVIOUS Log Shipping and Ordered Apply NEXT Quorum Reads, Writes, and Tunable Consistency