Observability for Replicated Data Systems
LESSON
Observability for Replicated Data Systems
At 09:37, Harbor Point's reservation API is green. Median write latency is normal. Traders receive successful responses. Yet the WAL archive has failed for twelve minutes because its object-store credentials expired. The compliance replica is also falling behind in replay.
The service is available at this instant, but two promises are already weakening: a recovery point after the last archived WAL segment may be unavailable, and a compliance query may be older than its stated freshness bound. A generic “database healthy” light cannot distinguish this state from a genuinely healthy system.
The useful model is simple: every client-visible guarantee needs an internal signal that can support or falsify it. Observability is not a bigger dashboard. It is evidence connecting a promise to the storage transition that makes the promise true.
Core Insight
Start with the claim, then choose the signal.
claim: “A confirmed reservation is recoverable within the RPO.”
evidence: continuous off-host archive progress and a proven restore path
claim: “The compliance view is at most five seconds stale.”
evidence: replica replay progress relative to the source history
claim: “A session can read its own completed write.”
evidence: a serving replica has replayed through the session commit token
CPU, TCP connection count, and median API latency can be useful context. They are not direct evidence for these claims. A signal is useful when a changed value changes the operational decision.
Map Promises to Engine State
The reservation cluster moves a write through several stages. The names differ by product, but the boundaries matter more than the brand.
client command
-> local log append and flush
-> client acknowledgement
-> archive of durable history off-host
-> replica receive and flush
-> replica replay into query-visible state
-> backup, restore, validation, and promotion when recovery is needed
Each stage earns a different claim. A local log flush can support a local crash-recovery claim. An archived log segment extends the recovery story beyond one host. Replica replay supports a read-freshness claim. A restore drill supports an RTO claim; a successful backup upload alone does not.
| Promise | Direct evidence | A tempting but insufficient proxy |
|---|---|---|
| Write latency and local durability | Commit/WAL flush latency, synchronous-ack wait time. | API median latency alone. |
| PITR recovery point | Age and continuity of successfully archived log history. | “Last backup succeeded.” |
| Replica-read freshness | Replay position, byte gap, or a comparable commit timestamp. | Replica TCP connection is open. |
| Read-your-writes | Replica visible position versus this session's token. | Average lag is below a threshold. |
| Restore time | Measured restore-to-serving drill including validation. | Size of the backup file. |
| Repair health | Divergence and repair completion by range. | Cluster has all nodes up. |
PostgreSQL's replication and statistics views expose distinct sender and receiver progress, which is exactly why a single “replication lag” number is often insufficient. The dashboard should preserve whether a problem lies in transport, durable receipt, replay, archive, or restore.
A Worked Investigation
At 09:37, the on-call dashboard shows these illustrative values:
reservation API p95 write latency: 24 ms (normal)
last successful WAL archive: 09:25 (12 minutes ago)
archive retry count: rising
compliance replica receive position: 9C/80
compliance replica replay position: 9C/30
source commit position: 9C/90
replay byte gap: growing
The initial, wrong inference is: “the API is green, so this is a low-priority storage alert.” The trace supplies better evidence.
archive age grows
-> off-host history is no longer continuous at the intended RPO
-> a host loss now exposes more recent accepted changes than promised
receive is near source, replay is far behind
-> transport is working; apply is the bottleneck
-> compliance reads may violate their freshness contract
The correct actions differ. Archive failure needs credential, destination, capacity, or archiver investigation; the team should protect local WAL space while it recovers. Replay delay needs inspection of standby I/O, replay conflicts, long queries, apply workers, or read workload. Restarting the WAL sender is not a fix for a replica that already received the data but cannot replay it.
This is why a signal should carry a question. “Replication lag is high” compresses several hypotheses into one vague alarm. “Receive is current but replay is 600 MB behind and growing” points to an apply-side investigation.
Build a Guarantee-Oriented Dashboard
Harbor Point groups its dashboard by the promise at risk, not by process name or machine.
1. Acknowledged-write path
Track the commit latency distribution, local log-flush time, and any synchronous replication wait. Break them down by endpoint or acknowledgement class when those contracts differ. A rising p99 with normal application CPU can reveal storage saturation or a remote-ack wait before users report a generic outage.
2. Recovery posture
Track the most recent successful archive, archive failure and retry counts, retained local log space, the newest verified base backup, and the result and duration of a restore drill. The important view is not “backup job green.” It is “can we restore to the RPO and return within the RTO?”
3. Replica serving posture
Track receive, flush, and replay positions separately; expose both time and byte/position gaps; and show the trend. A ten-second gap shrinking quickly after a burst is a different operational state from a two-second gap growing through the market open.
4. Derived-state and repair posture
Track indexer or change-stream position relative to the base source, repair backlog, mismatch count, and the age of the last successful reconciliation. A global index may be intended to lag, but that lag must stay inside the query contract it serves.
5. Authority and topology posture
Track current leader or epoch, quorum reachability, membership changes, fencing failures, and stale-route rejects. This does not teach consensus protocol internals; it verifies that the write and read paths still use one authority history.
The trade-off is specificity. These signals are more engine- and topology-specific than a generic host dashboard, and they require the team to keep the guarantee map current. The payoff is that on-call can answer “which promise is threatened?” instead of searching hundreds of charts after customers observe the failure.
From Signal to Alert to Action
An alert should express a decision boundary, not merely a number that looks unusual.
| Signal condition | What it means | First action | Escalate when |
|---|---|---|---|
| Archive age exceeds half the RPO and rises | Recovery window is shrinking. | Inspect archive errors, credentials, and destination capacity. | Archive age approaches the RPO or local WAL space is threatened. |
| Receive close to source, replay gap grows | Standby apply cannot keep up. | Inspect I/O, replay conflicts, and read workload. | A freshness-dependent endpoint lacks an eligible replica. |
| Commit p99 rises with synchronous wait time | Remote acknowledgement path is slowing writes. | Check required standby and link/disk health. | The write SLO or availability policy is violated. |
| Index watermark trails base commits | Alternate-key results may be stale. | Route sensitive reads to base authority or wait for catch-up. | The stated index freshness bound is exceeded. |
| Restore drill fails validation | RTO/RPO is unproven even if backups exist. | Stop treating the recovery objective as met; diagnose the failing stage. | Next drill or test cannot establish a recoverable path. |
The alert does not replace diagnosis. It identifies the guarantee at risk and gives the operator a first discriminating check. Keep it close to a runbook that names the affected endpoints, allowed degraded behavior, and a safe fallback. For example, a lagging compliance replica might cause the UI to display a freshness label or route a sensitive query to the authoritative source; it should not silently return a stale answer under a “current” label.
Boundaries of the Evidence
Metrics are evidence about the system's internal state. They are not a proof that every distributed guarantee holds under every fault.
- A continuous WAL archive supports the recovery input, but only a restore drill verifies that backup, archive, target, validation, and promotion compose into a usable recovery path.
- A small replay gap supports a freshness estimate, but it does not prove a specific session saw its own write. That needs a commit-token comparison.
- A healthy quorum metric supports a topology claim, but it does not prove that a business invariant was never violated by an application bug.
- A repair job finishing supports convergence progress, but it does not state which conflict policy decided concurrent values.
The next lesson turns these boundaries into fault experiments and client-visible histories. Observability tells the team where to look and which claim is at risk. Failure testing asks whether the claim survives an adversarial schedule at all.
Check Your Understanding
Check: The API p95 is normal, but the last successful archive is older than the RPO. Is the reservation service meeting its recovery promise?
Answer: No. The request path can be healthy while the off-host recovery history is incomplete. Escalate the archive failure and treat the RPO as violated or at risk according to the defined threshold.
Check: A replica's receive position matches the source, but its replay position is far behind. Which subsystem is the first suspect?
Answer: Apply, not transport. Inspect replay conflicts, standby I/O, long-running reads, and apply capacity before changing the sender or network path.
Practice: Write One Useful Alert
Harbor Point promises that its compliance dashboard is at most five seconds stale during market hours. Define an alert using a signal, a threshold, a first action, and a user-facing degraded behavior.
A strong answer may alert when the compliance replica's replay-age or comparable position gap exceeds five seconds for a sustained window; inspect whether replay is blocked or I/O-limited; and show an explicit freshness warning or route the sensitive query to a stronger read path. It should not alert only because the replica is connected, nor silently preserve a “current” label while the contract is exceeded.
Connections
- Backup, Snapshots, and Recovery Semantics defines the base backup, archive, replay, validation, and promotion steps whose health this lesson makes observable.
- Failure Testing Replication Claims uses metrics as investigation context, then validates the claim itself against concurrent operations and injected faults.
Resources
- [DOC] PostgreSQL: Monitoring Database Activity — Focus: Identify the statistics views that separate WAL, replication, archive, and I/O evidence.
- [DOC] PostgreSQL: Continuous Archiving and PITR — Focus: Relate archive continuity and retained WAL pressure to a recovery objective.
- [BOOK] Google SRE Book: Service Level Objectives — Focus: Start from a user-visible promise, then select a measurable indicator and an action threshold.
Key Takeaways
- A replicated-data dashboard should map every important client promise to the engine state transition that supports it.
- Archive continuity, replay position, commit waits, repair state, and restore drills reveal risks that generic API health cannot.
- A useful alert names a guarantee at risk, a discriminating signal, a first action, and an honest degraded behavior.
- Metrics narrow the investigation; failure testing is still needed to verify that the system preserves its claims under faults.