Reconfiguration and Disaster Recovery Beyond the Happy Path

LESSON

Consensus and Coordination

022 30 min intermediate

Reconfiguration and Disaster Recovery Beyond the Happy Path

By the end of this lesson, you will be able to...

  • Distinguish a protocol-backed member replacement from forced recovery after quorum loss.

  • Trace what evidence must remain connected while a voting configuration changes.

  • Review a recovery runbook for its continuity claim, data-loss boundary, and stop conditions.

Idea in one sentence: Changing membership changes who may decide history; when quorum is gone, recovery can restore service from chosen evidence but cannot honestly claim protocol-proven continuity without it.

Core Insight

Atlas has three voters: A, B, and C, one per zone. Zone C fails. The dashboard is yellow, and an operator wants it green again quickly by editing membership until a replacement appears.

The tempting model is that membership is an inventory list. Under that model, replacing C with D is routine administration. In consensus, membership defines which nodes can form the quorums that carry decision evidence. Changing it changes the safety argument.

The first question is therefore not “how do we add a server?” It is:

Does a healthy quorum still exist to commit a transition from the old decision set to the new one?

If A and B can still form the old majority, the cluster can use its protocol to preserve continuity while it replaces C. If only A survives, A cannot make ordinary progress for the old three-voter configuration. Starting a new cluster from A may be necessary for the business, but it is a forced recovery choice with a stated data-loss boundary, not the same safe operation under a different name.

The Production Symptom

An operator sees one member unavailable. Deployment controllers still work, but their recovery margin is gone: another voter failure would remove quorum. At this moment the cluster is degraded, not dead.

The initial response should protect the existing evidence path:

These checks are not paperwork. They distinguish “the protocol can still authorize a transition” from “operators are about to choose a new authority source.” The previous lesson's disk, network, and placement signals matter here: a slow B may look alive but be too unstable for a safe maintenance operation.

The Initial Model: Make the Member List Look Healthy

The fastest-looking repair is often dangerous:

C is gone
remove C from configuration files
start D with copied state
declare A, B, D healthy

This bypasses the question of who authorized the transition and which log/history D represents. Reusing a stale data directory can also revive an old member identity or configuration. If different partitions perform similar “repair” steps independently, two groups can each appear healthy while carrying incompatible histories.

The fact that a new group responds to clients does not prove it preserved the old committed prefix. A green dashboard is liveness evidence. It is not continuity evidence.

The stronger model separates two operational paths:

Path Preconditions Honest claim
Normal replacement A valid quorum of the old configuration is healthy The protocol commits the configuration transition and preserves its history rules
Forced recovery The old quorum cannot make progress Operators choose a surviving snapshot/log source and state which history may be missing

Both paths can be valuable. Confusing their claims is the safety failure.

The Safe Replacement Path

Assume A and B remain healthy after zone C fails. They are a majority of the old A,B,C configuration. Atlas can now change membership through its protocol.

The exact commands and joint-configuration rules are implementation-specific. The generic evidence sequence is not:

old voters A, B, C can still commit
    -> introduce replacement D without giving it unsafe decision power
    -> D receives a snapshot and the committed log tail
    -> protocol commits the transition involving old and new membership rules
    -> D becomes a voter only when protocol and catch-up conditions allow
    -> protocol removes C
    -> new voters A, B, D continue from the committed history

Many systems support a learner or non-voting stage. When available, it separates copying state from deciding state: D can catch up without changing the quorum calculation. Do not assume every product exposes the same role or exact transition. Use the protocol's documented reconfiguration procedure.

The intermediate state is the important part. D is not useful merely because a process exists. It needs a known snapshot boundary, log catch-up, and a committed configuration role before it can contribute safely to availability.

Worked Trace: Replacing a Lost Voter

The following members and indices are illustrative.

Step Configuration and state Operator/protocol action Evidence preserved
1 Voters A,B,C; committed index 700; C is permanently unavailable A and B confirm they still form the old quorum Old configuration can authorize changes
2 A,B continue committed history through 704 Add D using the protocol's replacement path D has no independent authority yet
3 D installs a snapshot through 690, then replays committed tail to 704 Verify catch-up against the service's required boundary D reconstructs the same state
4 Transition record is committed under the protocol's membership rules Promote or include D according to those rules Old and new decision sets remain connected
5 Voters now include A,B,D Commit removal of C New quorum continues the prior history
6 Replacement completes Test a leader change and recovery while D is present The runbook verifies more than a member count

The trace deliberately avoids a vendor command. The safety property is that valid quorum evidence carries the membership transition forward. An implementation may use joint consensus, a view change, or another documented mechanism, but it must prevent an old configuration and a new configuration from independently deciding conflicting histories.

When Quorum Is Gone

Now change the incident: C is gone and B is permanently unavailable. A alone cannot form a majority of the old configuration. It may have the newest visible log, an old snapshot, or only a partial record. None of those facts gives A normal protocol authority to change membership for A,B,C.

The business might still decide to recover service from a backup or selected survivor. That decision needs an explicit recovery statement:

Recovery source: snapshot S at revision R, or identified surviving data directory.
Continuity claim: this source is the chosen recovery point, not proof that all
acknowledged old-cluster writes after R survived.
Risk boundary: clients may need reconciliation for work after R or with unknown outcome.

This is not a criticism of recovery. It is the difference between a truthful recovery plan and a false consensus claim. The team can choose availability and a bounded loss window when the alternative is a prolonged outage. The team should not tell downstream systems that no acknowledged state could be lost unless it has evidence for that stronger claim.

Recovery Runbook: Evidence Before Action

Before a forced recovery, the runbook should make the choice inspectable:

Question Why it matters
Which old configuration and quorum rule applied? Defines what ordinary protocol progress required
What surviving snapshots, logs, and backups exist, and at which revision/time? Identifies candidate recovery points
What evidence supports choosing one source over another? Separates a quorum-backed source from a best-effort survivor
Which clients may have unknown or lost results after the boundary? Defines reconciliation work and external side-effect risk
Who authorizes the data-loss boundary? Makes a business decision visible rather than accidental
How will the old cluster be fenced or prevented from returning? Avoids two competing authority domains

Restore and bootstrap details vary by implementation. Follow its documented disaster-recovery procedure; do not combine copied directories, hand-edited membership, and old identities from unrelated incidents. A forced bootstrap should establish a clearly new authority domain, rotate or fence client access as needed, and require controllers to resync rather than trusting cached ownership.

Mitigation and Prevention

The trade-off is time versus certainty. A protocol-backed replacement can be slower because the operator waits for catch-up and committed transition evidence. It preserves the continuity claim. A forced recovery may restore service faster after catastrophic loss, but it can introduce a chosen recovery point and reconciliation work.

Prevention reduces how often this choice appears:

Readiness Check

Check: Voters A and B remain healthy after C fails in a three-voter cluster. A replacement D is online but has not yet replayed the committed tail. Should the operator immediately remove C and count D as a voter?

Think first, then reveal.

Answer: No. A and B have the old quorum needed to perform a safe transition, but D has not demonstrated the required state catch-up or committed configuration role. Keep the old evidence path intact, let D catch up through the documented protocol, then commit the membership transition.

Practice: Name the Recovery Claim

Only one voter survives from a three-voter cluster. The latest verified backup is at revision 650; a surviving disk has local entries through 668, but no quorum evidence remains. Product leadership authorizes recovery from the disk after reviewing the risk.

Write the opening of the recovery record.

A good answer should mention:

Connections

The previous lesson showed how disk, network, load, and placement can push a healthy cluster toward a recovery boundary. This lesson makes the boundary explicit when membership and continuity are at stake.

The next lesson extends the evidence question to Byzantine faults, where a replica may lie rather than merely crash and certificates make quorum evidence independently verifiable.

Resources

Key Takeaways

PREVIOUS Operating Consensus Clusters: Latency, Disk, Network, and Sizing NEXT Byzantine Consensus and Quorum Certificates