Consensus Foundations: Safety, Liveness, and Fault Models

LESSON

Consensus and Coordination

001 30 min intermediate

Consensus Foundations: Safety, Liveness, and Fault Models

By the end of this lesson, you will be able to...

  • Distinguish replication from consensus by identifying which decision needs one authoritative answer.

  • Classify a failure outcome as a safety failure, a liveness failure, or neither.

  • Explain how a fault model changes the evidence a consensus design needs.

Idea in one sentence: Consensus protects one authoritative decision by requiring evidence that remains compatible across failures, even when waiting for that evidence temporarily stops progress.

Core Insight

Two Operators, Two Replacement Nodes

A three-replica control plane stores the voter set for a cluster:

current configuration: {A, B, C}

The network partitions:

A | B, C

Replica A is still running. It can talk to operator West, who proposes:

X = replace C with D

Replicas B and C can talk to operator East, who proposes:

Y = replace A with E

Both proposals look locally reasonable. Each side believes a remote node is unavailable. But the control plane cannot confirm both configurations. The next decision would depend on two incompatible answers to a basic question: who is allowed to vote?

Temporary disagreement is acceptable for some data. Two monitoring replicas may report different CPU samples and converge later. Membership is different. Once the system tells a client that a membership change is authoritative, later convergence cannot erase the fact that conflicting authority was granted.

This is the pressure that creates the need for consensus.

The First Model: Copy the Majority Answer

A tempting model is:

1. elect a leader
2. ask for a majority vote
3. copy the winning answer to every replica

This model works in a quiet cluster when there is one leader, one proposal, no delayed messages, and no change in the voter set. It also captures part of the real mechanism: practical crash-fault protocols commonly use majority quorums.

But “the majority voted” is not yet a safety argument.

Which configuration defined the voters? What did each voter persist? Could a delayed vote from an older attempt be reused? Did a later leader learn about an earlier accepted value? Is the client looking at an appended value, a committed value, or an applied value?

The split exposes the missing part. Consensus is not majority arithmetic followed by copying. It is a protocol for preserving compatible evidence across competing attempts.

The Stronger Model: Protect a Decision Across Attempts

In plain English, consensus makes a group behave as if one authoritative choice won, even though messages can be late and participants can fail.

In this scenario, the protected choice is the configuration command at one position in the control plane's history:

configuration slot 8 = X or Y, but not both

The technical term is consensus. A consensus protocol normally states three kinds of rule:

Quorums matter because they can intersect. In a fixed three-voter configuration, any majority contains two voters, and any two such majorities share at least one voter. Paxos uses that overlap together with proposal-number and promise rules. Raft uses majority evidence together with terms, voting restrictions, and log rules. The shared voter is useful only because the protocol constrains what that voter records and may later accept.

This gives us a better sentence:

Consensus = intersecting evidence + rules that preserve prior choices

The exact evidence changes by protocol. The invariant does not: two incompatible values must not both become authoritative for the same decision point.

Safety and Liveness Ask Different Questions

Safety asks whether a forbidden outcome ever happens.

For this lesson's simplified configuration slot, the safety property is:

X and Y are never both committed for slot 8.

Liveness asks whether a desired outcome eventually happens.

A liveness property might be:

If a quorum can communicate for long enough, some valid value
is eventually committed for slot 8.

The condition matters. A protocol cannot create communication across a permanent partition. Nor can it promise progress from a side that lacks the required evidence.

Consider four outcomes:

Outcome Safety Liveness Why
A waits because it cannot reach a quorum Preserved Temporarily lost on A's side No conflicting value is confirmed, but a request stalls
B and C commit one valid value Preserved Achieved on the quorum side One value gains the required evidence
X and Y are both reported committed Violated Irrelevant to repairing the violation The system created incompatible authority
A stores X locally but never reports it committed Preserved Depends on what happens next Local storage is not the same as commitment

This separation prevents a common operational mistake. An unavailable minority is painful, but it is not automatically a consistency failure. Conversely, a system that answers every request can still be unsafe.

Worked Trace: What Each Replica Can Prove

The following is a simplified teaching trace, not a complete Paxos or Raft execution. Assume a fixed voter set {A, B, C} and a protocol whose commit rule requires durable evidence from two voters.

Starting state

All replicas know that configuration slot 7 is committed:

Replica Can communicate with Local slot 8 Commit evidence for slot 8
A A empty none
B B, C empty none
C B, C empty none

Step 1: A receives proposal X

A may record X locally. It still has evidence from only one voter.

A: X is local
A: X is not committed

The missing reply is not proof that B or C crashed. It only means A cannot currently assemble the evidence required by the commit rule.

Step 2: B and C establish a valid newer attempt

The real protocol must first establish which participant may lead the newer attempt and what prior evidence it must preserve. Those mechanics come later in the track. For this trace, assume those rules permit proposal Y.

B and C durably accept Y:

Replica Local slot 8 Evidence known locally May report committed?
A X one voter for X no
B Y B and C for Y yes
C Y B and C for Y yes

The majority is not selecting truth by popularity. It is producing protocol-defined evidence that can overlap with evidence from a later majority.

Step 3: The partition heals

A receives the newer committed history. Its local X never became authoritative, so the protocol may discard or overwrite it according to its log-repair rules. A then applies Y.

before repair:
A sees local X
B and C know committed Y

after repair:
A, B, C apply Y

The evidence corrects the initial model in two ways:

  1. copying is not commitment; A copied X locally without making it authoritative;
  2. a majority is useful only with rules that make its evidence survive leadership changes and delayed messages.

So far, the cluster preserved one authoritative history by denying the minority a protected write. This matters because recovery has one safe direction: toward the committed history, not toward whichever local value looks newest.

The Fault Model Defines “Failure”

A fault model names the behavior the protocol promises to tolerate. Without it, “fault tolerant” is incomplete.

Crash and recovery faults

A process may stop, lose contact, or restart from durable state. It does not deliberately send contradictory claims to different peers.

Many majority-based crash-fault protocols use 2f + 1 voters to tolerate f unavailable voters. With three voters, the service can lose one and still form a majority of two. This statement assumes the protocol's persistence and recovery rules are correct; the number alone proves nothing about an implementation.

Omission and network faults

Messages may be delayed, lost, duplicated, or reordered. A timeout therefore creates suspicion, not knowledge. These faults can damage liveness by causing elections, retries, or stalls. The safety rules must remain correct even when the suspicion is wrong.

Byzantine faults

A Byzantine participant may equivocate: it can send incompatible claims to different peers. Ordinary crash-fault majority reasoning does not cover that behavior.

In the classical unauthenticated Byzantine agreement model, tolerating f traitors requires at least 3f + 1 participants. Other models, especially those with signatures or different synchrony assumptions, change the details. The important lesson here is the boundary: changing what faulty nodes may do changes the protocol and its evidence. Lesson 023 develops that distinction.

Trade-offs and Boundaries

Consensus improves the defensibility of authority. A client can treat a committed configuration, log entry, or coordination update as part of one protected history.

It costs communication, durable writes, quorum latency, and operational discipline. A minority partition may remain readable under a weaker contract, but it cannot safely confirm protected writes. Slow disks or unstable leadership can reduce progress even when no node has permanently failed.

Consensus also does not solve every distributed-systems problem:

The boundary becomes visible in signals such as loss of quorum, repeated elections, commit latency, unapplied committed entries, or disagreement between a local value and the committed index.

Check Your Understanding

Check: During a partition, a three-voter cluster rejects writes on the one-voter side. Is that a safety failure or a liveness failure?

Think first, then reveal.

Answer: It is a liveness or availability failure on that side. Safety is preserved because the isolated voter does not confirm a conflicting authoritative value.

Check: A leader stores a command locally and replies “committed” before any other voter persists it. What is the dangerous confusion?

Think first, then reveal.

Answer: The leader is confusing local replication state with protocol-defined commit evidence. If it fails, a later quorum may have no evidence that carries the command forward.

Practice: Classify a Five-Voter Partition

A fixed cluster has voters {A, B, C, D, E}. A partition creates sides {A, B} and {C, D, E}. Both receive a different proposal for the same slot.

Explain:

  1. which side can potentially commit under a majority rule;
  2. what the two-voter side may still do locally;
  3. whether a timeout on either side proves a crash;
  4. which outcome would be a safety violation.

A good answer should mention:

Resources

Key Takeaways

NEXT FLP Impossibility and Failure Detectors