Raft Design Principles and Strong Leadership

LESSON

Consensus and Coordination

005 30 min intermediate

Raft Design Principles and Strong Leadership

By the end of this lesson, you will be able to...

  • Trace a Raft leader election through follower, candidate, and leader states.

  • Use terms, vote rules, and log freshness to explain why an obsolete leader loses authority.

  • Explain what strong leadership simplifies and why it makes leader stability a progress concern.

Idea in one sentence: Raft makes one leader the only normal source of new log entries, while terms and majority elections make stale leaders step aside instead of turning uncertainty into competing writes.

Core Insight

A five-server coordination cluster has been healthy for hours. S1 is leader, clients send it configuration changes, and the other servers receive the same log entries from it. Then S1 pauses during a long disk stall. The other servers cannot know whether it crashed, is merely slow, or is separated by the network.

The tempting model is: “a timeout proves the leader failed, so any server that times out can take over.” It is useful as a trigger for trying again. It is not safe enough as an authority rule. Two servers may time out at nearly the same time, and the old leader may wake up holding an apparently valid local view.

Raft makes that uncertainty inspectable. Servers have explicit roles; a term numbers a leadership epoch; a candidate needs a majority of votes for its term; and a server that learns of a higher term stops acting on the old one. The protocol does not require clocks to prove failure. It uses timeouts to start an election, then uses durable term state and majority evidence to determine authority.

This is Raft's strong-leadership shape: once elected, the leader is the normal origin of new log entries, and entries flow from leader to followers. That limits the number of histories the system must reason about at once.

The Situation: Three Roles and a Logical Clock

Raft runs under non-Byzantine crash and communication failures. At any moment, a server is in one of three roles:

Role What it does
Follower Responds to leaders and candidates; normally does not start requests of its own.
Candidate Starts an election after an election timeout, increments its term, votes for itself, and asks peers for votes.
Leader Accepts client commands, replicates log entries to followers, and sends heartbeats.

Terms are consecutive numbers used as a logical clock. They are not timestamps and do not say when a machine is physically newest. They let a server recognize information from an older leadership attempt.

term 12: an election may have selected one leader
term 13: a later election attempt; term 12 requests are stale to a term 13 server

A term can end with no leader if candidates split the vote. The safety claim is narrower and more useful: at most one leader can be elected in a given term. A leader or candidate that sees a higher term updates its durable currentTerm and becomes a follower. A server rejects a request carrying a stale term.

This gives each message a question to answer before its content matters:

What term does this message claim, and is that term still current here?

The Initial Model: A Timeout Elects a Replacement

We might initially say: “When a follower stops hearing heartbeats, it becomes the new leader.” This works in a toy system with one clear failure and no competing timeouts.

It fails in the real pressure case. Silence has the same local appearance whether S1 crashed or a heartbeat was delayed. If S2 and S3 both time out, declaring either one leader from its own timeout creates two self-authorized writers. The old S1 might also resume and believe it is still leader in its old term.

Raft replaces that local conclusion with a sequence:

timeout
  -> candidate begins a newer term
  -> voters apply term and log rules
  -> majority vote establishes a leader for that term
  -> higher-term evidence demotes older authority

The timeout starts a liveness attempt. The majority vote and term discipline carry the safety argument.

The Mechanism: Election Before the Write Path

Plain meaning: Before one server directs new log entries, the cluster needs enough peers to recognize that server as the current leader and reject older leadership attempts.

In this scenario: A follower starts term 13, requests votes, and becomes leader only when it has three votes in the five-server cluster. When the old term-12 leader learns of term 13, it steps down.

Technical name: Raft uses leader election and strong leadership. Leader election establishes a term's authority; strong leadership makes log entries flow only from that leader to other servers.

What a candidate changes before requesting votes

When a follower's election timeout expires without valid leader communication, it:

  1. increments currentTerm;
  2. becomes a candidate;
  3. votes for itself and records that vote durably;
  4. resets its election timer; and
  5. sends RequestVote to the other servers.

A RequestVote includes the candidate's term and information about its last log entry: lastLogTerm and lastLogIndex. A voter grants at most one vote in a term, only if the request is not stale and the candidate's log is at least as up to date as its own.

The log check prevents a candidate with an obviously older log from winning an election simply because it timed out first. The next lesson will examine how entries replicate and become committed; here, the important connection is that leader authority must not strand already committed history on an excluded server.

What the winner does

After a majority grants votes, the candidate becomes leader for that term. It immediately sends empty AppendEntries RPCs—heartbeats—to establish its leadership and prevent followers from timing out. Later, the same RPC carries actual log entries.

leader -> followers: AppendEntries(term, previous-log position, entries or none)

Followers do not originate new log entries during normal operation. A client that contacts a follower is redirected or retried toward the leader. This one-way flow is the “strong” part of Raft leadership, not a claim that followers are unimportant: they vote, compare terms, verify log prefixes, store entries, and withhold the acknowledgements a leader needs for progress.

Worked Trace: S1 Pauses, S3 Wins Term 13

The states and timings below are illustrative. Five servers mean a majority of three. In term 12, S1 has been leader and sends regular heartbeats. All servers have the same last log entry: index 40, term 12.

Server Initial role currentTerm Last log entry
S1 leader 12 (index 40, term 12)
S2 follower 12 (index 40, term 12)
S3 follower 12 (index 40, term 12)
S4 follower 12 (index 40, term 12)
S5 follower 12 (index 40, term 12)

Step 1: a timeout creates a candidate, not a leader

S1 pauses and its heartbeats stop reaching S3. After S3's randomized election timeout expires, it changes its own durable state:

S3: currentTerm 12 -> 13
S3: follower -> candidate
S3: votedFor = S3

It sends:

S3 -> S1,S2,S4,S5: RequestVote(
  term = 13,
  lastLogIndex = 40,
  lastLogTerm = 12
)

At this moment S3 has only its own vote. It has a newer term, but it is not leader yet.

Step 2: voters make the authority visible

S2 receives the request. It sees that term 13 is newer than 12, updates its currentTerm, and checks two conditions: it has not already voted in term 13, and S3's log is at least as up to date as S2's. Both are true, so S2 records and grants its one vote.

S4 does the same. The tally is now:

votes for S3 in term 13: S3, S2, S4
majority required: 3

S3 becomes leader of term 13. If S5 had timed out at the same time, it might also have begun a candidacy, but it could not also gather the same majority's one vote in term 13. A split vote can delay progress and start a later term; it does not elect two leaders in one term.

Step 3: the new leader establishes the normal path

S3 sends empty AppendEntries(term=13) heartbeats. S2 and S4 accept the term-13 authority and remain followers. A new client command now follows one path:

client -> S3 (leader) -> S2,S4,... (followers)

The detail that matters is not that S3 has a special personality. It has majority evidence in the current term, and followers use that term in every later interaction.

Step 4: the old leader wakes up

When S1 resumes, it may still locally believe it is leader in term 12. It sends a heartbeat carrying term 12 or receives a reply carrying term 13. Either way, the higher term wins:

S1 observes term 13 > 12
  -> S1 updates currentTerm to 13
  -> S1 becomes follower

S1 cannot keep appending entries as a term-12 leader. It must receive the new leader's log and rejoin as a follower. The term did not prove that S1 crashed; it made the protocol reject stale authority even if S1 was merely slow.

So far: Raft turns an ambiguous timeout into a controlled election. The leader is not the first server to suspect failure; it is the candidate whose current-term request earned a majority under the vote and log rules.

What Strong Leadership Changes

Before this design, we might picture several possible writers trying to advance a log while readers infer which attempt matters. After it, there is one normal write entry point: the elected leader chooses where to append a client command, then replicates it outward.

That constraint reduces nondeterminism. Operators can inspect currentTerm, current role, leader heartbeats, vote requests, and leader changes instead of reconstructing a peer-to-peer proposal race. It also separates three jobs that can be reasoned about independently:

leader election -> who may direct the current log path
log replication -> how followers match and receive entries
safety -> why committed and applied history cannot conflict

The separation is a teaching and implementation choice, not a promise that the components are independent in every proof. For example, election voting consults log freshness precisely because later commit safety depends on leaders carrying the right history.

Trade-offs, Limits, and Signals

The trade-off is explicit. Strong leadership buys a simple normal write path, fewer competing writers, and clearer operational state. It costs leader-centered progress: the cluster needs a leader that can communicate with a majority, and all normal client writes pass through that leader.

This works well when one leader is stable long enough to replicate entries and a majority can communicate. It can still stall when the leader pauses, when a network partition prevents any candidate from reaching a majority, or when multiple candidates repeatedly split votes. Randomized election timeouts reduce the chance of repeated collisions; they do not prove which server is healthy.

Useful signals include an increasing currentTerm, frequent transitions to candidate, vote rejections due to stale logs, missed heartbeats, and repeated leader changes. They indicate liveness or performance trouble. They are not evidence that Raft's safety guarantees have been abandoned: a system may make no progress rather than permit two elected leaders in one term.

Strong leadership also does not tell us when a replicated entry is committed or when it is safe to apply it to the state machine. A leader may have an entry locally without enough follower evidence. That distinction is the next lesson's central problem.

Common Confusions

Confusion: “A timeout proves the leader failed.”

Why it is tempting: A missing heartbeat is the only observation a follower has.

Better model: A timeout is a local suspicion that starts an election. Terms and majority votes, not the timeout alone, establish new authority.

Confusion: “Each term always has a leader.”

Why it is tempting: Diagrams often show a clean election followed by a leader.

Better model: A split vote can leave a term with no leader. Raft promises at most one elected leader per term, not guaranteed immediate success.

Confusion: “Followers are passive and irrelevant.”

Why it is tempting: Entries flow one way in the normal path.

Better model: Followers persist terms and votes, judge vote requests, validate append requests, retain logs, and form the majorities that constrain every leader.

Check Your Understanding

Check: S2 has already voted for S3 in term 18. Later in term 18, S4 asks S2 for a vote with an equally up-to-date log. Should S2 grant it?

Think first, then reveal.

Answer: No. A server grants at most one vote per term. The rule prevents S2 from helping two candidates each assemble a majority in the same term.

Check: A paused leader from term 18 resumes and receives a message carrying term 19. What must it do before considering any client write?

Think first, then reveal.

Answer: It updates its term to 19 and becomes a follower. Its old leader role was valid only under term 18; a higher term makes that authority stale.

Practice: Diagnose a Slow Election

Five servers have just lost their leader. S2 times out first and starts term 27, but its log ends at (index 70, term 25). S3 and S4 each have (index 72, term 26). S2 asks both for votes. Before any candidate receives a majority, S3 later times out and starts term 28 with its newer log.

Which candidate can plausibly win, what must the old candidate do after it learns term 28, and what does this episode show about timeout-based progress?

Model answer: S2 should not receive votes from S3 or S4 because its last log term 25 is behind their term 26; it cannot gain a majority from itself and the one remaining server alone. S3 can plausibly win term 28 if it reaches a majority of voters that have not voted in that term and accept its up-to-date log. When S2 learns of term 28, it updates its term and reverts to follower. The timeouts created election attempts, but log freshness, one vote per term, and a majority determine safe authority; liveness may take more than one term.

Connections

Multi-Paxos used stable leadership to amortize Phase 1. Raft turns that operational shape into the protocol's main interface: roles, terms, one normal writer, and an election that demotes stale authority.

The next lesson follows the new leader's entries. It distinguishes a value that the leader merely appended from one that enough followers replicated and the state machine may safely apply.

Resources

Key Takeaways

PREVIOUS Multi-Paxos and Leader-Based Optimization NEXT Raft Log Replication and Commit Semantics