Raft Log Replication and Commit Semantics

LESSON

Consensus and Coordination

006 30 min intermediate

Raft Log Replication and Commit Semantics

By the end of this lesson, you will be able to...

  • Trace how AppendEntries proves a shared prefix and repairs a follower's divergent suffix.

  • Distinguish an entry that is local, replicated, committed, and applied.

  • Explain why a Raft leader counts replicas directly only for entries from its current term.

Idea in one sentence: Raft lets a leader copy and repair tentative log entries, but exposes them to the state machine only after current-term quorum evidence makes the ordered prefix authoritative.

Core Insight

The leader of a five-node cluster receives grant lease to worker-12. It appends the command to its own log and sends it to four followers. One follower stores it; two replies are delayed; one server is unreachable.

At first, the entry looks real. It is on disk in two places, including the leader. But the leader may crash before reaching a majority. A new leader might not contain that entry and may repair those two copies toward a different suffix. If the old leader had already applied the lease to the state machine, the service would have made a decision visible that the consensus protocol was not yet forced to preserve.

Raft therefore separates four states:

local       -> stored only on this server's log
replicated  -> known stored on one or more followers too
committed   -> part of the authoritative log prefix
applied     -> executed by the state machine

The key correction is that “stored by a majority” is not a complete commit rule after a leadership change. A leader advances commit by counting a majority for an entry in its current term. That restriction protects old entries until a current leader has established an authoritative prefix around them.

The Situation: A Leader Replicates a Log, Not Individual Facts

Continue with the five-node Raft cluster from the previous lesson. L is the elected leader in term 9; F1 through F4 are followers. A log entry contains a state-machine command and the term in which its leader received it:

(index 41, term 9, command = grant lease to worker-12)

The log is an ordered history. If two replicas apply different commands at index 41, they can reach different states even if every later command is identical. Raft's safety target is therefore not merely “copies exist.” It is that a state machine never applies a different command at an already applied index.

The leader keeps two useful progress values for each follower:

State Meaning
nextIndex[F] The next log index the leader will try to send to follower F.
matchIndex[F] The highest index the leader knows follower F has stored.
commitIndex Highest index the leader knows is committed.
lastApplied Highest index already applied to the state machine.

The first two guide network repair. The last two guard what becomes externally meaningful. All four move forward monotonically in the normal protocol; they are not interchangeable counters.

The Initial Model: Acknowledged Copy Means Commit

We might first think: “If the leader and one follower have an entry, the entry is replicated, so the client can treat it as complete.” This fits a two-node thought experiment where those two nodes happen to be the required quorum.

It fails in a five-node cluster. Leader plus one follower is only two copies, not the required majority of three. More importantly, a later leader may have to remove an uncommitted suffix from a follower whose log diverged during an old leadership period. Local existence and even partial replication are evidence about storage, not yet a promise about the authoritative history.

The stronger model has two parts:

  1. first, prove a follower shares the leader's prefix before extending it;
  2. then, advance the committed prefix only under Raft's term-aware quorum rule.

The Mechanism: Prove the Prefix, Then Extend It

Plain meaning: Do not append after an unknown past. Find the last point where leader and follower agree, replace only the follower's uncommitted conflicting suffix, and apply only the prefix the leader has committed.

In this scenario: L sends AppendEntries with the index and term immediately before its new entries. A follower accepts only when that prior entry matches its own log.

Technical name: This is Raft's log-matching rule, carried by AppendEntries RPCs.

AppendEntries is a consistency check

Each request includes:

term          leader's current term
leaderId      current leader identity
prevLogIndex  index immediately before the new entries
prevLogTerm   term at that index in the leader's log
entries[]     entries to append; empty for a heartbeat
leaderCommit  leader's commitIndex

The follower first rejects a stale term. It then asks one narrow question:

Do I have prevLogIndex, and does its term equal prevLogTerm?

If no, the leader lowers nextIndex for that follower and tries an earlier prefix. If yes, the follower can safely compare the new entries: a conflict at the same index but a different term means it deletes that entry and everything after it, then appends the leader's suffix.

Deleting here is not arbitrary data loss. A follower may discard only a divergent uncommitted suffix; a committed entry must be present in every later leader's log. The repair rule makes the visible log relation precise:

same index and same term
  -> same complete prefix through that index

Worked Trace A: Repair a Follower Before Counting It

The following logs are illustrative. L is leader in term 9. It shares entries 1 through 3 with F4, but F4 kept an uncommitted suffix from an older leader.

L:  [1:t4] [2:t4] [3:t5] [4:t8] [5:t9]
F4: [1:t4] [2:t4] [3:t5] [4:t6] [5:t6]

Step 1: the optimistic append fails

L initially believes F4 may already have index 5, so it sends an AppendEntries request whose previous entry is (index 5, term 9). F4 has index 5, but its term is 6, not 9. It rejects the request.

L -> F4: prevLogIndex=5, prevLogTerm=9
F4 -> L: failure

The rejection does not tell L to invent another history. It tells it that its assumed shared prefix was too long.

Step 2: the leader finds the shared prefix

L decreases nextIndex[F4] and retries. It eventually reaches (index 3, term 5), which F4 does contain. Now the leader sends entries 4:t8 and 5:t9 after that verified point.

L -> F4: prevLogIndex=3, prevLogTerm=5,
         entries=[4:t8, 5:t9]

F4 sees a conflict at index 4: it has term 6, whereas the leader's entry has term 8. It removes its entries 4:t6 and 5:t6, then appends the leader's entries.

F4 after repair: [1:t4] [2:t4] [3:t5] [4:t8] [5:t9]

Only now may L advance matchIndex[F4] to 5. The follower is a repaired copy of the leader's log, not simply a machine that received a packet.

The Mechanism Continues: Commit Is Term-Aware Quorum Evidence

Replication says what a server has stored. The leader commits conservatively. It may advance commitIndex to an index N when:

a majority has matchIndex >= N
and log[N].term == the leader's current term

The first condition gives quorum evidence. The second is Raft's current-term restriction. Once a leader commits a current-term entry, all earlier entries in its log also become committed as part of the same authoritative prefix.

Why not count a majority for every old entry directly? Consider an entry from term 7 that happens to sit on a majority after a new leader starts term 8. Without the restriction, a later sequence of failures can let different leaders infer incompatible things about an old entry's status. Raft instead requires the term-8 leader to commit an entry of its own term; the protocol's leader-completeness and election rules then make the earlier prefix safe to treat as committed too.

This is a safety rule, not a performance preference. It may delay confirmation of old work, but it prevents a new leader from calling an ambiguous old suffix committed merely because it sees several copies at one moment.

Worked Trace B: An Old Entry Waits for a Current-Term Anchor

Five servers need three acknowledgements for a majority. New leader L8 begins term 8 with this entry already in its log:

index 42: (term 7, set owner = node-6)

L8 sees that S2 and S3 also store index 42. Counting L8,S2,S3 gives three copies. Yet L8 does not advance commitIndex to 42 merely by counting that old-term entry.

Instead, L8 appends a current-term entry at index 43. A no-op is sufficient for the example; a real client command could serve the same purpose.

Index Entry Replicas known by L8 Commit result
42 (term 7, set owner = node-6) L8,S2,S3 not counted directly by term-8 leader
43 (term 8, no-op) L8,S2,S3 majority plus current term: commit 43

When L8 advances commitIndex to 43, entry 42 becomes committed as part of the prefix. State machines may now apply 42 and then 43 in order.

The no-op is a teaching model of a common protocol action, not a claim that every implementation uses exactly this message at exactly this moment. The invariant is the important part: a leader uses a current-term committed entry to establish the prefix that includes earlier entries.

So far: a follower's matching log can still contain tentative entries, and a majority of copies can still require a term-aware commit rule. Commitment is the point where the protocol, not an individual disk, makes the ordered prefix authoritative.

Apply Only the Committed Prefix

After the leader commits, it sends its leaderCommit in later AppendEntries messages. A follower advances its own commitIndex only up to the leader's announced committed index and no farther than the entries it has. Each server then advances lastApplied one entry at a time until it reaches commitIndex.

stored in log -> commitIndex moves -> lastApplied moves -> state machine changes

This ordering means a state machine never sees a suffix that leadership repair may still replace. It also means “committed” does not mean every follower has already applied the entry; a slow follower can learn and apply the committed prefix later.

Trade-offs, Limits, and Signals

The trade-off is explicit. Raft gains a clear boundary between tentative copying and authoritative state. It pays in quorum latency, stable storage, follower-progress bookkeeping, and retry traffic while lagging followers find a shared prefix.

This works well when a leader can communicate with a majority. It can still stall when acknowledgements cannot reach a quorum, even if many replicas hold a new entry. It does not solve client retry semantics or external side effects by itself: an applied state-machine command still needs an idempotency or operation-identity design when a client misses its reply.

Useful signals include growing nextIndex backtracking, a large gap between lastLogIndex and matchIndex for a follower, a stagnant commitIndex, and a widening commitIndex - lastApplied gap. They reveal different conditions: divergent or lagging replicas, unavailable quorum progress, and slow local application.

Common Confusions

Confusion: “A follower receiving an entry makes it committed.”

Why it is tempting: The entry is now durable in more than one place.

Better model: It is replicated. Commitment requires the leader's term-aware majority rule; then followers learn the committed index separately.

Confusion: “Repairing a follower deletes committed data.”

Why it is tempting: The follower can remove entries from its local log.

Better model: Raft repairs only an uncommitted conflicting suffix. A committed entry is protected by the protocol's election and log-completeness rules.

Confusion: “A majority for any term lets the new leader commit immediately.”

Why it is tempting: Majority counting was sufficient in the simple current-leader case.

Better model: The leader counts replicas directly for entries from its current term. Committing a current-term entry makes its earlier prefix committed too.

Check Your Understanding

Check: F4 rejects AppendEntries(prevLogIndex=12, prevLogTerm=9) because its entry at index 12 has term 8. What should the leader infer?

Think first, then reveal.

Answer: The leader and follower do not share a prefix through index 12. It lowers nextIndex[F4] and retries with an earlier previous entry; it must find the matching prefix before replacing the follower's conflicting suffix.

Check: A term-20 leader finds that a majority stores entry 77 from term 19, but no term-20 entry has been committed. May it commit 77 solely by counting those copies?

Think first, then reveal.

Answer: No. Raft's leader rule advances commit by counting replicas for an entry from the leader's current term. Once it commits a current-term entry after 77, 77 becomes committed as part of the preceding prefix.

Practice: Inspect a Client Acknowledgement Boundary

In a five-server cluster, term-31 leader L appends rotate key K at index 105. L, F1, and F2 store it; F3 and F4 are unreachable. The entry is from term 31, and L knows the two follower acknowledgements. F1 is slow to apply entries.

  1. Can L advance commitIndex to 105?
  2. Can the state machine on F1 remain temporarily behind while the entry is committed?
  3. If F4 later returns with a conflicting uncommitted suffix, what must happen before it can apply index 105?

Model answer:

  1. Yes. L,F1,F2 are a majority, and index 105 is from the leader's current term.
  2. Yes. Committed and applied are distinct. F1 applies only as lastApplied catches up to its committed prefix.
  3. The leader must use AppendEntries prefix checks to find the common point, replace F4's conflicting uncommitted suffix, replicate index 105, and communicate the committed index. Only then can F4 apply it.

Connections

The previous lesson established who may direct the log. This lesson shows what that leader must do before a command becomes externally meaningful: prove prefixes, repair followers, gather term-aware commit evidence, then apply in order.

The next lesson changes who counts as a quorum. Joint consensus builds on this same commit boundary because a configuration change is itself a log entry whose authority must survive different membership rules.

Resources

Key Takeaways

PREVIOUS Raft Design Principles and Strong Leadership NEXT Raft Membership Changes and Joint Consensus