Final Capstone: Consensus-Backed Control Plane Architecture

LESSON

Consensus and Coordination

024 45 min intermediate FINAL CAPSTONE

Final Capstone: Consensus-Backed Control Plane Architecture

By the end of this lesson, you will be able to...

  • Defend a control-plane authority boundary, including the state that earns consensus cost and the state that does not.

  • Trace a deployment change from conditional write through commit, deterministic apply, watch recovery, lease authority, and fenced external action.

  • Review the architecture against safety invariants, operational signals, recovery boundaries, and an explicit crash-fault versus Byzantine decision.

Idea in one sentence: A consensus-backed control plane is credible only when every authority claim has a committed history, a recoverable API path, a protected side-effect boundary, and a failure test that could disprove it.

Core Insight

Aurora's scheduler may receive committed deployment revision 901 while a former scheduler is still paused with lease token 8. The important design move is to follow authority all the way to its consequence. Agreement inside a metadata cluster is not enough if a controller applies the wrong command, a watch misses an update, a former lease holder can still write outside the cluster, or recovery silently changes which history is trusted. The architecture is coherent only when its authority records, state-machine rules, API results, protected targets, operating limits, and failure tests tell one compatible story.

The Scenario

Aurora is a regional compute platform. Operators submit desired deployments. One scheduler per shard turns those deployments into assignments in an external workload database. Agents watch configuration changes. The control plane must survive one zone loss, recover controllers after pauses, and make it possible to explain what happened after a failed rollout.

The team proposes a simple slogan: “Put the control plane in consensus.” That is a useful beginning, not an architecture. It leaves unresolved which facts enter the log, what a successful API response means, whether a returning controller can still act, how watches recover, and what recovery can honestly claim after quorum loss.

This capstone turns the slogan into a defensible design. It does not prescribe a vendor or implement a full protocol. It asks for evidence at each boundary.

Constraints

Aurora must meet these constraints:

The non-goals matter too. Aurora is not a globally ordered log for every agent event, a general analytics warehouse, or a proof that every external workload operation happened exactly once. Those needs use other systems and their own contracts.

Design Goal

The architecture must make one narrow promise:

For each protected decision, the platform can identify
the committed authority record, the current actor allowed to use it,
the recovery path if that actor is stale, and the evidence that rejects
an older action at the protected resource.

The reasonable initial design stores every deployment object, heartbeat, pod status, log, lease, and scheduler event in the consensus service. Controllers read whichever local cache they have and call the workload database after seeing an owner key.

It fails in two ways. Bulk data makes every authority decision pay unnecessary storage, replication, compaction, and recovery cost. More importantly, a local cache and an owner key do not stop a paused scheduler from acting after a successor takes over.

The stronger design separates authority, observation, and execution. Consensus produces a small ordered history for decisions whose disagreement would create conflicting authority. Controllers derive work from that history. External resources enforce the authority generation they accept.

Proposed Model

Aurora runs three crash-fault voters, one in each of three zones. A majority can continue after one voter/zone loss if the remaining two are healthy and the network path between them is usable. The exact consensus implementation may be Raft-like, Paxos-like, or another correct crash-fault design; the API contract below does not assume a product.

operator API
   | conditional desired-state command
   v
consensus-backed metadata and replicated state machine
   | committed revisions, lease/token generations, operation records
   v
controllers ---- fenced, idempotent requests ----> workload database
   ^                                              |
   | snapshot + watch recovery                    | accepted-token high-water mark
   |                                              v
status, logs, metrics, artifacts, traces <--- observation systems

Authority Boundary

State Inside consensus? Why Outside-boundary rule
Desired deployment generation, immutable spec reference/hash Yes Controllers need one target to reconcile Store the bulky manifest payload in artifact storage
Shard ownership, lease state, fencing generation Yes Conflicting owners can make unsafe assignments Workload database checks the generation on every protected write
Conditional-update revision and membership metadata Yes They define current preconditions and who can decide history Retain and compact according to documented recovery rules
Durable operation record for dangerous retries Yes or transactionally durable with the effect Retries need one stable outcome/in-progress state Bound retention to the supported retry horizon
Pod health, capacity sample, and current agent status Usually no They are frequent observations that can be rebuilt Use them as reconciliation input, not authority alone
Logs, traces, metrics, image layers, large manifests No They need throughput, storage, and query features rather than a single decision order Keep a reference/hash in authority state when integrity matters

The question behind every row is: “Would two different answers grant incompatible authority or change a protected decision?” Importance alone is not enough to earn consensus cost.

Service and API Contracts

The state machine applies only the committed command prefix. A controller never treats a proposal or a local leader receipt as an applied service result.

Contract Aurora's rule
Desired-state update put_if_revision(key, R, command) succeeds only if the inspected revision remains current; success returns committed revision R+1
State-machine apply Same committed command plus same replicated state produces the same internal result on every replica
Current authority read A leader uses its documented current-read path and serves only state applied through that read boundary
Lease loss Renewal failure, closed session, or uncertainty means the controller stops protected work until a new acquisition succeeds
Watch recovery Read snapshot at revision R, watch after R, and resync from a new snapshot after disconnect/compaction gap
External action Include fencing token and stable operation ID; target rejects older tokens and returns stored result or in-progress state for retries
Reconfiguration Change voting membership through the protocol while a valid old quorum exists; do not use a normal-replacement claim after quorum loss

These are design contracts, not implementation details to postpone. Without them, an API user cannot tell what evidence a revision, lease, response, or error carries.

Walkthrough: One Deployment Change

The following identifiers and revisions are illustrative.

Step Committed or observed state Action Evidence and boundary
1 payments is desired generation 41, revision 900 Operator reads revision 900 and prepares immutable spec hash h42 The operator has a precondition, not permanent ownership
2 Command set payments=42, hash=h42, expected=900 enters the log Consensus commits it at revision 901; replicas deterministically apply it Desired generation 42 is now shared service state
3 Scheduler B holds lease generation/token 8 B snapshots through 901, sees observed 41, and creates operation_id=deploy:payments:42 B has a recoverable state base and current authority evidence
4 Operation record says IN_PROGRESS B calls workload database with generation 42, hash h42, token 8, and operation ID Target can distinguish stale authority and a retry from a new operation
5 Database accepts the action but reply is lost B treats the client outcome as unknown, not failed Consensus did not make the network reply atomic
6 B retries the same operation ID and token Database returns stored accepted result One protected effect, repeatable controller result
7 B's watch later disconnects after revision 905 and cannot resume due to compaction B snapshots at 912, replaces cache, then watches after 912 Recovery has a known boundary instead of an assumed gap

This trace connects several lessons. Quorum evidence preserves the command history across future leaders. Deterministic apply makes revision 901 mean the same state to each replica. CAS prevents an operator from silently writing over a newer desired state. Snapshot plus watch makes controller recovery bounded. Fencing makes the workload database refuse a stale scheduler. Idempotency handles a lost reply without pretending that a timeout identifies the result.

Failure Review

An architecture becomes credible when each failure has a safety statement, an expected behavior, and a signal that reveals a broken assumption.

Failure pressure Expected behavior Violation or signal
Leader crashes after a client request but before commit evidence Client does not receive a durable-success claim; later history may omit the proposal An acknowledged desired revision disappears after recovery
Old scheduler pauses; successor gets token 9 after token 8 Workload database accepts 9 and rejects later 8 A protected write succeeds with token below its high-water mark
Controller watch disconnects and history is compacted Controller fetches new snapshot and reconciles Controller continues from an unproven old cache
Slow leader disk or overloaded write path Tail commit latency, retries, and watch pressure rise; controllers fail safe on lease uncertainty Quorum exists but lease/control-loop budget is exceeded
One zone fails Remaining two voters may continue if their path is healthy Minority side claims it can commit new authority
One voter is lost while a quorum remains Replacement catches up and membership transition is committed by the protocol An operator file edit creates an independent voter set
Quorum is permanently lost Ordinary progress stops; recovery source and data-loss boundary are authorized and recorded Forced bootstrap is described as fully continuous without evidence

The values in a production alert policy depend on the lease duration, client workload, and recovery objective. What must not vary is the causal interpretation: a leader indicator alone does not prove the cluster can still supply authority within the needed time.

Trade-offs and Fault Model

The trade-off is deliberate. Consensus gives a compact, defensible history for authority. It costs quorum latency, durable storage behavior, careful client recovery, and operational discipline around member placement and replacement. The small metadata boundary reduces that cost; it does not remove it.

Three voters across three zones are a reasonable crash-fault choice when Aurora trusts the operators and machines as non-malicious participants and wants to survive one zone loss. The design must still handle software bugs, slow disks, partitions, stale controllers, and recovery mistakes; crash-fault consensus does not make those disappear.

Aurora should consider Byzantine protocols only if its actual trust boundary includes replicas or operators that may intentionally equivocate, forge operational claims through compromised identities, or act adversarially across organizations. That choice requires stable identities, key lifecycle control, authenticated votes, certificate verification, different quorum thresholds, and a more complex operational model. BFT is not a remedy for an oversized watch workload or an unfenced workload database.

The situated decision for the stated regional platform is crash-fault consensus plus strong client and target-side enforcement. It becomes risky when the administration boundary is no longer trusted; the signal is a threat model that includes arbitrary participant behavior rather than ordinary crash and delay.

Evidence and Readiness

Aurora must test the claims, not merely check that a leader election occurred. Its verification plan records client invocation/result histories, revisions, commit/apply boundaries, lease generations, operation IDs, and workload-database acceptance decisions.

Invariant Workload and fault Evidence of success
Desired state is not lost after acknowledgement Conditional writes during leader crash and delayed replies Every acknowledged revision appears in a later authoritative snapshot
Only current scheduler can change assignments Pause old scheduler, issue new lease, resume old scheduler Target rejects any action with an older token after newer token succeeds
Controllers converge after watch gaps Disconnect, compact history, make rapid desired-state changes Cache rebuilt from snapshot equals authoritative state
Retry does not duplicate protected work Crash/lost reply after database acceptance Same operation identity returns one stored result/effect
Reconfiguration preserves continuity with quorum Replace a voter after snapshot/log catch-up New configuration retains committed history through a leader change
Forced recovery is honest Simulate permanent quorum loss and bootstrap chosen source Runbook records source, recovery revision, clients needing reconciliation, and old-authority fencing

Jepsen-style fault injection is useful because it treats the history and checker as first-class evidence. A dashboard becoming green after a restart is liveness evidence. It does not prove that an old token was rejected or that an acknowledged desired revision survived.

Final Challenge

A later proposal makes three changes to “simplify” Aurora:

  1. Store all pod logs and image manifests in the consensus service so every controller reads one system.
  2. Let controllers continue after a watch disconnect because their caches were current recently.
  3. Remove workload-database token checks because the lease record is already strongly consistent.

Review the proposal.

Model answer: Keep desired generation, immutable artifact reference/hash, conditional revision, controller lease/token generation, and dangerous-operation identity inside the authority boundary. Put logs and artifact payloads in observation and artifact systems; their volume would make every consensus commit and recovery more expensive without adding authority. A controller with a watch gap must resync from a snapshot because “recently current” does not identify missed changes. Removing target-side fencing is unsafe: an old controller can resume after a successor obtains a newer token. The failure test pauses token 8, lets token 9 successfully assign a workload, then resumes token 8; the database must reject the old request.

Readiness Rubric

The team may call the architecture ready only when it can answer all of these with a contract, trace, and test:

Question Ready answer
Which state earns consensus cost? It names a conflicting authority or protected decision, not merely importance
What does an acknowledged write mean? Commit/apply boundary and client result semantics are explicit
How does a controller recover? Snapshot revision, watch resume/resync, and idempotent reconciliation are specified
What stops stale work? Fencing token reaches every protected write path
Which operational envelope is required? Tail latency, disk, network, watch, and recovery signals have service-specific budgets
How does membership change? Healthy-quorum transition is distinct from forced recovery after quorum loss
Which fault model is assumed? Crash-fault or Byzantine choice matches the real trust boundary and costs
How can claims fail a test? Workloads, faults, histories, and checkers make violations observable

Resources

Key Takeaways

PREVIOUS Byzantine Consensus and Quorum Certificates