Final Capstone: Consensus-Backed Control Plane Architecture
LESSON
Final Capstone: Consensus-Backed Control Plane Architecture
By the end of this lesson, you will be able to...
Defend a control-plane authority boundary, including the state that earns consensus cost and the state that does not.
Trace a deployment change from conditional write through commit, deterministic apply, watch recovery, lease authority, and fenced external action.
Review the architecture against safety invariants, operational signals, recovery boundaries, and an explicit crash-fault versus Byzantine decision.
Idea in one sentence: A consensus-backed control plane is credible only when every authority claim has a committed history, a recoverable API path, a protected side-effect boundary, and a failure test that could disprove it.
Core Insight
Aurora's scheduler may receive committed deployment revision 901 while a former scheduler is still paused with lease token 8. The important design move is to follow authority all the way to its consequence. Agreement inside a metadata cluster is not enough if a controller applies the wrong command, a watch misses an update, a former lease holder can still write outside the cluster, or recovery silently changes which history is trusted. The architecture is coherent only when its authority records, state-machine rules, API results, protected targets, operating limits, and failure tests tell one compatible story.
The Scenario
Aurora is a regional compute platform. Operators submit desired deployments. One scheduler per shard turns those deployments into assignments in an external workload database. Agents watch configuration changes. The control plane must survive one zone loss, recover controllers after pauses, and make it possible to explain what happened after a failed rollout.
The team proposes a simple slogan: “Put the control plane in consensus.” That is a useful beginning, not an architecture. It leaves unresolved which facts enter the log, what a successful API response means, whether a returning controller can still act, how watches recover, and what recovery can honestly claim after quorum loss.
This capstone turns the slogan into a defensible design. It does not prescribe a vendor or implement a full protocol. It asks for evidence at each boundary.
Constraints
Aurora must meet these constraints:
- An acknowledged desired deployment change must survive a leader change under the chosen crash-fault protocol.
- A scheduler cannot use stale authority to overwrite a newer scheduler's assignment.
- Controllers reconnecting after a missed watch event rebuild state rather than guessing what they missed.
- External assignment requests are safe to retry after a lost response.
- The coordination cluster remains useful under one zone loss and exposes when latency is outside the lease and control-loop budget.
- Member replacement with quorum preserves continuity through the protocol; recovery without quorum names its data-loss boundary.
- Telemetry, logs, and artifacts do not become part of the critical quorum write path.
- The team states whether its trust boundary needs crash-fault or Byzantine protection.
The non-goals matter too. Aurora is not a globally ordered log for every agent event, a general analytics warehouse, or a proof that every external workload operation happened exactly once. Those needs use other systems and their own contracts.
Design Goal
The architecture must make one narrow promise:
For each protected decision, the platform can identify
the committed authority record, the current actor allowed to use it,
the recovery path if that actor is stale, and the evidence that rejects
an older action at the protected resource.
The reasonable initial design stores every deployment object, heartbeat, pod status, log, lease, and scheduler event in the consensus service. Controllers read whichever local cache they have and call the workload database after seeing an owner key.
It fails in two ways. Bulk data makes every authority decision pay unnecessary storage, replication, compaction, and recovery cost. More importantly, a local cache and an owner key do not stop a paused scheduler from acting after a successor takes over.
The stronger design separates authority, observation, and execution. Consensus produces a small ordered history for decisions whose disagreement would create conflicting authority. Controllers derive work from that history. External resources enforce the authority generation they accept.
Proposed Model
Aurora runs three crash-fault voters, one in each of three zones. A majority can continue after one voter/zone loss if the remaining two are healthy and the network path between them is usable. The exact consensus implementation may be Raft-like, Paxos-like, or another correct crash-fault design; the API contract below does not assume a product.
operator API
| conditional desired-state command
v
consensus-backed metadata and replicated state machine
| committed revisions, lease/token generations, operation records
v
controllers ---- fenced, idempotent requests ----> workload database
^ |
| snapshot + watch recovery | accepted-token high-water mark
| v
status, logs, metrics, artifacts, traces <--- observation systems
Authority Boundary
| State | Inside consensus? | Why | Outside-boundary rule |
|---|---|---|---|
| Desired deployment generation, immutable spec reference/hash | Yes | Controllers need one target to reconcile | Store the bulky manifest payload in artifact storage |
| Shard ownership, lease state, fencing generation | Yes | Conflicting owners can make unsafe assignments | Workload database checks the generation on every protected write |
| Conditional-update revision and membership metadata | Yes | They define current preconditions and who can decide history | Retain and compact according to documented recovery rules |
| Durable operation record for dangerous retries | Yes or transactionally durable with the effect | Retries need one stable outcome/in-progress state | Bound retention to the supported retry horizon |
| Pod health, capacity sample, and current agent status | Usually no | They are frequent observations that can be rebuilt | Use them as reconciliation input, not authority alone |
| Logs, traces, metrics, image layers, large manifests | No | They need throughput, storage, and query features rather than a single decision order | Keep a reference/hash in authority state when integrity matters |
The question behind every row is: “Would two different answers grant incompatible authority or change a protected decision?” Importance alone is not enough to earn consensus cost.
Service and API Contracts
The state machine applies only the committed command prefix. A controller never treats a proposal or a local leader receipt as an applied service result.
| Contract | Aurora's rule |
|---|---|
| Desired-state update | put_if_revision(key, R, command) succeeds only if the inspected revision remains current; success returns committed revision R+1 |
| State-machine apply | Same committed command plus same replicated state produces the same internal result on every replica |
| Current authority read | A leader uses its documented current-read path and serves only state applied through that read boundary |
| Lease loss | Renewal failure, closed session, or uncertainty means the controller stops protected work until a new acquisition succeeds |
| Watch recovery | Read snapshot at revision R, watch after R, and resync from a new snapshot after disconnect/compaction gap |
| External action | Include fencing token and stable operation ID; target rejects older tokens and returns stored result or in-progress state for retries |
| Reconfiguration | Change voting membership through the protocol while a valid old quorum exists; do not use a normal-replacement claim after quorum loss |
These are design contracts, not implementation details to postpone. Without them, an API user cannot tell what evidence a revision, lease, response, or error carries.
Walkthrough: One Deployment Change
The following identifiers and revisions are illustrative.
| Step | Committed or observed state | Action | Evidence and boundary |
|---|---|---|---|
| 1 | payments is desired generation 41, revision 900 |
Operator reads revision 900 and prepares immutable spec hash h42 |
The operator has a precondition, not permanent ownership |
| 2 | Command set payments=42, hash=h42, expected=900 enters the log |
Consensus commits it at revision 901; replicas deterministically apply it |
Desired generation 42 is now shared service state |
| 3 | Scheduler B holds lease generation/token 8 |
B snapshots through 901, sees observed 41, and creates operation_id=deploy:payments:42 |
B has a recoverable state base and current authority evidence |
| 4 | Operation record says IN_PROGRESS |
B calls workload database with generation 42, hash h42, token 8, and operation ID |
Target can distinguish stale authority and a retry from a new operation |
| 5 | Database accepts the action but reply is lost | B treats the client outcome as unknown, not failed | Consensus did not make the network reply atomic |
| 6 | B retries the same operation ID and token | Database returns stored accepted result | One protected effect, repeatable controller result |
| 7 | B's watch later disconnects after revision 905 and cannot resume due to compaction |
B snapshots at 912, replaces cache, then watches after 912 |
Recovery has a known boundary instead of an assumed gap |
This trace connects several lessons. Quorum evidence preserves the command history across future leaders. Deterministic apply makes revision 901 mean the same state to each replica. CAS prevents an operator from silently writing over a newer desired state. Snapshot plus watch makes controller recovery bounded. Fencing makes the workload database refuse a stale scheduler. Idempotency handles a lost reply without pretending that a timeout identifies the result.
Failure Review
An architecture becomes credible when each failure has a safety statement, an expected behavior, and a signal that reveals a broken assumption.
| Failure pressure | Expected behavior | Violation or signal |
|---|---|---|
| Leader crashes after a client request but before commit evidence | Client does not receive a durable-success claim; later history may omit the proposal | An acknowledged desired revision disappears after recovery |
Old scheduler pauses; successor gets token 9 after token 8 |
Workload database accepts 9 and rejects later 8 |
A protected write succeeds with token below its high-water mark |
| Controller watch disconnects and history is compacted | Controller fetches new snapshot and reconciles | Controller continues from an unproven old cache |
| Slow leader disk or overloaded write path | Tail commit latency, retries, and watch pressure rise; controllers fail safe on lease uncertainty | Quorum exists but lease/control-loop budget is exceeded |
| One zone fails | Remaining two voters may continue if their path is healthy | Minority side claims it can commit new authority |
| One voter is lost while a quorum remains | Replacement catches up and membership transition is committed by the protocol | An operator file edit creates an independent voter set |
| Quorum is permanently lost | Ordinary progress stops; recovery source and data-loss boundary are authorized and recorded | Forced bootstrap is described as fully continuous without evidence |
The values in a production alert policy depend on the lease duration, client workload, and recovery objective. What must not vary is the causal interpretation: a leader indicator alone does not prove the cluster can still supply authority within the needed time.
Trade-offs and Fault Model
The trade-off is deliberate. Consensus gives a compact, defensible history for authority. It costs quorum latency, durable storage behavior, careful client recovery, and operational discipline around member placement and replacement. The small metadata boundary reduces that cost; it does not remove it.
Three voters across three zones are a reasonable crash-fault choice when Aurora trusts the operators and machines as non-malicious participants and wants to survive one zone loss. The design must still handle software bugs, slow disks, partitions, stale controllers, and recovery mistakes; crash-fault consensus does not make those disappear.
Aurora should consider Byzantine protocols only if its actual trust boundary includes replicas or operators that may intentionally equivocate, forge operational claims through compromised identities, or act adversarially across organizations. That choice requires stable identities, key lifecycle control, authenticated votes, certificate verification, different quorum thresholds, and a more complex operational model. BFT is not a remedy for an oversized watch workload or an unfenced workload database.
The situated decision for the stated regional platform is crash-fault consensus plus strong client and target-side enforcement. It becomes risky when the administration boundary is no longer trusted; the signal is a threat model that includes arbitrary participant behavior rather than ordinary crash and delay.
Evidence and Readiness
Aurora must test the claims, not merely check that a leader election occurred. Its verification plan records client invocation/result histories, revisions, commit/apply boundaries, lease generations, operation IDs, and workload-database acceptance decisions.
| Invariant | Workload and fault | Evidence of success |
|---|---|---|
| Desired state is not lost after acknowledgement | Conditional writes during leader crash and delayed replies | Every acknowledged revision appears in a later authoritative snapshot |
| Only current scheduler can change assignments | Pause old scheduler, issue new lease, resume old scheduler | Target rejects any action with an older token after newer token succeeds |
| Controllers converge after watch gaps | Disconnect, compact history, make rapid desired-state changes | Cache rebuilt from snapshot equals authoritative state |
| Retry does not duplicate protected work | Crash/lost reply after database acceptance | Same operation identity returns one stored result/effect |
| Reconfiguration preserves continuity with quorum | Replace a voter after snapshot/log catch-up | New configuration retains committed history through a leader change |
| Forced recovery is honest | Simulate permanent quorum loss and bootstrap chosen source | Runbook records source, recovery revision, clients needing reconciliation, and old-authority fencing |
Jepsen-style fault injection is useful because it treats the history and checker as first-class evidence. A dashboard becoming green after a restart is liveness evidence. It does not prove that an old token was rejected or that an acknowledged desired revision survived.
Final Challenge
A later proposal makes three changes to “simplify” Aurora:
- Store all pod logs and image manifests in the consensus service so every controller reads one system.
- Let controllers continue after a watch disconnect because their caches were current recently.
- Remove workload-database token checks because the lease record is already strongly consistent.
Review the proposal.
Model answer: Keep desired generation, immutable artifact reference/hash, conditional revision, controller lease/token generation, and dangerous-operation identity inside the authority boundary. Put logs and artifact payloads in observation and artifact systems; their volume would make every consensus commit and recovery more expensive without adding authority. A controller with a watch gap must resync from a snapshot because “recently current” does not identify missed changes. Removing target-side fencing is unsafe: an old controller can resume after a successor obtains a newer token. The failure test pauses token 8, lets token 9 successfully assign a workload, then resumes token 8; the database must reject the old request.
Readiness Rubric
The team may call the architecture ready only when it can answer all of these with a contract, trace, and test:
| Question | Ready answer |
|---|---|
| Which state earns consensus cost? | It names a conflicting authority or protected decision, not merely importance |
| What does an acknowledged write mean? | Commit/apply boundary and client result semantics are explicit |
| How does a controller recover? | Snapshot revision, watch resume/resync, and idempotent reconciliation are specified |
| What stops stale work? | Fencing token reaches every protected write path |
| Which operational envelope is required? | Tail latency, disk, network, watch, and recovery signals have service-specific budgets |
| How does membership change? | Healthy-quorum transition is distinct from forced recovery after quorum loss |
| Which fault model is assumed? | Crash-fault or Byzantine choice matches the real trust boundary and costs |
| How can claims fail a test? | Workloads, faults, histories, and checkers make violations observable |
Resources
- [DOC] etcd Documentation — Focus: Concrete API, watch, lease, recovery, and operational semantics for a coordination system.
- [DOC] Kubernetes API Concepts — Focus: Resource versions, watches, and reconciliation boundaries in a control-plane API.
- [PAPER] The Chubby Lock Service for Loosely-Coupled Distributed Systems — Focus: Coordination API semantics and the client-visible authority contract.
- [DOC] Jepsen Analyses — Focus: Fault histories and checkers as evidence for distributed-system claims.
- [PAPER] HotStuff: BFT Consensus in the Lens of Blockchain — Focus: The additional authenticated evidence and quorum structure needed for an adversarial trust boundary.
Key Takeaways
- Consensus belongs on small authority decisions; observations, telemetry, and bulky artifacts need other storage paths.
- A credible control plane connects committed history, deterministic apply, guarded APIs, watch recovery, and fenced external action.
- Quorum survival is not enough when the cluster cannot meet the latency and recovery budgets of its authority contracts.
- Normal replacement preserves continuity through a healthy quorum; forced recovery chooses and documents a recovery boundary.
- The architecture is complete only when its authority claims, fault model, and operational limits are testable under the failures it expects.