Control Plane Consensus Boundary Design Review
LESSON
Control Plane Consensus Boundary Design Review
By the end of this lesson, you will be able to...
Defend which control-plane facts need one authoritative history and which must remain outside consensus.
Walk through a controller update across revision, lease, fencing, replay, and external side-effect boundaries.
Review the architecture with explicit invariants, recovery rules, operational signals, and failure tests.
Idea in one sentence: A control plane earns consensus cost only for small facts whose disagreement grants unsafe authority; controllers then turn those facts into idempotent, fenced, recoverable action outside the consensus path.
Core Insight
Consensus is not a general-purpose importance marker. It is a costly way to establish one history when disagreement would allow incompatible authority. A control plane therefore keeps its desired decision, ownership, and anti-stale-action evidence small and explicit. It lets status, telemetry, and large payloads use systems optimized for observation and transport, then makes the join between those systems safe through versions, hashes, epochs, and replay rules.
The practical test is ownership of a harmful transition. If two answers could safely coexist—for example, two independent health samples—replication into an observation system may be enough. If two answers would let separate controllers roll out different generations or both delete the same protected resource, the system needs an authoritative decision and an enforcement point. This test keeps the boundary comprehensible as requirements grow.
The Scenario
Atlas runs deployments across three clusters. Operators submit a desired rollout: which version should run, how many replicas belong in each cluster, and which generation supersedes the last one. Controllers read that desired state and make changes in external systems.
The platform has a tempting first design: put every important thing in one consensus-backed store. Deployment manifests, live pod status, request logs, metrics, image metadata, controller leases, and every emitted event all feel important.
It works in a small demo. Then one controller pauses, a watch disconnects, a new controller takes over, and telemetry spikes during the recovery. Now the store is expected to be the deployment database, event bus, audit archive, health system, lock service, and source of controller authority at once.
The stronger model is narrower. Consensus protects decisions that must have one authoritative story. Other systems hold bulk data, observations, artifacts, and derived views. Controllers bridge those boundaries with revisions, idempotency, fencing, and recovery rules.
This capstone does not introduce a new consensus algorithm. It asks whether the architecture already built from the preceding lessons keeps the same safety story from API request to recovery and verification.
Constraints
Atlas must satisfy these requirements:
- An acknowledged desired-state change must not disappear after a leader change.
- At most one current controller may perform role-specific, dangerous actions.
- A paused former controller must not be able to overwrite work from its successor.
- Controllers must recover from restart, reconnect, and watch gaps without trusting an old cache.
- Retrying a controller action must not create a second irreversible effect.
- The control plane must remain usable when logs grow and when a minority of nodes is unavailable.
- Operators must be able to test the important claims under pause, partition, dropped reply, restart, and replay.
Some constraints are intentionally outside this capstone. Atlas does not try to make every workload event globally ordered, provide a general analytics system, or teach a complete protocol implementation. Those are separate scalability and data-system problems.
The Boundary We Need to Draw
The decisive question is not “is this data important?” It is “would disagreement about this fact grant conflicting authority or change a protected decision?”
| State | Consensus-backed? | Why | Outside-boundary rule |
|---|---|---|---|
| Desired deployment generation and spec hash | Yes | Controllers need one target to reconcile | Store large manifests elsewhere; retain a content reference and hash here |
| Controller lease owner and fencing epoch | Yes | Two accepted owners can cause conflicting changes | Every protected target rejects older epochs |
| Membership and rollout ownership metadata | Yes | The control plane needs one defensible view | Keep membership observations and logs distinct from authority |
| Idempotency record for a dangerous action | Usually yes or transactionally durable with its effect | Retries must converge on one result | Retain it for the supported retry horizon |
| Current pod status and health observations | Usually no | They are frequent, partial, and can be rebuilt | Treat them as inputs to reconciliation, not authority by themselves |
| Request logs, traces, metrics, and clickstream events | No | They need throughput, retention, and queryability | Preserve them in observability or event systems |
| Images, manifests, and large artifacts | No | Size and transfer do not need a serialized decision path | Consensus stores immutable references and expected hashes |
This boundary makes one subtle distinction visible: a controller can use health evidence to decide what to attempt, but a health report alone does not grant exclusive authority. The authoritative decision and the observation that informs it have different jobs.
The Naive Design and Its Failure
The naive controller keeps a local cache from watch events and treats a lease record as permission forever:
watch desired state
if local lease says "I am leader":
apply every cached change
This works while the watch is continuous and the controller stays connected. It breaks under ordinary partial failure.
Assume controller A read desired generation 41 and acquired fencing epoch 7. A then pauses. Its lease expires, controller B acquires epoch 8, and an operator submits generation 42. B applies the new generation. When A wakes up, its cache may still contain generation 41 and it may still believe it has permission to act.
Two repairs are needed:
- Resynchronize state. A controller rebuilds from a current authoritative snapshot at revision
R, then watches changes afterR. If its watch disconnects or cannot resume after compaction, it starts from a new snapshot rather than trusting its old cache. - Fence external action. The deployment target accepts action only when the supplied epoch is current enough. A's request with epoch
7must be rejected after B's epoch8is accepted.
Neither repair replaces the other. A current cache does not stop a stale actor, and fencing does not tell a controller which desired state it missed.
Proposed Model
Atlas separates authority, observation, and execution:
operator API
|
| conditional write: desired generation + immutable spec reference
v
consensus-backed metadata
| revisions, leases, fencing epochs, idempotency records
v
controllers ----> external deployment API / clusters
| ^
| watches and resync | epoch + operation identity required
v |
observability, health, status, artifacts
The metadata store is authoritative for the desired generation, controller role, and coordination records. It does not become the canonical home of every observed status update. Controllers continuously compare desired state with observed state and issue safe, repeatable actions.
The relevant contracts are explicit:
| Contract | Rule |
|---|---|
| Desired-state write | Update generation only through a conditional write against the revision the operator reviewed; return the committed revision. |
| Controller recovery | Read a complete snapshot at revision R, reconcile, then watch from R + 1; resync after a gap. |
| Leader-only action | Attach the current fencing epoch; the target persists its highest accepted epoch and rejects an earlier one. |
| Retry | Reuse a stable operation ID and return a stored outcome or an explicit in-progress state. |
| External artifact | Store its immutable reference and expected hash in authoritative metadata; fetch the payload outside consensus. |
These are teaching contracts, not a vendor configuration. A real system must choose the exact read mode, transaction semantics, failure detector, retention, and target enforcement that its API documents.
Walkthrough: One Rollout Through the Boundary
The following trace uses illustrative revisions and epochs.
| Step | Authoritative state | Controller action | Safety evidence |
|---|---|---|---|
| 1 | Operator writes payments generation 42, spec hash h42, revision 901 |
API returns success after the store's commit rule | Acknowledged revision 901 |
| 2 | Controller B reads snapshot at 901 and owns role epoch 8 |
It sees desired 42 differs from observed 41 |
Snapshot revision and lease/epoch record |
| 3 | B creates operation_id=deploy:payments:42 |
It records IN_PROGRESS durably before retrying external work |
Stable identity across retries |
| 4 | B calls the deployment API with generation 42, hash h42, epoch 8, and the operation ID |
Target accepts and stores its highest epoch | Fenced target response |
| 5 | Reply is lost after acceptance | B cannot infer success from its timeout | Operation remains unknown or in progress |
| 6 | B retries with the same identity and epoch | Target returns the stored outcome | One external effect, repeatable client result |
| 7 | Watch disconnects during later updates | B reads a new snapshot and reconciles before resuming | New revision boundary, not an assumed event gap |
Notice where consensus stops. It made the desired generation and epoch authoritative. It did not atomically deploy a container, deliver a network reply, or make the external API idempotent by itself. The deployment target and operation record must carry those semantics.
So far, the architecture has a complete first-pass story: consensus decides a small desired state; controllers observe and reconcile; fencing limits stale authority; stable identity makes retries safe; snapshots and resync make recovery bounded.
Failure Review
An architecture is not ready until it states the safety properties it relies on and the fault that could reveal each one.
| Invariant | Failure pressure | What counts as a violation |
|---|---|---|
| Acknowledged desired writes survive recovery | Leader restart, dropped reply, slow disk | A later snapshot omits an acknowledged generation |
| Stale leaders cannot change deployment state | Pause old leader, expire lease, resume it after a successor acts | Target accepts an older fencing epoch |
| Controllers converge after watch gaps | Disconnect, compact history, rapid writes | Final controller cache differs from authoritative state |
| Retries do not duplicate deployment effects | Crash after target call, retry after lost reply | Same operation ID creates two accepted deployments |
| Quorum loss prevents unsafe authority | Partition the metadata cluster | Conflicting authoritative generations are acknowledged |
Use Jepsen-style thinking here. Record client invocations and results, the fault schedule, revisions, epochs, and operation identities. A green dashboard after failover is useful liveness evidence, but it does not disprove a stale action or lost acknowledged write.
Trade-offs and Operational Evidence
The trade-off is deliberate: consensus supplies a defensible history for small control facts, but it adds replication latency, quorum dependence, compaction and snapshot operations, and careful client recovery. Putting bulk state into the same path turns those costs into the platform's hot path.
The boundary becomes risky when these signals appear:
- increasing time from desired-state write to committed revision;
- watch cancellations, compaction responses, or controllers acting from an old snapshot;
- lease keepalive failures, leadership churn, or rejected stale fencing epochs;
- operations stuck in
IN_PROGRESSafter a lost reply or controller restart; - recovery time from snapshot plus log tail exceeding the operational budget;
- metadata-store growth caused by payloads, unbounded event history, or unexpired dedupe records;
- a verification run with unknown outcomes that its checker cannot model.
The right response is not automatically to reduce safety. For example, a large manifest can move to artifact storage while the control plane keeps its immutable reference and hash. A slow watch client can resync from snapshot. A dangerous external target can add epoch enforcement. Each mitigation preserves the authority boundary instead of hiding it.
Evidence and Readiness
Use this review rubric before calling the design ready:
| Question | Ready when | Not ready when |
|---|---|---|
| Does each consensus record earn its cost? | It names authority, ownership, or one desired decision | It is stored there only because it is convenient or important |
| Can a controller restart safely? | It has a snapshot, revision, watch-resume, and resync rule | It relies on a perfect event stream or unbounded cache |
| Can an old controller harm the system? | The target enforces fencing, and retries carry stable identity | The lease record is the only protection |
| Can effects be replayed safely? | The operation boundary returns a durable outcome | The controller assumes a timeout means failure |
| Can the team test the claim? | Invariants, workloads, faults, and checker evidence are named | The plan says only “test failover” |
This is a design defense, not an implementation certification. The next lessons deepen what makes state-machine application, commit evidence, and safe reads work internally. This capstone only requires that the API and operational boundaries already make those future proof obligations visible.
Final Challenge
Atlas proposes one change: store live pod status, ten-minute debug logs, and full image manifests in the consensus store so controllers can read everything from one place. In exchange, it removes fencing because the lease record is “already strongly consistent.”
Review the proposal. State what remains inside consensus, what moves outside, what authoritative references remain, and which invariant would catch the fencing regression.
Model answer: Keep desired deployment generation, the immutable manifest hash/reference, controller role, fencing epoch, and idempotency record for the dangerous action in authoritative metadata. Keep live status, debug logs, and image payloads in systems designed for observation and artifact storage; controllers may read them as reconciliation inputs. Removing fencing is unsafe because a paused former controller can still act after a new lease owner exists. Test by pausing epoch 7, allowing epoch 8 to deploy, then resuming epoch 7; the protected deployment API must reject the older epoch. A successful stale action is the violation.
Connections
The previous lesson supplied the verification loop: contract, workload, fault, history, checker. This capstone applies that loop to a full control-plane boundary.
The next lesson begins the deeper pass. It explains how the committed command history inside the authority boundary becomes deterministic shared service state.
Resources
- [DOC] etcd API — Focus: Revisions, transactions, watches, compaction, and leases for a coordination API.
- [DOC] Kubernetes API Concepts — Focus: Resource versions, watches, desired-state APIs, and client recovery patterns.
- [PAPER] In Search of an Understandable Consensus Algorithm — Focus: The replicated log and commit model behind the authority boundary.
- [DOC] Jepsen Analyses — Focus: Histories, faults, and checkers for production-facing safety claims.
Key Takeaways
- Consensus should protect small authority decisions, not every important byte in the platform.
- A safe controller combines current authoritative state, watch resync, fencing, and idempotent external action.
- A lease transfers coordination authority but does not stop a stale process; the protected target must enforce epochs.
- Snapshots, compaction, and operation identity make recovery and retries bounded rather than hopeful.
- The design is ready to implement only when its authority claims have concrete failure tests and observable evidence.