Partition-Time Guarantees: CAP and PACELC
LESSON
Partition-Time Guarantees: CAP and PACELC
By the end of this lesson, you will be able to...
Explain why CAP applies to one operation during a network partition, not to a whole product as a permanent label.
Trace the choice between accepting a local request and preserving one authoritative answer when replicas cannot communicate.
Use PACELC to state the normal-case latency cost of a stronger replicated-data guarantee.
Idea in one sentence: A replicated service earns a strong answer by coordinating before it replies; when communication breaks, it must sometimes refuse the answer instead.
Core Insight
One concert ticket remains. The service has a healthy replica in Madrid and another in Virginia, and both can receive a checkout request. Its most important promise is simple:
one seat -> at most one successful purchase
It is tempting to say: “Both servers are running, so both should keep serving customers.” That works while the replicas can exchange the information needed to decide who gets the final seat.
Now the link between Madrid and Virginia stops carrying messages. It may be broken or only very slow; the replicas cannot tell which.
The real question is: what may this checkout endpoint promise while it cannot learn what the other side has decided? CAP names that pressure. PACELC adds the normal-case latency question.
The Small Situation
Before the link fails, each replica has the same state:
Madrid: seat 42 = available, version 17
Virginia: seat 42 = available, version 17
Lucía reaches Madrid and Jordan reaches Virginia. The link fails just before either replica learns about the other request.
Lucía -> Madrid: buy seat 42
Madrid X partition / indistinguishable delay X Virginia
Jordan -> Virginia: buy seat 42
Neither replica can infer that the other request did not happen. A timeout is missing information, not proof that the other side is idle.
The Initial Model Breaks
The initial model says: “If a local replica is alive, let it accept the local checkout.” That can work for a wishlist or cached display, where a later conflict is cheap to repair. It fails for this invariant. If both replicas say “purchase confirmed,” the service has already made two incompatible promises; reconciliation cannot make both true.
Separate two questions:
- Can this replica return a response without waiting for the other side?
- Can this replica still prove that its response preserves the single-seat promise?
During a partition, an endpoint may not be able to answer both with “yes.”
CAP in This Scenario
The familiar letters become clearer when they stay attached to this one operation.
- Consistency, in the CAP result, is a single-copy or atomic story: completed reads and writes behave as though the service had one copy that respects real-time order. It is not ACID's umbrella term.
- Availability, in the formal result, means a request received by a non-failing replica eventually receives a response. A checkout path that must refuse because it cannot establish authority gives up this property for that request.
- Partition tolerance means the service keeps facing a model where messages between replicas may be lost or delayed. It is not a switch that a multi-site system can simply turn off.
The CAP theorem says a replicated read/write service cannot guarantee both atomic consistency and availability while it tolerates a partition.
A Worked Partition Trace
Here are two possible checkout designs. The state values are a teaching model, not production measurements.
Design A: accept both local requests
Starting state
Madrid: available @17
Virginia: available @17
1. The link fails.
2. Madrid accepts Lucía's request and writes sold-to-Lucía @18.
3. Virginia cannot see @18. It accepts Jordan's request and writes sold-to-Jordan @18.
4. The link heals. Repair finds two incompatible sales.
Both replicas answered, but checkout no longer has one single-copy answer. The system now needs compensation after making incompatible promises.
Design B: require current authority before confirmation
Starting state
Madrid: available @17
Virginia: available @17
1. The link fails.
2. Madrid cannot reach the quorum or current authority needed to confirm a sale.
3. Madrid returns “try again later” rather than “confirmed.”
4. Virginia may confirm only if its own authority rule remains provable; otherwise it also refuses.
This protects the seat invariant but costs availability: a customer can receive an error while a nearby server is alive. The exact mechanism comes later; the contract is already clear: do not say “confirmed” without the evidence that earns it.
So far: CAP makes a hidden product choice visible. For checkout, accepting every local request and preserving one final answer are incompatible while replicas are separated.
PACELC: The Cost Before the Partition
Partitions are not the only time coordination matters. Suppose the link is healthy again.
- A Madrid checkout can return after one local durable write. It is fast, but Virginia learns about it later.
- The same checkout can wait for a remote acknowledgment or quorum. That gives stronger cross-replica evidence but adds network delay.
For illustration, a local durable write may take 8 ms and a remote acknowledgment may add 70 ms. That may be worth it for a scarce allocation but wasteful for a “like” that can converge later.
PACELC is a design framework, not a replacement formalization of the CAP proof:
If there is a Partition: choose Availability or Consistency.
Else: choose Latency or Consistency.
So “we are CP” is only a partial answer: normal-case latency and freshness choices still remain.
Turn the Letters into Endpoint Contracts
One product can make different choices for different promises:
| Operation | Must never happen | Partition-time behavior | Normal-case behavior |
|---|---|---|---|
| Confirm final-seat purchase | Two customers both receive confirmation | Refuse if authority is unclear | Wait for the required acknowledgment |
| Show inventory estimate | A customer sees a count that is a few seconds old | Serve a labelled stale value | Prefer a nearby replica with a freshness budget |
| Save a wishlist item | Two temporary versions exist | Accept locally and reconcile | Replicate asynchronously |
| Start a refund workflow | Two irreversible refunds start from one request | Require durable authority | Wait for a recorded workflow decision |
The table records the harm to avoid, the failure response, and the normal-case latency cost.
What CAP Does Not Decide
CAP does not choose a shard key, define transaction isolation, guarantee durability, or resolve every conflict. A design still needs a timeout policy, authority rule, and client contract. The next lesson makes those API contracts explicit.
Check Your Understanding
Check: During a partition, Madrid confirms Lucía's purchase from version 17. Virginia cannot see that write and confirms Jordan's purchase from its own version 17. Which promise failed?
Think first, then reveal.
Answer: The final-seat operation lost its single-copy consistency. Both replicas were responsive, but their responses created two incompatible confirmed sales. Repair may resolve the records later; it cannot undo the fact that the service made both promises.
Check: A dashboard can display inventory that is up to two seconds old, but it must show the writer's newly confirmed purchase when the writer refreshes. Is “eventual consistency” alone a sufficient API contract?
Answer: No. “Eventually” does not promise read-your-writes. The writer needs a commit token, a caught-up replica, or a fallback to an authoritative read path.
Practice: Write the Promise Before Choosing the Database
An event-registration service has two endpoints:
POST /registrations
GET /event-page
There are two seats left. The event page may be slightly stale; a successful registration must never exceed capacity.
For each endpoint, state the invariant or allowed staleness, partition-time response, normal-case latency choice, and evidence needed before replying.
A good answer should mention:
- Capacity authority and a retryable failure for
POST /registrationswhen authority cannot be established. - Bounded staleness for
GET /event-page; its count is not proof that registration succeeded. - A client-visible confirmation token or freshness timestamp.
Connections
- Consistency Contracts and API Semantics names the client-visible histories an API may promise after this lesson establishes the partition-time choice.
- Later lessons show mechanisms that can earn these promises: ordered apply, synchronous acknowledgment, quorums, repair, and safe reads.
Resources
- [PAPER] Brewer's Conjecture and the Feasibility of Consistent, Available, Partition-Tolerant Web Services — Focus: Read the formal definitions of atomic consistency, availability, and partition tolerance before applying the slogan to a product.
- [ARTICLE/PAPER] Consistency Tradeoffs in Modern Distributed Database System Design: CAP Is Only Part of the Story — Focus: Use PACELC to separate partition-time availability/consistency choices from normal-case latency/consistency choices.
- [PAPER] Linearizability: A Correctness Condition for Concurrent Objects — Focus: Compare the single-copy real-time model with weaker client-visible guarantees.
- [BOOK] Designing Data-Intensive Applications — Focus: Connect the formal models to replication, leader, quorum, and application-contract decisions.
Key Takeaways
- CAP applies to a replicated operation under a partition; it is not a personality label for an entire product.
- When two replicas cannot communicate, a final-seat checkout must choose between accepting every local request and preserving one authoritative confirmation.
- PACELC adds the everyday question: how much latency should a healthy system spend to make a stronger replicated-data promise?
- The useful output is an endpoint contract that names the invariant, failure response, normal-case cost, and evidence behind success.
Partition as a Product Decision
The ticketing service has an invariant:
one physical seat must map to at most one successful purchase
That invariant is stronger than "replicas should converge eventually." If Madrid and Virginia both accept the final-seat purchase during a partition, reconciliation can notice the conflict later, but it cannot make both customers happy without compensation.
So the team has to decide which behavior is acceptable during the broken link:
Madrid replica partition Virginia replica
buy final seat X---------------X buy final seat
One design chooses consistency over availability for checkout. It may route final-seat writes through a quorum, a single leader, or a lease holder. If the service cannot prove that the write is still authorized, it rejects or delays the operation. Some requests fail, but the seat does not get double-sold.
Another design chooses availability over immediate consistency. Both regions can accept local writes, then reconcile conflicts later. That can be reasonable for a shopping cart, a "like" counter, or a wishlist. It is dangerous for the final purchase unless the business has a clear compensation policy.
The important habit is to attach CAP to a specific operation. A single product can make different choices for different APIs: checkout may stop under partition, while browsing, recommendations, and saved searches keep serving stale or local data.
CAP in Precise Terms
In the CAP framing, the three words are narrower than their everyday meanings:
- Consistency means a strong single-copy story: reads and writes behave as if there were one correct copy of the data.
- Availability means every request to a non-failed node receives a non-error response.
- Partition tolerance means the system has a defined behavior even when messages between nodes are lost or delayed.
The common mistake is treating partition tolerance as a feature toggle. In practice, if the system spans more than one failure domain, the network can delay, drop, or reorder messages. The architecture can ignore that possibility, but the production system still has to live through it.
During a partition, a replica cannot know whether a missing message means:
- the other side accepted a conflicting write
- the other side is slow
- the other side is isolated
- the local side is the isolated one
That uncertainty is why CAP bites. A replica that keeps answering every request may be available, but it cannot also guarantee the same strong single-copy story for every operation. A replica that protects the single-copy story must sometimes refuse to answer.
PACELC: The Normal-Case Trade-Off
CAP is essential, but most days the network is not fully partitioned. Requests still cross regions, leaders still wait for followers, and quorums still add latency. PACELC captures that broader shape:
if Partition: choose Availability or Consistency
Else: choose Latency or Consistency
For the ticketing service, the healthy-network question might be:
- Should checkout wait for cross-region confirmation before returning success?
- Should reads of remaining inventory go to a local replica even if it may be slightly stale?
- Should the app show "almost sold out" from a cached projection while the purchase path uses stricter coordination?
Those choices are not just performance tuning. They define what users are allowed to observe. A low-latency local read may show a seat as available after another region has already sold it. A strongly consistent read may avoid that surprise, but it can cost extra round trips and become more sensitive to slow replicas.
PACELC keeps the design discussion honest because it prevents a team from saying "we are CP" and stopping there. A system can reject writes during partitions and still make many different latency-versus-consistency choices when the network is healthy.
Designing the Guarantee
A useful design review starts with a small guarantee table instead of a slogan:
Operation Partition behavior Healthy-network behavior
-------------------- ----------------------------- ------------------------
final purchase reject if authority is unclear wait for authorized write
inventory display serve stale with clear budget prefer local low-latency read
wishlist update accept locally and reconcile local write, async replicate
refund initiation require durable coordination wait for workflow record
This table does three things.
First, it separates operations by business invariant. Selling the final seat has a different correctness budget from updating a wishlist. Second, it names the user-visible behavior during failure. Third, it records where the team is spending latency to buy stronger consistency.
The weakest acceptable guarantee is usually the best one. Strong consistency is valuable when it protects an invariant the business cannot repair cheaply. Availability and low latency are valuable when users benefit more from progress than from an immediate global truth. The skill is not choosing one word for the whole system; the skill is assigning the right promise to each boundary.
Failure Modes
Mistaking CAP for a database label. A product page saying "CP" or "AP" hides the operation-level decision. Ask what a specific API does when replicas cannot communicate.
Calling every stale read a CAP problem. CAP is about partitions. Stale reads during healthy operation often belong to the PACELC side of the discussion: the team chose latency, locality, caching, or asynchronous replication over a stronger read guarantee.
Assuming reconciliation fixes every conflict. Reconciliation can merge counters, refresh projections, or repair caches. It cannot magically undo a double-booked seat, an irreversible payment, or a violated legal workflow without a compensating process.
Ignoring the client contract. Users do not experience "CAP"; they experience success, failure, waiting, stale state, duplicate confirmations, and later corrections. The architecture should name those outcomes explicitly.
Resources
- [PAPER] Brewer's Conjecture and the Feasibility of Consistent, Available, Partition-Tolerant Web Services
- Focus: Use this as the formal anchor for the CAP result and its assumptions.
- [ARTICLE] CAP Twelve Years Later: How the "Rules" Have Changed
- Focus: Notice how Brewer reframes CAP as a nuanced design discussion, not a slogan.
- [ARTICLE] Problems with CAP, and Yahoo's Little Known NoSQL System
- Focus: Read for the PACELC framing and the normal-case latency/consistency trade-off.
- [BOOK] Designing Data-Intensive Applications
- Focus: Review the chapters on replication and consistency models for practical examples.
Key Takeaways
- CAP is about partition-time behavior: when replicas cannot communicate, a strongly consistent operation cannot also remain fully available everywhere.
- Partition tolerance is not the meaningful knob in real distributed systems; the meaningful choice is what each operation does when communication fails.
- PACELC extends the question to normal operation, where stronger consistency usually costs latency even without a partition.
- Good designs assign guarantees per API boundary, not as one vague label for the whole database or product.