Partition-Time Guarantees: CAP and PACELC

LESSON

Consistency and Replication

001 30 min intermediate

Partition-Time Guarantees: CAP and PACELC

By the end of this lesson, you will be able to...

  • Explain why CAP applies to one operation during a network partition, not to a whole product as a permanent label.

  • Trace the choice between accepting a local request and preserving one authoritative answer when replicas cannot communicate.

  • Use PACELC to state the normal-case latency cost of a stronger replicated-data guarantee.

Idea in one sentence: A replicated service earns a strong answer by coordinating before it replies; when communication breaks, it must sometimes refuse the answer instead.

Core Insight

One concert ticket remains. The service has a healthy replica in Madrid and another in Virginia, and both can receive a checkout request. Its most important promise is simple:

one seat -> at most one successful purchase

It is tempting to say: “Both servers are running, so both should keep serving customers.” That works while the replicas can exchange the information needed to decide who gets the final seat.

Now the link between Madrid and Virginia stops carrying messages. It may be broken or only very slow; the replicas cannot tell which.

The real question is: what may this checkout endpoint promise while it cannot learn what the other side has decided? CAP names that pressure. PACELC adds the normal-case latency question.

The Small Situation

Before the link fails, each replica has the same state:

Madrid:   seat 42 = available, version 17
Virginia: seat 42 = available, version 17

Lucía reaches Madrid and Jordan reaches Virginia. The link fails just before either replica learns about the other request.

Lucía -> Madrid:   buy seat 42

Madrid       X  partition / indistinguishable delay  X       Virginia

Jordan -> Virginia: buy seat 42

Neither replica can infer that the other request did not happen. A timeout is missing information, not proof that the other side is idle.

The Initial Model Breaks

The initial model says: “If a local replica is alive, let it accept the local checkout.” That can work for a wishlist or cached display, where a later conflict is cheap to repair. It fails for this invariant. If both replicas say “purchase confirmed,” the service has already made two incompatible promises; reconciliation cannot make both true.

Separate two questions:

  1. Can this replica return a response without waiting for the other side?
  2. Can this replica still prove that its response preserves the single-seat promise?

During a partition, an endpoint may not be able to answer both with “yes.”

CAP in This Scenario

The familiar letters become clearer when they stay attached to this one operation.

The CAP theorem says a replicated read/write service cannot guarantee both atomic consistency and availability while it tolerates a partition.

A Worked Partition Trace

Here are two possible checkout designs. The state values are a teaching model, not production measurements.

Design A: accept both local requests

Starting state
  Madrid:   available @17
  Virginia: available @17

1. The link fails.
2. Madrid accepts Lucía's request and writes sold-to-Lucía @18.
3. Virginia cannot see @18. It accepts Jordan's request and writes sold-to-Jordan @18.
4. The link heals. Repair finds two incompatible sales.

Both replicas answered, but checkout no longer has one single-copy answer. The system now needs compensation after making incompatible promises.

Design B: require current authority before confirmation

Starting state
  Madrid:   available @17
  Virginia: available @17

1. The link fails.
2. Madrid cannot reach the quorum or current authority needed to confirm a sale.
3. Madrid returns “try again later” rather than “confirmed.”
4. Virginia may confirm only if its own authority rule remains provable; otherwise it also refuses.

This protects the seat invariant but costs availability: a customer can receive an error while a nearby server is alive. The exact mechanism comes later; the contract is already clear: do not say “confirmed” without the evidence that earns it.

So far: CAP makes a hidden product choice visible. For checkout, accepting every local request and preserving one final answer are incompatible while replicas are separated.

PACELC: The Cost Before the Partition

Partitions are not the only time coordination matters. Suppose the link is healthy again.

For illustration, a local durable write may take 8 ms and a remote acknowledgment may add 70 ms. That may be worth it for a scarce allocation but wasteful for a “like” that can converge later.

PACELC is a design framework, not a replacement formalization of the CAP proof:

If there is a Partition: choose Availability or Consistency.
Else: choose Latency or Consistency.

So “we are CP” is only a partial answer: normal-case latency and freshness choices still remain.

Turn the Letters into Endpoint Contracts

One product can make different choices for different promises:

Operation Must never happen Partition-time behavior Normal-case behavior
Confirm final-seat purchase Two customers both receive confirmation Refuse if authority is unclear Wait for the required acknowledgment
Show inventory estimate A customer sees a count that is a few seconds old Serve a labelled stale value Prefer a nearby replica with a freshness budget
Save a wishlist item Two temporary versions exist Accept locally and reconcile Replicate asynchronously
Start a refund workflow Two irreversible refunds start from one request Require durable authority Wait for a recorded workflow decision

The table records the harm to avoid, the failure response, and the normal-case latency cost.

What CAP Does Not Decide

CAP does not choose a shard key, define transaction isolation, guarantee durability, or resolve every conflict. A design still needs a timeout policy, authority rule, and client contract. The next lesson makes those API contracts explicit.

Check Your Understanding

Check: During a partition, Madrid confirms Lucía's purchase from version 17. Virginia cannot see that write and confirms Jordan's purchase from its own version 17. Which promise failed?

Think first, then reveal.

Answer: The final-seat operation lost its single-copy consistency. Both replicas were responsive, but their responses created two incompatible confirmed sales. Repair may resolve the records later; it cannot undo the fact that the service made both promises.

Check: A dashboard can display inventory that is up to two seconds old, but it must show the writer's newly confirmed purchase when the writer refreshes. Is “eventual consistency” alone a sufficient API contract?

Answer: No. “Eventually” does not promise read-your-writes. The writer needs a commit token, a caught-up replica, or a fallback to an authoritative read path.

Practice: Write the Promise Before Choosing the Database

An event-registration service has two endpoints:

POST /registrations
GET  /event-page

There are two seats left. The event page may be slightly stale; a successful registration must never exceed capacity.

For each endpoint, state the invariant or allowed staleness, partition-time response, normal-case latency choice, and evidence needed before replying.

A good answer should mention:

Connections

Resources

Key Takeaways

Partition as a Product Decision

The ticketing service has an invariant:

one physical seat must map to at most one successful purchase

That invariant is stronger than "replicas should converge eventually." If Madrid and Virginia both accept the final-seat purchase during a partition, reconciliation can notice the conflict later, but it cannot make both customers happy without compensation.

So the team has to decide which behavior is acceptable during the broken link:

Madrid replica        partition        Virginia replica
buy final seat   X---------------X     buy final seat

One design chooses consistency over availability for checkout. It may route final-seat writes through a quorum, a single leader, or a lease holder. If the service cannot prove that the write is still authorized, it rejects or delays the operation. Some requests fail, but the seat does not get double-sold.

Another design chooses availability over immediate consistency. Both regions can accept local writes, then reconcile conflicts later. That can be reasonable for a shopping cart, a "like" counter, or a wishlist. It is dangerous for the final purchase unless the business has a clear compensation policy.

The important habit is to attach CAP to a specific operation. A single product can make different choices for different APIs: checkout may stop under partition, while browsing, recommendations, and saved searches keep serving stale or local data.

CAP in Precise Terms

In the CAP framing, the three words are narrower than their everyday meanings:

The common mistake is treating partition tolerance as a feature toggle. In practice, if the system spans more than one failure domain, the network can delay, drop, or reorder messages. The architecture can ignore that possibility, but the production system still has to live through it.

During a partition, a replica cannot know whether a missing message means:

That uncertainty is why CAP bites. A replica that keeps answering every request may be available, but it cannot also guarantee the same strong single-copy story for every operation. A replica that protects the single-copy story must sometimes refuse to answer.

PACELC: The Normal-Case Trade-Off

CAP is essential, but most days the network is not fully partitioned. Requests still cross regions, leaders still wait for followers, and quorums still add latency. PACELC captures that broader shape:

if Partition: choose Availability or Consistency
Else:         choose Latency or Consistency

For the ticketing service, the healthy-network question might be:

Those choices are not just performance tuning. They define what users are allowed to observe. A low-latency local read may show a seat as available after another region has already sold it. A strongly consistent read may avoid that surprise, but it can cost extra round trips and become more sensitive to slow replicas.

PACELC keeps the design discussion honest because it prevents a team from saying "we are CP" and stopping there. A system can reject writes during partitions and still make many different latency-versus-consistency choices when the network is healthy.

Designing the Guarantee

A useful design review starts with a small guarantee table instead of a slogan:

Operation              Partition behavior             Healthy-network behavior
--------------------   -----------------------------  ------------------------
final purchase         reject if authority is unclear  wait for authorized write
inventory display      serve stale with clear budget   prefer local low-latency read
wishlist update        accept locally and reconcile    local write, async replicate
refund initiation      require durable coordination    wait for workflow record

This table does three things.

First, it separates operations by business invariant. Selling the final seat has a different correctness budget from updating a wishlist. Second, it names the user-visible behavior during failure. Third, it records where the team is spending latency to buy stronger consistency.

The weakest acceptable guarantee is usually the best one. Strong consistency is valuable when it protects an invariant the business cannot repair cheaply. Availability and low latency are valuable when users benefit more from progress than from an immediate global truth. The skill is not choosing one word for the whole system; the skill is assigning the right promise to each boundary.

Failure Modes

Mistaking CAP for a database label. A product page saying "CP" or "AP" hides the operation-level decision. Ask what a specific API does when replicas cannot communicate.

Calling every stale read a CAP problem. CAP is about partitions. Stale reads during healthy operation often belong to the PACELC side of the discussion: the team chose latency, locality, caching, or asynchronous replication over a stronger read guarantee.

Assuming reconciliation fixes every conflict. Reconciliation can merge counters, refresh projections, or repair caches. It cannot magically undo a double-booked seat, an irreversible payment, or a violated legal workflow without a compensating process.

Ignoring the client contract. Users do not experience "CAP"; they experience success, failure, waiting, stale state, duplicate confirmations, and later corrections. The architecture should name those outcomes explicitly.

Resources

Key Takeaways

  1. CAP is about partition-time behavior: when replicas cannot communicate, a strongly consistent operation cannot also remain fully available everywhere.
  2. Partition tolerance is not the meaningful knob in real distributed systems; the meaningful choice is what each operation does when communication fails.
  3. PACELC extends the question to normal operation, where stronger consistency usually costs latency even without a partition.
  4. Good designs assign guarantees per API boundary, not as one vague label for the whole database or product.
NEXT Consistency Contracts and API Semantics