Conflict Resolution and Convergence Policies

LESSON

Consistency and Replication

009 30 min advanced

Conflict Resolution and Convergence Policies

By the end of this lesson, you will be able to...

  • distinguish a stale replica from two concurrent versions by inspecting causal history;

  • choose among deterministic selection, merge, rejection, and compensation from a business invariant;

  • design an API and an operational signal that make information loss visible.

Idea in one sentence: Making replicas agree is a mechanical goal; deciding which user intent survives is a domain decision.

Core Insight

Harbor Point stores reservation holds in a leaderless replicated database. During a regional network break, two workflows update hold H-8821 from the same version:

parent v18: status=active, expires_at=09:35

Lisbon trader:
  v19-lis: status=active, expires_at=09:40

Baltimore approval workflow:
  v19-iad: status=confirmed, reservation_id=R-88421

Each write is reasonable from the writer's local view. Neither writer has seen the other write. When the network heals, the database has two children of v18.

The tempting model is: “repair should copy the newest value everywhere.” That model works when one replica is merely behind. It breaks here because neither child is newer in the causal sense. A wall-clock timestamp can select one child, but it cannot show that the selected writer observed the other child's intent.

Harbor Point needs two different decisions:

  1. Detect the relationship: Is one version an ancestor of the other, or are they concurrent?
  2. Apply a policy: If they are concurrent, may the system choose, merge, reject, or compensate?

The first decision belongs to replication metadata. The second belongs to the invariant and to the information available when the branches meet.

First Decide Whether There Is a Conflict

Replica disagreement does not always mean concurrent user intent.

stale replica
v18 -> v19

concurrent branches
v18 -> v19-lis
    `-> v19-iad

In the stale case, v19 descends from v18. Replacing v18 with v19 does not discard a separate decision. In the concurrent case, neither child descends from the other. Choosing either child discards information unless the policy explicitly says that information is replaceable.

A version vector, revision tree, or another causal marker can support this comparison. This lesson uses a simplified ancestry interface:

def compare(left, right):
    if left.version.descends_from(right.version):
        return "left_supersedes_right"
    if right.version.descends_from(left.version):
        return "right_supersedes_left"
    return "concurrent"

The code is a teaching model, not a complete version-vector implementation. Its important claim is smaller: a system needs evidence of observed history if it wants to distinguish succession from concurrency.

Suppose Baltimore's clock reads 09:34:22.900 and Lisbon's reads 09:34:23.100. Lisbon has the later timestamp. That does not prove the trader saw the confirmation. Clock order answers “which timestamp sorts later?” Causal order answers “did this write include knowledge of the other write?” Those are different questions.

This correction matters. If Harbor Point treats v19-iad as stale because its timestamp is smaller, it can erase a confirmed reservation. If it treats a delayed copy of v19-iad as concurrent, it can create needless retries and review work.

State the Invariant Before Choosing a Winner

Harbor Point writes the rule for the hold lifecycle before it writes resolver code:

Once a hold is confirmed, ordinary hold updates must not return it to active or released. A confirmation also creates an externally visible reservation record.

This is a scenario assumption, not a universal reservation rule. Under this rule, confirmed is an absorbing lifecycle state. The later-expiry branch still records real trader intent, but that intent cannot reverse confirmation.

Other Harbor Point data has different meaning:

Data Invariant or product tolerance Candidate policy
Dashboard layout An older preference may be lost Deterministic last-writer-wins
Desks that viewed a hold Membership only grows for this audit field Add-only set union
Reservation lifecycle Confirmation must not be reversed Domain resolver or reject
Free-form audit notes Every note may matter Preserve siblings for review

The set-union row is safe only because this particular field is add-only. If users can remove a desk, plain union can resurrect a removed member. Merge safety comes from the data's rules, not from the word “merge.” A downstream CRDT track develops the algebra needed to make broader merge claims.

Build a Conflict Ledger

Before implementing a resolver, Harbor Point reviews each policy against the same questions. This is the worked design artifact for H-8821.

Policy Information used Result for H-8821 Information lost or deferred Safe under this invariant?
Last-writer-wins (LWW) A deterministic timestamp and tie-breaker active may win because Lisbon's stamp sorts later Confirmation and its side effect can be hidden No
Field merge Both JSON documents Could create status=confirmed with an expiry copied from the active branch The merged object may never have existed or passed validation Not without a domain rule
Domain resolver Causal branches plus lifecycle rule Preserve confirmed; record the extension as superseded intent Requires code, metadata, and tests Yes, under the stated assumption
Reject or preserve siblings Both branches Return a conflict or enqueue review Resolution work moves to a caller or operator Yes, but slower
Compensate Branches plus evidence of external effects Keep the valid state and undo or offset an invalid side effect Compensation can fail and is not instant rollback Only with an owned workflow

LWW is deterministic: replicas given the same candidates and tie-breaker can choose the same value. Determinism gives convergence. It does not give semantic correctness.

A generic field merge is also suspicious. Combining status=confirmed from one branch with expires_at=09:40 from the other produces a document that no workflow wrote. If validation and side effects depend on whole-state transitions, a field-wise merge can invent an impossible state.

The domain resolver uses more information:

def resolve_hold(left, right):
    relation = compare(left, right)
    if relation == "left_supersedes_right":
        return Resolved(left)
    if relation == "right_supersedes_left":
        return Resolved(right)

    if left.status == right.status == "active":
        return resolve_active_extensions(left, right)

    confirmed = branch_with_status("confirmed", left, right)
    if confirmed and confirmation_evidence_is_valid(confirmed):
        return Resolved(
            value=confirmed,
            superseded_intent=other_branch(confirmed, left, right),
        )

    return NeedsReview([left, right])

This pseudocode deliberately refuses to say “confirmed always wins.” It first checks the causal relation. For concurrent branches, it also requires the evidence that Harbor Point's confirmation rule depends on. If that evidence is missing or contradictory, the resolver preserves the branches and asks the owning workflow to decide.

Trace One Resolution from Detection to Repair

Now follow the complete path. The values and times are illustrative.

09:34:20  Lisbon reads parent v18.
09:34:20  Baltimore reads parent v18.
09:34:22  Baltimore writes v19-iad: confirmed.
09:34:23  Lisbon writes v19-lis: active until 09:40.
09:34:30  The network path recovers.
09:34:31  Anti-entropy finds different leaf versions.
09:34:31  Causal comparison says: concurrent.
09:34:32  Lifecycle resolver validates confirmation evidence.
09:34:32  Resolver emits v20: confirmed, based on both branches.
09:34:33  Repair copies v20 to the replica set.
09:34:34  Metric records one domain-resolved conflict and one
          superseded extension intent.

The result version must descend from both branches. Otherwise a later repair pass can rediscover the same sibling and repeat the conflict.

Notice what each layer contributed:

So far, we have not proved that every Harbor Point conflict is safely mergeable. We have shown a narrower result: this conflict has enough causal and domain evidence for one declared policy. That is the level at which a resolver should earn trust.

Make the Policy Part of the API

Conflict handling changes what a successful write means. It should not remain hidden inside storage repair.

API style Suitable use Client-visible behavior
Blind overwrite Replaceable preference Accept that an older preference may disappear under LWW
Conditional command Lifecycle transition Require an expected version; return conflict if the observed version has changed
Merge-aware command Add-only or otherwise proven mergeable data Return the merged version or its identity
Review-required command Ambiguous, high-cost conflict Return or enqueue a durable conflict reference

A condition such as If-Match: v18 prevents a caller from overwriting a newer version that the receiving authority already knows. It does not by itself prevent two isolated replicas from both accepting a write against v18. Preventing that concurrent acceptance requires a write path with stronger authority or coordination. This is the boundary between optimistic conflict handling and invariants that must be protected before acknowledgment.

Harbor Point should therefore ask one question for every invariant:

May two sides temporarily accept conflicting states, provided a later resolver handles them?

If the answer is no—for example, because both sides could allocate the same scarce asset—the service must coordinate or reject writes during the partition. A clever resolver cannot undo every external effect after the fact.

Trade-offs, Failure Boundaries, and Signals

Automatic resolution improves availability and reduces client retries. It costs causal metadata, deterministic policy code, cross-version testing, and visibility for discarded intent.

Preserving or rejecting conflicts protects ambiguous intent. It costs latency, client complexity, queue growth, and possibly human work. Compensation helps after a side effect escaped, but compensation is another fallible distributed workflow. It is not a time machine.

Useful signals connect the mechanism to the user promise:

Signal Question it answers
Concurrent branches by entity type Where are users making decisions from disconnected state?
Outcomes by policy and resolver version Which rules choose, merge, reject, or compensate?
Age of oldest unresolved conflict Is deferred work becoming a correctness or support risk?
Superseded-intent count How often does convergence hide an accepted action?
Compensation failures and age Did an external effect remain inconsistent after resolution?

The boundary signal is not simply “all replicas now match.” Agreement shows convergence. To evaluate the policy, Harbor Point also needs to know which branches existed, which rule ran, what the rule produced, and whether any promised side effect remains unresolved.

Check Your Understanding

Check: Two replicas return different values. Version metadata shows that v22 descends from v21. Should the system invoke the concurrent-conflict policy?

Think first, then reveal.

Answer: No. This is stale-versus-descendant disagreement. Repair may replace v21 with v22. A conflict policy is needed only when neither version causally supersedes the other.

Check: Two concurrent profile-layout writes are safe under deterministic LWW. Does that make LWW safe for a reservation confirmation?

Think first, then reveal.

Answer: No. The layout policy accepts lost preference information. The reservation invariant does not accept losing confirmation or hiding its side effect. The same convergence mechanism can be suitable for one data type and unsafe for another.

Practice: Review a Conflict Policy

A replicated inventory service accepts two concurrent updates for the last camera:

branch A: allocate camera to order O-71; payment captured
branch B: allocate camera to order O-84; payment captured

Design the smallest honest policy. State:

  1. whether the service may choose, merge, reject, or compensate;
  2. which evidence the resolver needs;
  3. what the API reports;
  4. one operational signal;
  5. whether the write contract should change to prevent recurrence.

A strong answer should reject field merge and should not claim that LWW repairs the double allocation. If both payments already happened, the workflow must choose an allocation by a declared business rule, compensate the losing order, and expose both outcomes. The resolver needs causal branches plus payment and allocation identities. The API needs a durable command-status result rather than a silent overwrite. Compensation age or failure rate is a useful signal. Because one physical item cannot satisfy two accepted allocations, the stronger long-term design is to coordinate allocation before acknowledging success, even if that reduces write availability during a partition.

Connections

Resources

Key Takeaways

PREVIOUS Read Repair, Anti-Entropy, and Merkle Divergence Checks NEXT Replication Lag and Read-Your-Writes