Exactly-Once Semantics, Idempotency, and Deduplication

LESSON

Consensus and Coordination

013 30 min intermediate

Exactly-Once Semantics, Idempotency, and Deduplication

By the end of this lesson, you will be able to...

  • Trace why a lost response turns a retry into an ambiguous operation.

  • Design a bounded retry contract with stable identity, durable outcomes, idempotency, and deduplication.

  • Find where an “exactly once” claim stops applying and name the state needed to support it.

Idea in one sentence: Safe retries come from making one logical operation recognizable and repeatable inside a declared boundary, not from assuming that the network will deliver it once.

Core Insight

A checkout service asks a payment provider to charge EUR 40 for order 731. The provider completes the charge, but the reply is lost. After a timeout, the checkout service knows only this:

request sent
reply not observed

It does not know whether the provider received nothing, charged the card and lost the reply, or is still working. Retrying with a fresh request may collect the missing payment. It may also create a second charge.

The first model is tempting: send each command once, acknowledge it after success, and call the workflow “exactly once.” That model fits a process with no crashes and a network with no ambiguous timeouts. The lost reply is evidence that the model is too weak. Delivery count is not the same as effect count, and absence of a reply is not evidence that an effect did not happen.

A stronger design gives all attempts for one logical operation the same identity. The receiver stores one authoritative outcome for that identity and returns that outcome to later attempts. This makes retries converge on one business result within a named scope.

The Promise We Need to Keep

For order 731, the checkout service needs a precise client-visible contract:

Repeating the same charge operation with the same key and parameters will not create another charge while that key remains valid. A retry will return the first operation's stored outcome or a clear in-progress result.

This promise says more than “the queue delivers once.” It names the logical operation, repeated behavior, parameter rule, returned result, and retention boundary.

It also says less than “the entire business workflow happens exactly once.” The payment provider may protect the charge while an email provider sends two receipts. Each effect boundary needs its own contract.

Three ideas support the promise:

These ideas cooperate, but they are not synonyms. A dedupe cache can recognize a request without making an external charge idempotent. An idempotent set status = paid operation may be safe even if two messages reach it. A stream engine may update its managed state exactly once while an external API call remains outside that guarantee.

Why “Acknowledge After Success” Breaks

Consider this initial worker design:

1. receive ChargeOrder(731, EUR 40)
2. call payment provider
3. write PAID in the order database
4. acknowledge the queue message

Each local step looks sensible. The failure appears between steps.

The following trace uses illustrative times:

Time Checkout worker sees Payment provider sees Queue sees
10:00:00.000 Receives message m-88 No request yet m-88 is in flight
10:00:00.030 Sends charge request Starts a EUR 40 charge No acknowledgement
10:00:00.110 Still waiting Charge succeeds as pay-501 No acknowledgement
10:00:02.000 Times out; result unknown Charge remains successful No acknowledgement
10:00:05.000 Worker crashes Still has pay-501 Makes m-88 available again

The queue is correct to redeliver. It cannot infer that the external charge succeeded. The worker is also correct to distrust the missing reply. Neither component has enough evidence to decide whether a second, newly identified charge is safe.

Acknowledging before the provider call only reverses the danger: a crash can now lose the charge. Acknowledging after the call risks repetition. Changing the acknowledgement position cannot atomically join two independent systems.

The missing design element is a durable identity that crosses the retry boundary and names the business operation, not merely one delivery attempt.

Build the Bounded Retry Contract

1. Name the logical operation

The checkout service creates one stable key before its first attempt:

operation_key = "charge:order-731:v1"

Every retry carries that same key. A random key generated inside the retry loop would identify attempts, not the logical operation, and deduplication would fail.

The receiver must also bind the key to the relevant parameters. Reusing charge:order-731:v1 with EUR 55 should return a conflict, not silently reuse the EUR 40 result. Stable identity means “the same intended operation,” not “any request that happens to reuse this string.”

2. Store one durable state for the key

A useful teaching model is a small state machine:

UNSEEN
  -> IN_PROGRESS(amount=40)
  -> SUCCEEDED(provider_id=pay-501, amount=40)

UNSEEN
  -> IN_PROGRESS(amount=40)
  -> RETRYABLE_FAILURE(reason=temporary_unavailable)

The record must live long enough to cover realistic late retries. If it exists only in worker memory, a restart erases the evidence. If it expires after ten minutes while clients may retry for a day, an old attempt can look new.

The state transition that claims an unseen key must be conditional. Two concurrent attempts must not both observe UNSEEN and start separate effects. A uniqueness constraint, compare-and-swap, or transaction can make one attempt the owner and make the other observe IN_PROGRESS.

3. Make the effect repeatable at its own boundary

If the payment provider accepts idempotency keys, the checkout service forwards the same operation key on every provider call. The provider can then associate retries with one stored result. Official Stripe documentation, for example, specifies that retries using the same key return the saved result and that keys have a retention policy. That is a concrete API contract, not a property of HTTP delivery.

If the provider offers no idempotent operation or status lookup, the checkout database cannot prove whether a timed-out external charge happened. A local dedupe row closes the local race; it does not make an uncontrolled external side effect atomic.

This is the boundary to state honestly:

checkout database transaction | payment provider API
          controlled           | separately controlled

The design needs an idempotent provider contract, a provider operation identifier that can be queried, or a reconciliation path that can resolve UNKNOWN. Calling the local transaction “exactly once” does not bridge the gap.

4. Return the stored outcome

A duplicate request should not merely receive “ignored.” The client may have missed the only successful reply. Return the stored response or an explicit current state:

same key + same parameters + SUCCEEDED
  -> return pay-501, already completed

same key + same parameters + IN_PROGRESS
  -> return in progress; retry or poll according to contract

same key + different parameters
  -> return conflict; key is bound to another operation

Persisting the outcome makes the API deterministic for retries. It also turns the dedupe record into part of the client-visible contract rather than an invisible performance cache.

Worked Path: The Lost Reply Revisited

Assume the provider supports stable idempotency keys and retains the first result. This is an explicit assumption for the example.

Step Attempt Durable evidence Decision
1 Worker receives m-88 No local record for charge:order-731:v1 Insert IN_PROGRESS with amount EUR 40
2 Worker calls provider with the key Provider records SUCCEEDED(pay-501) Create one charge
3 Provider reply is lost Provider still has pay-501; local row is IN_PROGRESS Outcome is locally unknown; do not invent success or failure
4 Queue redelivers m-88 Same key and same amount are present Reuse the key instead of creating a new operation
5 Worker retries provider call Provider finds the stored result Return pay-501; create no new charge
6 Worker commits local outcome and acknowledges Local row is SUCCEEDED(pay-501) Later retries return the same result

The retry still happens. The message may still be delivered more than once. What converges is the business effect and its observable result inside the declared retention and provider contract.

So far, the stronger model has replaced “one delivery” with “one durable operation identity.” That matters because crashes and timeouts can repeat attempts, but repeated attempts no longer need to create repeated effects.

Where Exactly-Once Semantics Fits

Some systems control a larger transactional boundary. A stream processor may coordinate replayable input positions, managed operator state, and a transactional output sink. After recovery, the system can replay an event while ensuring it affects managed state or the sink once.

That is a meaningful exactly-once guarantee. It remains scoped. Apache Flink's documentation makes the distinction explicit: exactly once for managed state does not mean each event is physically processed once, and end-to-end exactly once requires replayable sources plus transactional or idempotent sinks.

Use this review question for any exactly-once claim:

Which source, state, and sink share the commit or recovery boundary?
What happens at the first side effect outside it?

If the answer is only “the broker supports exactly once,” the client-visible contract is incomplete.

Compare the Design Options

Design What it can promise State and cost Boundary
Blind retry Eventual attempts Little extra state Non-idempotent effects may repeat
Natural idempotency Repetition converges, such as set status = paid Usually simple Only operations with that effect shape
Durable deduplication Repeated keys can reuse or suppress an outcome Identity, storage, atomic claim, retention Fails after expiry or outside the protected effect
Coordinated transaction One commit across participating state and output Coordination, latency, compatible participants Cannot include arbitrary external systems
Idempotent external API Safe retry at the provider boundary Provider key and result retention Limited to the provider's documented scope

The preferred design depends on the effect. For a database status update, a conditional write may be enough. For a payment, stable identity and the provider's retry contract are essential. For a managed stream pipeline, coordinated checkpoints and a transactional or idempotent sink may justify an exactly-once claim.

Trade-offs, Limits, and Signals

The central trade-off is that safe retries buy predictable recovery by turning operation identity and outcomes into durable state, but that state adds storage, coordination, and lifecycle obligations:

Useful signals include duplicate-hit rate, key-conflict rate, operations stuck in IN_PROGRESS, age of the oldest unresolved operation, dedupe-store write failures, provider lookups for ambiguous outcomes, and retries arriving after key expiry.

The most dangerous signal is an unresolved operation older than the provider's lookup or key-retention window. At that point, automated retry may no longer distinguish “not done” from “done but forgotten.” The system needs reconciliation or human review rather than optimistic repetition.

Check Your Understanding

Check: A client times out and retries with a new random idempotency key. Will the receiver reliably recognize the retry?

Think first, then reveal.

Answer: No. The new key identifies a new operation. A retry must reuse the identity created for the original logical operation.

Check: A stream processor checkpoints its local state but sends emails through a non-idempotent external API. Is the email workflow exactly once?

Think first, then reveal.

Answer: No. The checkpoint can protect managed processor state, but the email call crosses into a separate effect boundary. Recovery can repeat it unless the sink is transactional, idempotent, or reconciled through another durable protocol.

Practice: Review a Reservation API

An inventory API receives Reserve(sku-9, quantity=2). Clients may retry for 48 hours. The team proposes keeping request IDs in an in-memory cache for one hour and returning 200 ignored for a duplicate.

Redesign the smallest safe retry contract. State the identity, durable states, concurrency rule, result behavior, retention rule, and one boundary the design cannot protect.

Model answer: Create one reservation key per logical request, scoped by tenant and operation, and require retries to reuse it with identical parameters. Store IN_PROGRESS, SUCCEEDED(reservation_id), and clearly classified failure states durably. Claim an unseen key with a uniqueness constraint or compare-and-swap so concurrent attempts cannot both reserve stock. Return the stored reservation result, not only “ignored.” Retain the record for at least the supported 48-hour retry horizon with an operational margin. This contract still cannot make a separately called shipping or notification API idempotent; that boundary needs its own identity and outcome record.

Connections

The previous lesson tied recovery to snapshots, checkpoints, and log positions. This lesson adds the client boundary: after replay, repeated work must still converge on a safe effect.

The next lesson compares production coordination systems. Their conditional writes and authoritative metadata can help claim an operation key, but consensus does not make an external payment or email atomic. The effect contract still matters.

Resources

Key Takeaways

PREVIOUS Snapshotting, Checkpointing, and Log Compaction NEXT Consensus Systems in Production: etcd, Consul, and ZooKeeper