Exactly-Once Semantics, Idempotency, and Deduplication
LESSON
Exactly-Once Semantics, Idempotency, and Deduplication
By the end of this lesson, you will be able to...
Trace why a lost response turns a retry into an ambiguous operation.
Design a bounded retry contract with stable identity, durable outcomes, idempotency, and deduplication.
Find where an “exactly once” claim stops applying and name the state needed to support it.
Idea in one sentence: Safe retries come from making one logical operation recognizable and repeatable inside a declared boundary, not from assuming that the network will deliver it once.
Core Insight
A checkout service asks a payment provider to charge EUR 40 for order 731. The provider completes the charge, but the reply is lost. After a timeout, the checkout service knows only this:
request sent
reply not observed
It does not know whether the provider received nothing, charged the card and lost the reply, or is still working. Retrying with a fresh request may collect the missing payment. It may also create a second charge.
The first model is tempting: send each command once, acknowledge it after success, and call the workflow “exactly once.” That model fits a process with no crashes and a network with no ambiguous timeouts. The lost reply is evidence that the model is too weak. Delivery count is not the same as effect count, and absence of a reply is not evidence that an effect did not happen.
A stronger design gives all attempts for one logical operation the same identity. The receiver stores one authoritative outcome for that identity and returns that outcome to later attempts. This makes retries converge on one business result within a named scope.
The Promise We Need to Keep
For order 731, the checkout service needs a precise client-visible contract:
Repeating the same charge operation with the same key and parameters will not create another charge while that key remains valid. A retry will return the first operation's stored outcome or a clear in-progress result.
This promise says more than “the queue delivers once.” It names the logical operation, repeated behavior, parameter rule, returned result, and retention boundary.
It also says less than “the entire business workflow happens exactly once.” The payment provider may protect the charge while an email provider sends two receipts. Each effect boundary needs its own contract.
Three ideas support the promise:
- Idempotency is the effect property: repeating the logical operation has no additional effect after the first successful application.
- Deduplication is the recognition mechanism: stable identity plus remembered state lets the receiver detect another attempt.
- Exactly-once semantics is a scoped claim: under stated failure assumptions, one input affects the managed state or output once inside a coordinated boundary.
These ideas cooperate, but they are not synonyms. A dedupe cache can recognize a request without making an external charge idempotent. An idempotent set status = paid operation may be safe even if two messages reach it. A stream engine may update its managed state exactly once while an external API call remains outside that guarantee.
Why “Acknowledge After Success” Breaks
Consider this initial worker design:
1. receive ChargeOrder(731, EUR 40)
2. call payment provider
3. write PAID in the order database
4. acknowledge the queue message
Each local step looks sensible. The failure appears between steps.
The following trace uses illustrative times:
| Time | Checkout worker sees | Payment provider sees | Queue sees |
|---|---|---|---|
| 10:00:00.000 | Receives message m-88 |
No request yet | m-88 is in flight |
| 10:00:00.030 | Sends charge request | Starts a EUR 40 charge | No acknowledgement |
| 10:00:00.110 | Still waiting | Charge succeeds as pay-501 |
No acknowledgement |
| 10:00:02.000 | Times out; result unknown | Charge remains successful | No acknowledgement |
| 10:00:05.000 | Worker crashes | Still has pay-501 |
Makes m-88 available again |
The queue is correct to redeliver. It cannot infer that the external charge succeeded. The worker is also correct to distrust the missing reply. Neither component has enough evidence to decide whether a second, newly identified charge is safe.
Acknowledging before the provider call only reverses the danger: a crash can now lose the charge. Acknowledging after the call risks repetition. Changing the acknowledgement position cannot atomically join two independent systems.
The missing design element is a durable identity that crosses the retry boundary and names the business operation, not merely one delivery attempt.
Build the Bounded Retry Contract
1. Name the logical operation
The checkout service creates one stable key before its first attempt:
operation_key = "charge:order-731:v1"
Every retry carries that same key. A random key generated inside the retry loop would identify attempts, not the logical operation, and deduplication would fail.
The receiver must also bind the key to the relevant parameters. Reusing charge:order-731:v1 with EUR 55 should return a conflict, not silently reuse the EUR 40 result. Stable identity means “the same intended operation,” not “any request that happens to reuse this string.”
2. Store one durable state for the key
A useful teaching model is a small state machine:
UNSEEN
-> IN_PROGRESS(amount=40)
-> SUCCEEDED(provider_id=pay-501, amount=40)
UNSEEN
-> IN_PROGRESS(amount=40)
-> RETRYABLE_FAILURE(reason=temporary_unavailable)
The record must live long enough to cover realistic late retries. If it exists only in worker memory, a restart erases the evidence. If it expires after ten minutes while clients may retry for a day, an old attempt can look new.
The state transition that claims an unseen key must be conditional. Two concurrent attempts must not both observe UNSEEN and start separate effects. A uniqueness constraint, compare-and-swap, or transaction can make one attempt the owner and make the other observe IN_PROGRESS.
3. Make the effect repeatable at its own boundary
If the payment provider accepts idempotency keys, the checkout service forwards the same operation key on every provider call. The provider can then associate retries with one stored result. Official Stripe documentation, for example, specifies that retries using the same key return the saved result and that keys have a retention policy. That is a concrete API contract, not a property of HTTP delivery.
If the provider offers no idempotent operation or status lookup, the checkout database cannot prove whether a timed-out external charge happened. A local dedupe row closes the local race; it does not make an uncontrolled external side effect atomic.
This is the boundary to state honestly:
checkout database transaction | payment provider API
controlled | separately controlled
The design needs an idempotent provider contract, a provider operation identifier that can be queried, or a reconciliation path that can resolve UNKNOWN. Calling the local transaction “exactly once” does not bridge the gap.
4. Return the stored outcome
A duplicate request should not merely receive “ignored.” The client may have missed the only successful reply. Return the stored response or an explicit current state:
same key + same parameters + SUCCEEDED
-> return pay-501, already completed
same key + same parameters + IN_PROGRESS
-> return in progress; retry or poll according to contract
same key + different parameters
-> return conflict; key is bound to another operation
Persisting the outcome makes the API deterministic for retries. It also turns the dedupe record into part of the client-visible contract rather than an invisible performance cache.
Worked Path: The Lost Reply Revisited
Assume the provider supports stable idempotency keys and retains the first result. This is an explicit assumption for the example.
| Step | Attempt | Durable evidence | Decision |
|---|---|---|---|
| 1 | Worker receives m-88 |
No local record for charge:order-731:v1 |
Insert IN_PROGRESS with amount EUR 40 |
| 2 | Worker calls provider with the key | Provider records SUCCEEDED(pay-501) |
Create one charge |
| 3 | Provider reply is lost | Provider still has pay-501; local row is IN_PROGRESS |
Outcome is locally unknown; do not invent success or failure |
| 4 | Queue redelivers m-88 |
Same key and same amount are present | Reuse the key instead of creating a new operation |
| 5 | Worker retries provider call | Provider finds the stored result | Return pay-501; create no new charge |
| 6 | Worker commits local outcome and acknowledges | Local row is SUCCEEDED(pay-501) |
Later retries return the same result |
The retry still happens. The message may still be delivered more than once. What converges is the business effect and its observable result inside the declared retention and provider contract.
So far, the stronger model has replaced “one delivery” with “one durable operation identity.” That matters because crashes and timeouts can repeat attempts, but repeated attempts no longer need to create repeated effects.
Where Exactly-Once Semantics Fits
Some systems control a larger transactional boundary. A stream processor may coordinate replayable input positions, managed operator state, and a transactional output sink. After recovery, the system can replay an event while ensuring it affects managed state or the sink once.
That is a meaningful exactly-once guarantee. It remains scoped. Apache Flink's documentation makes the distinction explicit: exactly once for managed state does not mean each event is physically processed once, and end-to-end exactly once requires replayable sources plus transactional or idempotent sinks.
Use this review question for any exactly-once claim:
Which source, state, and sink share the commit or recovery boundary?
What happens at the first side effect outside it?
If the answer is only “the broker supports exactly once,” the client-visible contract is incomplete.
Compare the Design Options
| Design | What it can promise | State and cost | Boundary |
|---|---|---|---|
| Blind retry | Eventual attempts | Little extra state | Non-idempotent effects may repeat |
| Natural idempotency | Repetition converges, such as set status = paid |
Usually simple | Only operations with that effect shape |
| Durable deduplication | Repeated keys can reuse or suppress an outcome | Identity, storage, atomic claim, retention | Fails after expiry or outside the protected effect |
| Coordinated transaction | One commit across participating state and output | Coordination, latency, compatible participants | Cannot include arbitrary external systems |
| Idempotent external API | Safe retry at the provider boundary | Provider key and result retention | Limited to the provider's documented scope |
The preferred design depends on the effect. For a database status update, a conditional write may be enough. For a payment, stable identity and the provider's retry contract are essential. For a managed stream pipeline, coordinated checkpoints and a transactional or idempotent sink may justify an exactly-once claim.
Trade-offs, Limits, and Signals
The central trade-off is that safe retries buy predictable recovery by turning operation identity and outcomes into durable state, but that state adds storage, coordination, and lifecycle obligations:
- Retention costs storage. The key must remain available for the longest supported retry horizon, plus clock and queue-delay margins appropriate to the system.
- Atomic claims cost coordination. Conditional writes or transactions serialize competing attempts for the same key.
- Stored failures need policy. A validation failure may be final, while a temporary provider outage may be retryable. Caching every failure forever can freeze a recoverable operation.
- Identity has a scope. Include tenant, operation type, and version when the same business identifier can name different actions.
- Downstream effects remain separate. One protected charge does not deduplicate an email, shipment, or analytics event.
Useful signals include duplicate-hit rate, key-conflict rate, operations stuck in IN_PROGRESS, age of the oldest unresolved operation, dedupe-store write failures, provider lookups for ambiguous outcomes, and retries arriving after key expiry.
The most dangerous signal is an unresolved operation older than the provider's lookup or key-retention window. At that point, automated retry may no longer distinguish “not done” from “done but forgotten.” The system needs reconciliation or human review rather than optimistic repetition.
Check Your Understanding
Check: A client times out and retries with a new random idempotency key. Will the receiver reliably recognize the retry?
Think first, then reveal.
Answer: No. The new key identifies a new operation. A retry must reuse the identity created for the original logical operation.
Check: A stream processor checkpoints its local state but sends emails through a non-idempotent external API. Is the email workflow exactly once?
Think first, then reveal.
Answer: No. The checkpoint can protect managed processor state, but the email call crosses into a separate effect boundary. Recovery can repeat it unless the sink is transactional, idempotent, or reconciled through another durable protocol.
Practice: Review a Reservation API
An inventory API receives Reserve(sku-9, quantity=2). Clients may retry for 48 hours. The team proposes keeping request IDs in an in-memory cache for one hour and returning 200 ignored for a duplicate.
Redesign the smallest safe retry contract. State the identity, durable states, concurrency rule, result behavior, retention rule, and one boundary the design cannot protect.
Model answer: Create one reservation key per logical request, scoped by tenant and operation, and require retries to reuse it with identical parameters. Store IN_PROGRESS, SUCCEEDED(reservation_id), and clearly classified failure states durably. Claim an unseen key with a uniqueness constraint or compare-and-swap so concurrent attempts cannot both reserve stock. Return the stored reservation result, not only “ignored.” Retain the record for at least the supported 48-hour retry horizon with an operational margin. This contract still cannot make a separately called shipping or notification API idempotent; that boundary needs its own identity and outcome record.
Connections
The previous lesson tied recovery to snapshots, checkpoints, and log positions. This lesson adds the client boundary: after replay, repeated work must still converge on a safe effect.
The next lesson compares production coordination systems. Their conditional writes and authoritative metadata can help claim an operation key, but consensus does not make an external payment or email atomic. The effect contract still matters.
Resources
- [DOC] Stripe: Idempotent Requests — Focus: Saved responses, parameter checks, key reuse, and retention at an external API boundary.
- [DOC] Apache Flink: Fault Tolerance — Focus: The distinction between replay, exactly-once managed state, and end-to-end sink requirements.
- [DOC] Apache Kafka: Design — Message Delivery Semantics — Focus: Producer idempotence, transactions, consumer offsets, and the boundary of Kafka's guarantee.
- [BOOK] Designing Data-Intensive Applications — Focus: Retries, transactions, stream processing, and the limits of end-to-end delivery claims.
Key Takeaways
- A timeout creates an unknown outcome; it does not prove that an operation failed.
- Stable operation identity, durable outcome memory, and an atomic claim let retries converge on one result.
- Idempotency protects the effect, while deduplication recognizes repeated attempts; neither is a synonym for delivery count.
- Exactly-once semantics is meaningful only when its source, state, sink, failure assumptions, and external boundaries are named.
- Retention and unresolved-operation signals are correctness concerns because forgotten identity can make an old retry look new.