Blog · Data engineering
Idempotent ingestion: exactly-once is something you build
October 08, 2026 · The Murmurator team
Read the delivery guarantees of any queue, webhook provider or event bus and you'll find the same sentence in different words: messages are delivered at least once. Duplicates are not a failure mode to be handled someday. They are the documented, expected behaviour.
Most pipelines are written as though this weren't true, and get away with it until the day a consumer times out after doing its work but before acknowledging. Then the same payment posts twice.
"Exactly-once processing" is achievable, but not as a delivery guarantee. It's a property of the consumer: at-least-once delivery plus idempotent handling.
Where the duplicates come from
Worth knowing, because the sources need different handling.
Redelivery after a failed ack. The consumer did the work, then died before acknowledging. The broker cannot distinguish this from a consumer that died before doing anything, so it redelivers. This is the common case and it is unavoidable.
Producer retries. An HTTP call times out on the client side. The server may have processed it. The client retries because it has no other option.
Provider retries. Webhook senders retry on non-2xx, and many retry on slow responses they treat as failures. A handler that takes nine seconds and returns 200 may already have a duplicate in flight.
Replays. Someone reprocesses yesterday's data to fix a bug. Every downstream effect happens again.
Upstream at-least-once. Your source system has the same problem, so its outputs already contain duplicates before you see them.
The mechanism
Every write needs a key that identifies the work, not the attempt, and the write needs to be conditional on that key being new.
Picking the key is the part that takes thought:
- Provider-supplied event ID. Best option when available. Stripe, GitHub and most mature webhook senders include one and reuse it across retries.
- A natural key from the payload.
(account_id, invoice_number),(sensor_id, reading_timestamp). Works when the domain has a genuine uniqueness constraint. - A content hash. Hash the canonicalised payload. Legitimate distinct events that happen to be byte-identical will collide — a temperature sensor reporting the same value twice in a second is a real event, not a duplicate — so this is a last resort, and canonicalisation must be stable.
What does not work: arrival timestamps, broker message IDs that change on redelivery, or an auto-increment assigned by the consumer. These identify the attempt.
Then enforce it in the database, not in application logic:
insert into events (event_id, account_id, kind, payload, received_at)
values ($1, $2, $3, $4, now())
on conflict (event_id) do nothing;
The unique constraint on event_id is the guarantee. A check-then-insert in application code is a race — two workers both check, both find nothing, both insert. This bug is invisible in testing and appears under load.
Idempotency for things that aren't database rows
Rows are the easy case. The hard cases are effects you don't own.
Outbound HTTP. Many APIs accept an idempotency key header and will return the original response for a repeat. Use it. Derive the key from your event ID plus the operation, so a retry of the same logical call carries the same key and a genuinely new call doesn't.
Email and messaging. Usually no idempotency support. Record the intent transactionally before sending, keyed on the event, and check that record first. A crash between the record and the send means a missing message, which is recoverable; the reverse ordering means duplicate sends, which is not.
Aggregates. Incrementing a counter is not idempotent. Either store the contributing events and compute the aggregate as a query, or make the counter update conditional on the event not already being counted. The first is usually better — it lets you recompute after a bug.
Stateful external systems. "Create issue" duplicates. "Set issue state to closed" doesn't. Prefer declarative operations over imperative ones; set_to(x) is naturally idempotent in a way that increment and create are not.
Deduplication windows
Storing every event ID forever is expensive. Most pipelines keep a dedupe table with a TTL — typically a multiple of the longest retry window the upstream provider uses.
Two things to be honest about. First, a window shorter than the real retry horizon will let duplicates through, rarely and unreproducibly. Check the provider's documented maximum, then double it. Second, replays and backfills routinely reach outside any reasonable window, so the dedupe table can't be your only protection for those — backfills need their own reasoning.
A short checklist
Before a pipeline runs unattended:
- Every write has a key derived from the work, not the attempt.
- Uniqueness is enforced by a database constraint.
- Outbound calls use idempotency keys where supported, and a recorded-intent pattern where not.
- Aggregates are derived, or conditionally updated.
- The dedupe window exceeds the longest upstream retry horizon.
- Reprocessing the last 24 hours of input is a safe operation you have actually tested.
Point six is the real test, and it's the one to run deliberately rather than discover. If replaying a day of events would cause a second round of emails, refunds or pages, the pipeline isn't idempotent yet — it's just been lucky.
Turn one person's AI process into the company's.
Murmurator is invite-only. Already invited? Sign in and build something real.