Skip to content
fullstackhero

Concept

Idempotency

Idempotency-Key header support with distributed-cache-backed replay protection across instances.

views 0 Last updated

Network is unreliable. A POST /orders request might submit, time out before the response arrives, and get retried. Without idempotency, the retry creates a second order. With it, the retry returns the same response as the first.

The kit ships an Idempotency-Key header convention. Callers send a fresh key with every request that mutates state; the kit caches the response (status + body) keyed on (tenant, key) for a configurable window. Duplicate requests within the window return the cached response without re-executing the handler.

Opt an endpoint in

The Web block ships a .WithIdempotency() extension on RouteHandlerBuilder:

endpoints.MapPost("/orders", handler)
.RequirePermission(perm)
.WithIdempotency();

An endpoint whose response goes stale on its own schedule passes its own window:

// A presigned URL is good for FilesOptions.UploadUrlTtlMinutes, so the replay must not out-live it.
var uploadUrlTtl = TimeSpan.FromMinutes(
endpoints.ServiceProvider.GetRequiredService<IOptions<FilesOptions>>().Value.UploadUrlTtlMinutes);
endpoints.MapPost("/upload-url", handler)
.RequirePermission(perm)
.WithIdempotency(uploadUrlTtl);

Without it the entry keeps DefaultTtl, and a retry an hour later would get a 200 carrying a URL that expired forty-five minutes earlier, with no way out except inventing a new key. Read the value at map time from the same option the handler mints the response with, so the two cannot drift. A non-positive TTL throws at startup: it would expire every entry on write, leaving the endpoint advertising an idempotency it no longer has.

That’s the whole opt-in. The kit’s IdempotencyEndpointFilter (an IEndpointFilter) wraps the handler:

  1. Before: read the Idempotency-Key header. If present, build the cache key from the tenant (the resolved tenant context, the caller’s tenant claim as a fallback, "global" when neither is available), the caller (the user id; anon when unauthenticated, which is why idempotency does not belong on an anonymous endpoint, see below), the operation (HTTP method + route pattern + resolved route values) and the key itself, then probe the cache.
  2. If cached: write the cached status, body and replayable headers, set an Idempotency-Replayed: true response header, and skip the handler entirely.
  3. If not: reserve the key so a concurrent duplicate can’t run the handler too, invoke the handler, capture what it actually put on the wire, store that, then release the reservation.

If no header is sent, the filter is a no-op - the endpoint behaves like an ordinary one. Keys longer than MaxKeyLength are rejected with a 400 before the handler runs.

Probe and store both go through IDistributedCache, on the same key and the same serializer. That symmetry is the whole game: a store that keys its entries under a different scheme than the probe reads makes replay silently never engage, with no error anywhere.

What gets replayed

The handler’s IResult is executed into an in-memory buffer first, so what’s captured is the real wire shape - the plain DTO body and the status the result actually sets, not the Ok<T> / Created<T> wrapper and not the status the response happens to carry before the result runs. Alongside them the filter captures an allow-list of headers that carry meaning for the caller: Location and ETag. A replayed 201 therefore still points at the created resource. Transport and host-owned headers (Content-Length, Transfer-Encoding, Date, Server) are deliberately not replayed - a stale value there corrupts the response.

Only successes are stored

A response is stored only when its status is 2xx. A 409, 429 or a 500-shaped result is not a record of a committed side effect, and storing it would lock the caller out of that key for the whole TTL after a transient downstream failure. A retry with the same key after a failure runs the handler again, which is what you want.

The store outlives the request

The response is written to the cache before the body goes to the client, on a cancellation token that can’t be cancelled. Client-times-out-then-retries is the single commonest way a duplicate request is produced; if the store were tied to the client’s connection, that exact retry would find nothing cached and run the handler a second time.

For the same reason the handler itself runs with the client’s abort token detached. Left attached, a disconnect right after the side effect committed cancels whatever the handler awaits next - an EF read, an outbox write, a Mediator behaviour - and the exception leaves the filter with nothing to store, so the retry re-executes. The trade-off is explicit: on an idempotent endpoint, a client hanging up no longer aborts the handler. Keep those handlers short, and don’t put .WithIdempotency() on a streaming or large-file endpoint - the response is buffered in memory to be captured, with no size ceiling.

Storing is best-effort, and so is reading: if the cache is down, the probe, the reservation and the store each log a warning and the request proceeds. Idempotency degrades to a convenience instead of 500ing requests - including the probe, which runs on every keyed request and would otherwise take every idempotent endpoint down for exactly the clients that send a key.

A handler that writes to HttpContext.Response itself is left alone: the response has already started, so nothing can be captured and nothing is stored. Return an IResult from an idempotent endpoint.

Concurrent duplicates

Two requests carrying the same key at the same time are serialized by an atomic in-flight reservation - Redis SET NX when an IConnectionMultiplexer is registered, an in-process set otherwise (single-instance hosts; a multi-instance host in this stack already runs Redis for the shared Data Protection key ring). The duplicate that loses the race re-probes once - the original may have finished in the meantime, in which case it replays - and otherwise gets 409 Conflict with Retry-After: 1: a request with this key is still being processed, try again in a second.

The reservation fails open. A Redis blip on reserve or release logs a warning and lets the request through rather than failing it; the stored response still dedupes later retries. A request that failed open holds no reservation, so its release deletes nothing - releasing is a compare-and-delete against the token the reservation was taken with, which keeps a request from freeing a lock another one owns.

Both branches expire on ReservationTtl: the in-process one hands the key over once the holder has outlived the TTL, mirroring what Redis does on its own, so a handler that never returns cannot strand the key until the process restarts.

Configuration

{
"IdempotencyOptions": {
"HeaderName": "Idempotency-Key", // default
"DefaultTtl": "1.00:00:00", // 24 hours (default) - how long a stored response replays
"ReservationTtl": "00:01:00", // 1 minute (default) - how long the in-flight lock survives
"MaxKeyLength": 128 // default
}
}

ReservationTtl is deliberately decoupled from DefaultTtl. It only has to outlast the handler’s execution: if the process is killed between reserving the key and releasing it, the lock frees itself in about a minute instead of stranding the key for the full response TTL, during which every retry would 409. Raise it if you have a handler slower than the default - if it lapses mid-request, a concurrent duplicate can slip past the lock.

All four are validated at startup: a non-positive TTL, a ReservationTtl above DefaultTtl, an empty header name or a non-positive key length fails the host rather than degrading silently (a zero DefaultTtl used to throw inside the best-effort store, which logs a warning and carries on - so nothing was ever cached and replay never engaged).

Idempotency is on by default in FshPlatformOptions (EnableIdempotency = true); AddHeroPlatform binds the options. The replay store is IDistributedCache, so a Valkey backing means cached responses survive across instances. Without Valkey you get per-instance idempotency, which is OK for dev but not for multi-instance production - and note that cross-instance dedup depends on the cache being on Valkey, not on the reservation: a host that points only quota at Redis gets a shared lock over a per-process response store, which still lets each instance run the handler once.

What clients send

POST /api/v1/orders HTTP/1.1
Authorization: Bearer <token>
tenant: acme
Idempotency-Key: 7f3d9a2c-1b8e-4f5a-9c0d-2e8f6b1a3d4c
Content-Type: application/json
{ "items": [...] }

The key should be a fresh UUID per logical request:

  • A retry of the same request uses the same key - gets the cached response (and the Idempotency-Replayed: true header, so clients can tell).
  • A different request (even from the same client) uses a fresh key - gets a fresh response.

Most HTTP client libraries can generate keys automatically; for HttpClientFactory consumers, a DelegatingHandler is a clean place to do it.

What it doesn’t do

  • It doesn’t dedupe content. If the client sends two requests with different keys and identical bodies, both create resources. Idempotency keys are per-request-attempt identifiers, not content hashes. (Use a content hash + lookup if you want content dedupe - but it’s a different feature.)
  • It doesn’t span the request boundary forever. After the entry’s TTL expires, a replayed request runs again. DefaultTtl is 24 h; set it to whatever is realistic for your client retry strategy, and give an individual endpoint a shorter window with WithIdempotency(ttl) when what it returns goes stale sooner than that.
  • It doesn’t catch partial failures. If the handler runs, mutates state, then throws before the response is captured, the cache won’t have an entry. The retry runs again. Combine idempotency with idempotent domain operations for true safety. (A client that hangs up after the handler committed is covered - the store happens before the body is written to the client.)
  • It doesn’t replay failures. Only 2xx responses are stored, so a retry after an error re-runs the handler rather than replaying the error.

Domain-level idempotency

The kit’s domain already has several naturally-idempotent operations:

  • Notification.MarkRead() - ReadAtUtc ??= now; second call is a no-op.
  • Ticket.Reopen() - guarded against illegal source states; second call from a valid state is harmless.
  • WebhookSubscription deactivation - sets IsActive = false; idempotent.
  • Find-or-create DM channels - DirectKey uniqueness makes “find or create” race-safe.

Layer Idempotency-Key on top of these for full retry safety across both transport and domain failure modes.

Gotchas

  • The cache key includes tenant, so two tenants using the same idempotency key get separate responses. This is correct; tenants are independent.
  • The tenant on the key is the one the request is scoped to, not the one in the caller’s token. That matters for a root operator using the tenant header to act on a specific tenant: the entry follows the target tenant, which is also the one the handler’s writes land in. An unresolved tenant header - one naming a tenant that does not exist - is never used to build the key; those requests fall back to the claim, then to "global".
  • The cache key includes the operation (method + route pattern + resolved route values), so one key reused against two different idempotent endpoints - or against two different resources on the same endpoint, PUT /tickets/1 then PUT /tickets/2 - no longer replays the first response on the second. Still generate a fresh key per logical request; the scoping is a guard rail, not a licence to share keys.
  • The cache key includes the caller. Two users of the same tenant who pick the same key value on the same endpoint get separate entries, instead of one receiving the other’s response body while their own request is silently suppressed.
  • The cache key does NOT include the request body. The same key sent to the same endpoint with a different payload replays the first response rather than rejecting the mismatch. If you need that check, hash the body yourself.
  • Do not put .WithIdempotency() on an anonymous endpoint. There is no user to scope by, so every unauthenticated caller resolves to the same anon caller: two people sending the same low-entropy key against the same endpoint and tenant would share one entry, and the second would replay the first’s response while their own request was silently suppressed. WithIdempotency() marks the endpoint with IdempotentEndpointMetadata, and an integration test walks the endpoint map to fail the build if an AllowAnonymous() endpoint carries it. Self-registration, the one endpoint this applied to, no longer does: a genuine retry there is already safe, because the unique-email constraint rejects the duplicate.
  • Cache miss after restart is normal. If Valkey isn’t configured, the in-memory store starts empty after a restart - replayed requests run again. Always use Valkey in production.
  • A duplicate still in flight gets 409, not the response. The original hasn’t produced one yet. Clients that retry aggressively should treat 409 on an idempotent endpoint as “retry shortly” - the response carries Retry-After: 1 - not as a business conflict.
  • Idempotency-Replayed is not exposed to cross-origin JavaScript. The CORS policy doesn’t list it in Access-Control-Expose-Headers, so a browser client cannot read it. Server-to-server callers see it normally.