API idempotency: retrying a request without a second order
A timeout does not confirm failure. Designing an idempotent operation, protecting concurrent requests in the database and testing duplicate writes.

A timeout leaves the outcome open
An online shop sends an order to an ERP. The ERP saves it, but the connection breaks before the response reaches the shop. The shop sees a timeout and retries the request. If the second attempt creates another order, someone must decide which one to cancel and check the stock reservation or subsequent invoice. This is an illustrative integration scenario; retrying the transfer alone does not resolve it.
A timeout means the caller did not receive an outcome within its time limit. It does not establish whether the remote system performed the operation. Your interface therefore needs a state for an outcome that is still unknown. The user can see a request awaiting confirmation while a technical process checks the result. Immediately declaring permanent failure and offering to create a new order could introduce another business intent.
An idempotent operation allows the same intent to be retried without another business effect within its agreed contract. You must define what counts as the same operation, how long the server remembers it and what happens under concurrency. Without those rules, a header containing a key is just additional text.
The HTTP method and the business effect
RFC 9110 defines idempotency in terms of a request’s intended effect. PUT, DELETE and safe methods have this property within their defined HTTP semantics. Their responses need not be identical: a second deletion may return a different status while the resource remains deleted. Recording each request in an operational log does not by itself violate this property either.
For POST /orders, however, you cannot assume that calling it twice only confirms the first call. A custom API can introduce idempotency for POST, but it must implement and document it. Changing the method’s name does not repair code that sends another invoice or starts another reservation on every attempt. During design, list all side effects, including those performed by another service.
The key represents intent, not form contents
In the proposed contract, the client generates a random identifier when one order begins and stores it with its local task. Every retry of that task uses the same key. When the customer actually orders another identical product, a new key is created even if the request body is identical. A hash of the JSON document alone could not distinguish these two intentions.
The server can scope the key by account and operation: tenant_id, operation, idempotency_key. Two accounts then cannot block each other with the same identifier. Authentication and authorisation must run on retries too. Knowing a key must not replace permission checks or grant access to an outcome belonging to another account.
Store a fingerprint of the normalised input with the key. Your normalisation rules must resolve defaults, field ordering and number representation unambiguously. If a client reuses a key with a different amount or address, this design returns a conflict and changes nothing. The client must not automatically “fix” that error with a new key; someone must first decide whether this is a new intent.
The local write and its result belong in one transaction
For an operation whose effects all remain in one database, reserving the key, creating the order and saving the response can share one transaction. The pseudocode below proposes a contract, not a complete implementation. It assumes a database unique constraint covering the key’s scope and an appropriate way to read the result after a conflict under the chosen isolation level.
Database enforcement matters when multiple servers are running. A “SELECT first, then INSERT” check without a unique constraint allows two processes to conclude simultaneously that a key is new. PostgreSQL can enforce uniqueness in the database. The application must then handle conflicts, waits for competing transactions and possible rollbacks; it must not convert them into another order creation.
This simplified model keeps the transaction open only for brief local work. Calling a slow external API in its middle would hold database locks while providing no way to undo a remote effect that has already happened. Such a process needs a separate state model, a stable identifier on the remote side too, and a way to investigate an unknown outcome.
authorize(account, create_order)
input = validate_and_normalize(request)
key = require_idempotency_key(request)
transaction:
claimed = claim_unique(account.id, 'create_order', key, hash(input))
if not claimed:
previous = load_committed_result(account.id, 'create_order', key)
require_same_payload(previous, hash(input))
return previous.status, previous.body
order = insert_order(input)
save_result(account.id, 'create_order', key, status=201, body={ order_id: order.id })
commit
return 201, { order_id: order.id }Concurrency and responses are part of the contract
When two attempts arrive together, one process obtains the right to perform the operation and the other waits or receives an agreed “processing” status. Waiting must have a limit. If the second attempt times out, it keeps the original key and checks the result later. An asynchronous operation may instead return a task identifier with a separate status endpoint.
The caller needs a usable response after losing the confirmation. In this illustrative model, a retry returns the original order’s identifier and the stored response status. Decide whether to return the original representation or a reference to the current resource; these behave differently after later changes. Avoid storing sensitive contents that are unnecessary for returning the outcome.
Retries need a time budget and a stopping condition
Set retry policies from the provider’s contract. Selected transient failures or rate limits may permit another attempt; permission errors or invalid input usually require intervention. The same HTTP status can mean different things in different APIs. A retry must preserve operation identity, respect a supported Retry-After and have limits on both total time and attempt count.
AWS describes exponential backoff with a random component as a way to spread retries. Also check whether an SDK, proxy or parent worker already retries requests. Multiple layers can multiply load. For this proposed integration, one owner would control attempts and hand the task over for investigation when its budget is exhausted, retaining a traceable outcome.
Key retention is finite. Stripe, for example, documents that keys can be removed after at least 24 hours; reusing a removed key can initiate a new request. That duration is not a universal rule. Your retention must cover delayed attempts, recovery after an outage and manual repairs. For important business records, a permanent external identifier independent of the temporary key may help.
Test the points where certainty is lost
A test that sends two identical requests consecutively covers only the easiest case. A useful suite deliberately interrupts the connection after commit, creates concurrency across instances and terminates a process before its transaction completes. For each test, verify the number of business records, the saved outcome and the next attempt’s behaviour. An HTTP success alone does not establish that no duplicate reservation occurred.
Operational records should include the operation identifier, attempt number and resulting state, with a correlation identifier where needed. Distinguish completed, pending and unresolved operations. The team can then repair an individual transfer without rerunning an entire export or looking for an order by an ambiguous customer name.
- After a lost response, the next attempt returns the same order identifier.
- Concurrent requests on two servers create only one business record.
- The same key with different normalised input produces the agreed conflict.
- Rolling back both the key reservation and the order permits a safe next attempt.
- A test after retention expires verifies the documented behaviour, including a delayed transfer.
- A retry under another account does not reveal someone else’s response, and every call checks permissions.
Sources and documentation
For implementation, consult the documentation for the version you use.
Put the topic into practice.
Related project: FaxCopy a.s.
