Production webhooks: signatures, duplicates and recovery
Separating event receipt from processing, controlling duplicate deliveries and repairing a transfer without repeating its business effects.

Receiving an event does not complete it
A payment system announces a successful payment. The endpoint returns HTTP 200 and only then adds a job to a queue. The process terminates between the response and that write. The provider has received confirmation, but the order is waiting for processing that never started. In the opposite sequence, completing a long process before replying may cause the provider to treat a slow response as failure and deliver the event again.
Separate receipt from business processing. In the proposed model, the endpoint verifies the request, durably saves the event to a database inbox and sends its agreed response only after commit succeeds. A separate worker performs the work later. Here, confirmation means “we have taken responsibility for further processing”, rather than “the order has already been fulfilled”.
The provider determines the response deadline. GitHub, for example, documents a 2XX response within ten seconds and recommends moving longer work into a queue. Another interface can have different deadlines and retry rules. Design for its documentation rather than a shared constant for every webhook.
Verify the signature before trusting the contents
Anyone can call a public URL. Before making a business change, verify the signature according to the provider’s protocol, limit input size and check that the event type is supported. A signature checks the message’s origin and integrity within that mechanism; it does not replace the rules governing which account may modify which order.
Stripe requires the original request body for signature verification. If middleware first parses JSON and serialises it again, it may change the bytes that were signed. Use a supported library and preserve the raw body for verification. Manage the secret outside code and distinguish test and production endpoints. Set timestamp tolerance and secret rotation according to the provider.
In this design, an invalidly signed message never enters the work inbox. Diagnostics may save a safe rejection reason and receipt time, but should not print secrets or complete payment details. Signature tests must cover a single changed byte and the actual middleware path; passing an isolated library test is insufficient.
The inbox needs a stable event identity
Distinguish an event from its delivery attempts. A business record can have multiple different events, and an event can have multiple attempts. An order identifier therefore is not a suitable universal deduplication key. In the proposed inbox, combine the integration owner, source and stable event identifier according to the provider’s contract.
CloudEvents defines identity through source and id. GitHub retains X-GitHub-Delivery when redelivery is requested. These examples explain why you must check the scope and stability of the actual identifier. If the source supplies no identity, receipt time alone cannot reliably replace it; the design needs rules grounded in that source’s business model.
The example table stores the payload, state and attempt count. Its primary key protects identity across multiple processes. The endpoint inserts a valid event atomically; on a conflict it acknowledges the existing event without creating another task. If an identical ID carries different contents, it records the discrepancy and follows provider-specific rules. Agree separate retention policies for the payload and identity record.
CREATE TABLE webhook_inbox (
tenant_id bigint NOT NULL,
source text NOT NULL,
event_id text NOT NULL,
payload jsonb NOT NULL,
received_at timestamptz NOT NULL DEFAULT now(),
state text NOT NULL DEFAULT 'pending'
CHECK (state IN ('pending', 'processing', 'done', 'failed')),
attempts integer NOT NULL DEFAULT 0 CHECK (attempts >= 0),
next_attempt_at timestamptz NOT NULL DEFAULT now(),
PRIMARY KEY (tenant_id, source, event_id)
);The worker must recover after a crash
The simplest worker for short database effects locks the selected record, performs the business change and marks the event complete in one transaction. If the process crashes before commit, the transaction rolls back and the event remains available. If commit succeeds, another worker sees the completed state. This protects a specific local change; it does not establish a general “exactly once” delivery guarantee.
Longer work can use a reservation with an expiry time, or lease. Its implementation needs an owner identity and rules governing who can confirm an outcome after the reservation expires. Merely setting processing can leave an event stuck forever after a crash. The inbox above is a schema foundation; a lease-based implementation must add these fields and concurrency protection.
The worker schedules transient failures for later attempts with a delay. Formatting errors or missing mappings go to investigation after an agreed limit. The operator needs access to the original event, a safe failure reason and any steps already performed. Correcting a mapping and deliberately replaying one event is often more useful than manually replaying a whole day’s traffic.
External effects cross the database transaction boundary
If a worker sends a message to an ERP and then marks the inbox complete, it can crash between those steps. A retry sends the message again. Reversing the order can lose the message. A database transaction cannot atomically confirm both its own commit and an external HTTP call without a separate distributed protocol.
For this proposed process, the local business change and an outbox entry can share one transaction. Another process sends the entry to its destination and retries under the same operation identifier. AWS notes that a transactional outbox can still publish duplicate messages; the consumer still needs idempotent processing. The outbox moves the responsibility boundary rather than removing it.
If the destination supports neither idempotency nor querying an outcome by a stable identifier, a timeout can leave the outcome unresolved. Introduce checking in the target system, approval or another business procedure. Automatically retrying forever could create further reservations or documents. This limitation should be understood before promising complete automation.
Delivery order does not determine order state
Stripe does not guarantee event delivery order. Check other providers independently. Imagine receiving “order fulfilled” before an older “order created” event. Blindly replacing state with the last delivered payload can return a completed order to the beginning of its process.
The solution depends on the available contract. If the source supplies a usable record version, compare it according to its rules. If it offers a current-state API, an event can trigger a fresh read. If you have only individual facts, design permitted state transitions and a way to handle a missing predecessor event. Neither receipt time nor an ordinary timestamp automatically establishes reliable global ordering.
Handle two different events concerning the same business effect separately too. Deduplicating event_id catches repeated delivery but may not catch distinct events that trigger the same fulfilment. A business rule may need another unique fulfilment identifier or a conditional database state transition. Derive this rule from event meaning rather than resemblance between JSON documents.
Operations need traceable exceptions
Monitor the age of the oldest pending event, backlog size and recurring failures by source. The number of HTTP 200 responses alone does not tell you whether workers are processing orders. Retain the necessary replay identity after sensitive payload contents have been removed. Access to details and manual replay must match operator permissions.
A test suite should combine receipt failures, worker crashes and problematic business transitions. Define the expected outcome of each scenario in advance. When a response is lost after the inbox commit, the next delivery should confirm the same record; when the connection fails before storage, the endpoint must not pretend it has accepted responsibility.
- An invalid signature or modified payload never triggers a business change.
- Concurrent deliveries with the same event_id create one inbox record.
- A crash before the local change commits permits another attempt without a partial outcome.
- Reverse delivery order does not move an order into an invalid older state.
- A crash after remote success exercises destination idempotency or the agreed manual check.
- Replay after a repair preserves event identity and records the operator’s decision.
Sources and documentation
For implementation, consult the documentation for the version you use.
Put the topic into practice.
Related project: FaxCopy a.s.
