Transactional outbox: when the database commits but the message fails
An order can exist while the warehouse never receives its event. Outbox design, message identity, acknowledgement boundaries and recovery after worker failure.

A gap exists between committing and sending
An illustrative ordering system saves an order and then tells the warehouse to reserve stock. If the process crashes after commit, the order remains but the message never leaves. Reversing the sequence creates a different problem: the warehouse may receive the request even though the subsequent order insert fails. The two operations need a clearly defined shared boundary.
The transactional outbox pattern stores the business record and the intent to publish an event in one local database transaction. A separate worker sends committed events. AWS documents this principle; the implementation depends on the database and transport. A local transaction does not undo a remote side effect and does not automatically give the entire integration exactly-once behaviour.
Assign event identity when the event is created
In this proposed model, an event carries its own event_id, source system, tenant, business object and object-change version. These values stay stable across attempts. The recipient can distinguish another delivery from a new event. CloudEvents uses the source and id pair to identify an event within its defined scope; the order's business identifier can remain a separate field.
The payload should contain what the recipient needs and a format version. Publishing an entire internal database row would unnecessarily couple an external integration to a private schema. When choosing between event data and a later API read, consider whether the recipient needs the object's state when the event occurred or its latest state.
An outbox insert must fail with the order
The following pseudocode illustrates only the atomic local write. Both tables are assumed to use the same transactional database. If the event insert fails, the transaction must not commit the order alone. The API also needs its own protection against repeated business intent; an outbox does not prevent a client from creating a second order through a new request.
The worker starts from committed rows. An outbox record can contain a status, attempt count, next-attempt time and a short processing lease. The lease needs an owner and an expiry so another worker can recover the job after a crash. Keeping a database transaction open during a network call would extend lock duration without solving a lost broker acknowledgement.
event_id = new_event_id()
transaction:
order = insert_order(validated_input)
insert_outbox(
event_id=event_id,
source='orders',
aggregate_id=order.id,
aggregate_version=1,
type='order.created.v1',
payload={ order_id: order.id }
)
commit
return order.idBroker acknowledgement differs from recipient completion
The worker sends an event, the broker accepts it and the connection breaks before the response arrives. The worker cannot determine the outcome and sends the same event after recovery. A crash after a successful response but before marking the outbox record as sent can also cause duplicate delivery. The design therefore allows another attempt with the same identity.
The recipient can record the processed identity alongside a local stock reservation in one transaction. Remote side effects need an additional contract. A sent status in the outbox represents the agreed transport acknowledgement, not an automatically completed order. If the business process requires confirmation that stock was reserved, it needs a separate business state and a result event from the warehouse.
Define ordering for the relevant business object
Changing the order of order.created and order.cancelled events can damage the result. Creation timestamps alone are insufficient when writes or sends run concurrently. This illustrative design assigns an increasing version to each order. Under its contract, the recipient processes the expected version, identifies an old event or pauses the object when a version is missing.
Decide whether an error affecting one order should block other orders too. Ordering within one object is often sufficient, although the business process may require a wider relationship. Moving a problematic event to a separate queue still leaves the question of what happens to its successors. Automatically unblocking the entire queue can violate the agreed ordering.
Operations need event age and a way to investigate
A queue's row count says little without context. Track the oldest pending event, errors by recipient and time to business acknowledgement. A short large batch may be healthy while a single old reservation needs attention. Retry needs a bounded time budget, delays with jitter and a distinction between temporary failure and an invalid payload.
Manual redelivery should preserve event identity and record the reason for intervention. Retention rules must cover outage recovery and the required audit trail. The following checks exercise points at which the outcome becomes uncertain. Inspect the order, outbox, transport and warehouse effect; a successful worker log is only one part of the evidence.
- A crash before commit saves neither the order nor the event.
- A crash after commit leaves the event available to another worker.
- A lost acknowledgement produces another attempt with the same event_id.
- Delivering one event twice cannot create a second local reservation.
- A delayed cancellation and a missing object version produce the agreed business outcome.
Sources and documentation
For implementation, consult the documentation for the version you use.
