Monitoring with OpenTelemetry: where did the order go?
The API reports success but the order is missing from the ERP. Connect metrics, logs and traces across a queue to find the failed processing step.

Successful HTTP does not mean a completed order
In an illustrative online shop, the server accepts an order and returns HTTP 202. A worker must pass it to the ERP. The customer sees confirmation, but the warehouse cannot find the order. Availability monitoring stays green because the website and API respond. The team needs to observe the whole process: accepting a request does not establish that later processing achieved its result.
First name the observable states: accepted, waiting, forwarded, confirmed or failed. Identify the authoritative system for each state. If the ERP provides no confirmation, the integration must not claim completion based only on an outgoing HTTP call. This model also helps support explain to a customer what actually happened to the order.
Metrics show scale; a trace shows a particular journey
Metrics describe how many orders are waiting and how long processing takes. Logs record individual events and failure details. A trace connects periods of work across services. OpenTelemetry propagates context identifiers so a receiver can associate its activity with an earlier step. Installing an instrumentation library does not automatically define the order's business states.
Record acceptance, queue insertion, each worker attempt and the ERP response. An error log can include a trace ID and internal attempt identifier without storing the entire customer payload. Measuring only the web request excludes two hours spent waiting in a queue. Also measure the oldest pending job's age and elapsed time from acceptance to confirmation.
Context must cross the asynchronous boundary
Carry the permitted trace context in message metadata and extract it through supported propagation when processing starts. Choose the relationship between producer and consumer according to the messaging instrumentation. Batch processing and repeated attempts may call for span links. Reusing one span for every retry would blur separate attempts and their timing.
An event identifier and a trace identifier serve different purposes. The event remains the same across attempts, and its identity can prevent a duplicate business effect. A trace supports diagnostics. In this illustrative design, the database records event state while telemetry explains individual attempts. A failed trace export therefore cannot remove the record of completed orders.
Do not turn every order number into a metric label
Each combination of metric label values creates a separate time series. Order IDs, email addresses and full URLs with parameters produce a growing number of combinations. Prometheus therefore advises against unbounded label values. Processing duration is better grouped by finite categories such as integration type and attempt outcome.
Look up an individual order in protected business records and permitted logs. Specify retention, access and telemetry volume limits. Sampling can control cost, but it does not preserve every problem. Verify the chosen strategy for rare errors and overload behaviour. The business record cannot depend on whether sampling happens to retain a trace.
Propagating context does not grant permission
OpenTelemetry baggage can carry additional values across services. It is unsuitable for passwords, tokens or a complete customer profile. Headers may continue into an external call. A tenant_id received in baggage also does not prove the user's membership in that tenant. Authorisation must come from verified identity and server rules.
Use an explicit attribute allowlist and remove sensitive values before export. Apply appropriate controls in both the application and telemetry pipeline. Check automatically captured URLs, exception text and database parameters. An external system's error message may contain personal information even when the application never deliberately adds it to logs.
An alert must lead to an action
For this illustrative integration, the team could alert when a waiting order exceeds its permitted age. Choose the time limit from warehouse requirements, rather than a dashboard default. Include the affected integration, problem scope and a link to the response procedure. Repeated notifications without an owner or an actionable next step soon lose their effect.
- A simulated ERP outage triggers an alert while the web API still responds.
- One event can be followed through its queue and every processing attempt.
- Test tokens and personal information do not appear in exported telemetry.
- A telemetry outage does not affect order completion or its business record.
Sources and documentation
For implementation, consult the documentation for the version you use.
