Digital Engineering
08 min read

OpenTelemetry provides vendor-neutral APIs, SDKs, semantic conventions, instrumentation libraries and a Collector for generating and transporting traces, metrics and logs. Distributed tracing follows a request across services by propagating trace context and recording spans that describe operations, timing, status and selected attributes.
A reliable implementation requires more than installing an agent. Teams must define service identity, propagate context across every supported boundary, control attribute cardinality, choose sampling deliberately, deploy resilient collectors, protect sensitive data and validate that traces answer real operational questions.
Project Supply designs observable digital platforms and production engineering systems: Project Supply Digital Engineering
What distributed tracing solves
In a distributed application, one customer request may cross an edge service, API gateway, authentication service, application services, queues, databases and third-party APIs. Local logs show fragments. A trace connects those fragments into a causal path so responders can see where time was spent, where errors originated and which dependencies were involved.
Tracing is most valuable for latency decomposition, failure propagation, dependency discovery, asynchronous workflows, release comparison and incident investigation. It is not a replacement for metrics or logs. Metrics reveal aggregate behaviour, logs capture detailed events, and traces explain individual execution paths.
Core OpenTelemetry concepts
Trace
A trace represents an end-to-end operation. It is identified by a trace ID shared across related spans. A trace may be complete, partial or unsampled depending on instrumentation and collection decisions.
Span
A span represents one timed operation such as an inbound HTTP request, database call, queue publish or internal computation. It includes a span ID, parent relationship, name, timestamps, status, attributes, events and links.
Context propagation
Propagation carries trace and baggage context across process and network boundaries. If an outbound client fails to inject context or an inbound service fails to extract it, the trace breaks into unrelated segments.
Resource
Resource attributes describe the entity producing telemetry, including service name, version, deployment environment, cloud or container information. Consistent resource identity is essential for ownership, filtering and comparison.
Instrumentation scope
Instrumentation scope identifies the library or component that generated telemetry. It helps teams distinguish application spans, framework instrumentation and custom instrumentation during upgrades and troubleshooting.
Collector
The OpenTelemetry Collector receives, processes and exports telemetry. It can centralise enrichment, filtering, batching, sampling, routing and credential management instead of embedding vendor-specific exporters in every service.
Reference architecture
Applications use OpenTelemetry SDKs and automatic or manual instrumentation. Telemetry is sent using OTLP to a local or gateway Collector. Collectors apply processors and export to one or more observability backends. The architecture should avoid making telemetry delivery a synchronous dependency of customer requests.
Small environments may send directly to a gateway Collector. Kubernetes deployments often use an agent pattern for node- or workload-local collection plus gateway collectors for central policy. Serverless and edge environments require platform-specific evaluation of cold starts, networking and export reliability.
Instrumentation strategy
Start with automatic instrumentation
Automatic instrumentation can capture common frameworks, HTTP clients, databases and messaging libraries quickly. Use it to establish coverage and discover gaps. Validate generated span names, attributes and status rather than assuming default output is operationally useful.
Add manual spans at business boundaries
Manual instrumentation should describe operations that matter to the business: price calculation, eligibility decision, inventory reservation, payment authorisation, fulfilment allocation or model inference. Instrument stable conceptual boundaries, not every function.
Use semantic conventions
Semantic conventions standardise attribute names and span meaning across technologies. Following them improves queries, dashboards and backend portability. Check the stability status of the relevant convention before building long-lived dependencies.
Avoid instrumentation duplication
Duplicate automatic and manual instrumentation can produce nested spans with little diagnostic value and increased cost. Document which library owns each boundary and suppress redundant instrumentation where supported.
Context propagation across boundaries
HTTP and RPC
Use supported propagators consistently across inbound and outbound clients. Validate proxies, gateways and service meshes because they can preserve, replace or drop headers. Never trust incoming trace identifiers without applying platform security controls.
Queues and event streams
Asynchronous systems may use parent-child relationships or span links depending on processing semantics. Preserve message context in supported metadata, but avoid creating misleading traces when one consumer processes batches or one event triggers many downstream operations.
Scheduled and batch jobs
Create a root span for the job or batch and child spans for meaningful stages. Record dataset or job identity using bounded attributes, not unbounded payload values. Link back to triggering events when causal context exists.
Third-party boundaries
External providers may not propagate compatible context. Represent the outbound call and its observed outcome without fabricating remote spans. Apply redaction and allow-list rules before sending headers or attributes across trust boundaries.
Span design
Span names should be stable, low-cardinality descriptions such as HTTP route templates or operation names. Do not put customer IDs, order numbers, full URLs or SQL values in span names. Dynamic values belong in governed attributes when genuinely needed.
Record status according to the operation’s result, not merely the transport response. Add span events for meaningful state transitions or exceptions. Keep attributes small and bounded; telemetry is not an alternate payload archive.
Sampling strategy
Head sampling
Head sampling decides near trace creation. It is simple and reduces volume early, but it cannot know whether a trace will later become slow or erroneous. Parent-based policies help maintain consistent decisions across services.
Tail sampling
Tail sampling decides after enough spans are collected to evaluate the trace. It can retain errors, high latency or selected business journeys while reducing normal traffic. It requires buffering, capacity planning and careful handling of incomplete traces.
Hybrid approach
Many platforms combine parent-based head sampling with Collector-level tail decisions or route different services through different policies. Preserve all traces for low-volume critical paths when justified, and sample high-volume routine traffic.
Sampling changes what evidence exists. Document policies, expose effective rates and avoid using sampled trace counts as exact business totals. Metrics should remain the source for aggregate service indicators.
Collector pipeline design
A Collector configuration contains receivers, processors, exporters, connectors and extensions assembled into pipelines. Use batching for efficient export, memory controls to protect the process, retry and queue behaviour appropriate to the backend, and health monitoring for the collector itself.
Separate tenant, environment or sensitivity boundaries when policy requires it. Restrict debug exporters in production. Store credentials through approved secret management and encrypt telemetry in transit. Treat Collector configuration as production code with review, versioning, testing and rollback.
Cardinality and cost control
High-cardinality attributes create expensive indexes and unreliable queries. Establish an attribute policy covering approved keys, data types, expected value ranges, retention and owners. Review route, database, messaging and business attributes before broad rollout.
Track spans per request, bytes per span, export volume, rejection, queue saturation, backend ingestion and query latency. Control cost by fixing noisy instrumentation, reducing redundant spans, filtering unsafe attributes and sampling intentionally—not by disabling telemetry during incidents.
Privacy and security
Telemetry can contain URLs, query parameters, database statements, user identifiers, IP addresses, headers, exception messages and business data. Use allow lists, redaction, hashing only where appropriate, and environment-specific policies. Do not collect secrets, tokens, payment data or unnecessary personal information.
Define access control, retention, residency, deletion and audit requirements with security and privacy owners. A vendor-neutral telemetry format does not remove responsibility for data handling in collectors or backends.
Need an observability architecture review covering performance, cost and sensitive-data controls? Contact Project Supply: Talk to Project Supply
Rollout plan
Phase 1: define questions and baselines
Choose important journeys and incident questions. Record current detection time, diagnostic effort, telemetry cost and service ownership. Define success before adding instrumentation.
Phase 2: instrument one vertical slice
Select a customer journey crossing several services. Instrument entry, internal calls, data stores, messaging and external dependencies. Validate context continuity and span meaning end to end.
Phase 3: deploy collectors and governance
Introduce resilient collection, attribute policy, environment tagging, security controls and backend routing. Load-test the pipeline and verify behaviour when the backend is slow or unavailable.
Phase 4: expand by service tier
Prioritise critical services and shared dependencies. Require service name, owner, version, environment and deployment metadata. Add instrumentation review to platform onboarding.
Phase 5: operationalise
Create trace-driven dashboards, exemplars where supported, investigation runbooks and release comparisons. Train incident responders to move between metrics, traces and logs.
Validation checklist
Send a known test request through every supported boundary and confirm one coherent trace. Verify parent relationships, timestamps, service identity, route naming, errors, retries, database spans, messaging links and deployment version. Test sampled and unsampled paths.
Restart services and collectors, simulate exporter failure, saturate queues and confirm application performance remains protected. Validate that forbidden attributes never reach the backend. Compare trace latency with independent measurements.
Operating model
The platform team should provide libraries, collector patterns, conventions and guardrails. Service teams own meaningful instrumentation and runbooks. Security and privacy teams approve data policy. FinOps or platform owners monitor consumption. A named observability council can resolve conventions without centralising every instrumentation decision.
Metrics for the telemetry system
Monitor SDK export failures, dropped spans, Collector accepted and refused spans, queue size, send failures, memory use, CPU, sampling rates, backend ingestion lag and query availability. The observability system needs its own service objectives.
Common implementation failures
Frequent failures include missing context on queues, inconsistent service names, dynamic span names, sensitive attributes, uncontrolled automatic instrumentation, duplicate spans, direct vendor exporters in every service, sampling without documentation, unmonitored collectors and dashboards that show traces but do not support decisions.
Another failure is declaring success from trace volume. Evaluate whether responders can answer specific questions faster and whether development teams use traces to improve releases.
Migration from vendor-specific tracing
Inventory existing agents, APIs, propagation formats, dashboards and alert dependencies. Introduce OpenTelemetry at a controlled boundary, compare trace completeness and semantics, then migrate services in stages. Keep backend-specific features explicit rather than pretending every capability is portable.
Use the Collector for dual export during validation when appropriate, but monitor duplicate ingestion and cost. Define an exit criterion before running parallel pipelines indefinitely.
Commercial decision framework
Evaluate OpenTelemetry as an operating model, not a free agent. Include engineering time, collector infrastructure, backend ingestion, retention, support, governance and training. Benefits include consistent instrumentation, backend choice, cross-service diagnosis and reusable platform controls.
Adopt when distributed complexity creates material diagnostic cost and the organisation can own instrumentation quality. Start smaller when the system is simple, service ownership is unclear or there is no capacity to operate the telemetry pipeline.
For OpenTelemetry implementation, Collector architecture and production observability support, contact Project Supply: Talk to Project Supply
FAQs
Web Personalisation
Framer is a design tool that allows you to design websites on a freeform canvas, and then publish them as websites with a single click.
UI and UX Design
Framer is a design tool that allows you to design websites on a freeform canvas, and then publish them as websites with a single click.
Search Engine Optimisation
Framer is a design tool that allows you to design websites on a freeform canvas, and then publish them as websites with a single click.
CRM and ERP Solutions
Framer is a design tool that allows you to design websites on a freeform canvas, and then publish them as websites with a single click.
Ecommerce
Framer is a design tool that allows you to design websites on a freeform canvas, and then publish them as websites with a single click.
Email Marketing
Framer is a design tool that allows you to design websites on a freeform canvas, and then publish them as websites with a single click.
Marketing Automation
Framer is a design tool that allows you to design websites on a freeform canvas, and then publish them as websites with a single click.
Chatbots and Conversational AI
Framer is a design tool that allows you to design websites on a freeform canvas, and then publish them as websites with a single click.
Chatbots and Conversational AI
Framer is a design tool that allows you to design websites on a freeform canvas, and then publish them as websites with a single click.
Related Blogs
We know your space
Explore our latest UI/UX Case Studies that showcase how our process-driven creativity transforms complex ideas into real, measurable business results, step by step.



