The Production Answer
A production OpenTelemetry Java design should treat telemetry as a distributed data pipeline, not as a collection of logging libraries added to an application. The application should create traces, metrics, and logs through the OpenTelemetry API, while carefully selected SDKs, exporters, processors, and collectors handle configuration, batching, sampling, retries, and delivery. For Java services, the usual production path is an auto-instrumentation agent for rapid adoption, supplemented by explicit API instrumentation around business transactions and AI operations. This approach gives teams visibility across HTTP, JDBC, messaging, and supported frameworks without forcing them to redesign every call site immediately.
Also worth reading: How Should AI Architects Implement OpenTelemetry for Agent Observability in 2026? · How Should You Design an OpenTelemetry Trace Architecture for AI Agents? · How Do You Design a Reliable AI Agent Architecture for Production in 2026?
The central design decision is where control belongs. Instrumented code should remain independent of a particular observability vendor, while the deployment environment should decide which signals are exported, where they go, and how much they cost. In 2026, a reasonable default is parent-based trace sampling, approximately 10% for ordinary production traffic, with higher sampling for errors and explicitly important AI requests. That percentage is a starting point rather than a universal rule: a service processing 1 million requests per day at 10% produces about 100,000 traces, so cardinality, trace size, retention, and incident needs must be evaluated before the setting is finalized.
OpenTelemetry is appropriate when a Java organization has microservices, queues, external APIs, or databases whose behavior must be followed across process boundaries. It is less justified for a small monolith with excellent structured logs and no operational need for distributed traces. The technology is open source under the CNCF ecosystem, but ingestion, storage, query, retention, and staffing are not automatically free. A credible architecture therefore includes an explicit monthly telemetry budget and an owner responsible for instrumentation quality, rather than assuming that enabling every signal is harmless.
Separate the API, SDK, and Collector
OpenTelemetry Java has three architectural layers that should not be confused. The API is the instrumentation surface used by application code or agent-generated bytecode. The SDK implements trace creation, context propagation, metrics aggregation, log correlation, sampling, and batching. The OpenTelemetry Collector is a separate, vendor-neutral service that receives, processes, and exports telemetry. Keeping the API in application dependencies and the SDK plus exporter configuration in deployment packaging makes it easier to test an application without requiring a live observability backend.
A typical request carries trace context from an HTTP server through the Java service, possibly across Kafka or another messaging system, into a downstream service and database client. W3C Trace Context is the default propagation format for modern HTTP work, and OpenTelemetry instrumentation normally injects and extracts standard headers. Explicit instrumentation can attach attributes such as tenant_tier, model_name, tool_name, prompt_tokens, or order_id, but high-cardinality values such as complete prompts, raw SQL text, email addresses, or unique customer identifiers should not become metric labels. They may be appropriate in restricted logs or traces, but only after privacy, security, and storage rules are settled.
The Collector should run as a gateway or sidecar according to the hosting model. A gateway centralizes policy and reduces duplicated configuration, while a per-node agent or sidecar can improve locality and isolate workloads. Neither placement removes the need for back-pressure handling. Java exporters should use batching, bounded queues, and timeouts; otherwise, a slow destination can increase heap use or block application requests. A useful rule is to bound memory before maximizing telemetry: queue sizes, batch sizes, and exporter timeouts should have documented ceilings, with dropped telemetry represented as an operational condition rather than silently ignored.
Build a Practical Java Instrumentation Path
Begin with a representative service and run OpenTelemetry Java agent-based instrumentation in a non-production environment. Auto-instrumentation can cover common libraries and frameworks, including supported HTTP servers, servlet or Spring ecosystems, HTTP clients, JDBC, and messaging clients, but its exact behavior depends on framework and agent versions. Record which libraries were detected, because an apparently enabled trace with missing JDBC or messaging spans often means a compatibility gap rather than an application failure. Agent attachment through the JAVA_TOOL_OPTIONS mechanism, a container entry point, or a service-management wrapper should be standardized so operators do not need application-specific knowledge.
The next step is to add explicit spans only where auto-instrumentation cannot express business meaning. An order-processing path might create a span for pricing, inventory reservation, payment authorization, and fulfillment orchestration, recording the stable operation name as inventory.reserve. Do not create one span per method; that can produce thousands of low-value spans per request. Instead, aim for roughly 5 to 30 spans per typical production request, reserving more detail for exceptional or explicitly selected AI workflows. For an AI agent, useful boundaries commonly include planning, model invocation, tool execution, retrieval, validation, and final response delivery.
Metrics should describe service health and behavior rather than reproduce every trace attribute. Good Java service metrics include request rate, error rate, latency percentiles, JVM heap, garbage-collection pause time, thread count, and exporter queue pressure. Histograms or explicit bucket boundaries should be selected because OpenTelemetry SDKs may not preserve arbitrary backend histogram behavior. For example, HTTP latency buckets can be tested around 5, 10, 25, 50, 100, 250, 500, 1,000, 2,500, and 5,000 milliseconds. These values are not universal; they should match the service SLO and the queries operators expect to run during an incident.
Control Sampling, Cardinality, and Data Volume
Sampling is the main cost control in high-volume trace systems. Parent-based TraceIdRatioBased sampling keeps related spans in the same trace and is a sensible default for ordinary requests. Tail sampling in the Collector can make a better decision when the complete distributed trace is available, but it requires buffering, adds state, and increases latency in the telemetry pipeline. It should not be introduced merely because it is more advanced. A straightforward ratio sampler is easier to predict, while tail sampling is useful when errors are rare and retaining all failed traces has clear incident value.
Errors deserve special treatment, but “always keep errors” must be defined carefully. A status code or exception event is not always equivalent to a user-visible failure. Retried operations, validation responses, and expected business rejections can inflate retained traces if classified too broadly. Teams should agree on criteria such as unhandled exceptions, 5xx responses, or failed policy checks, then measure the retained volume. At 100 million daily requests, sampling 1% still retains 1 million traces; at 10%, it retains 10 million. Those numbers can dominate storage even when the SDK and Collector are inexpensive.
Metrics have a different risk: label cardinality. A metric labeled with raw URL, SQL statement, exception message, or request ID can create millions of time series and destabilize the backend. Keep metric labels bounded to a small set such as method, route template, status class, region, and outcome. Put request IDs, full exceptions, SQL, and diagnostic attributes in traces or logs. Apply attribute allowlists in the Collector, redact secrets before export, and test that authorization tokens, passwords, personal data, and prompt contents do not appear accidentally. For AI workloads, record token counts, model identifiers, latency, tool status, and cost estimates, while treating full prompt and completion text as sensitive payloads.
Collector Reliability and Backend Choices
The Collector should be configured as a production state-handling service rather than a simple forwarding binary. Use memory limiter and batch processors, define sensible scrape or receiver timeouts, and configure exporters with retry behavior that respects backend rate limits. A resilient pipeline usually places a queue or Collector tier between application agents and the vendor, allowing short backend outages to be absorbed without blocking request processing. Persistent queues can protect against host restarts, but they introduce disk requirements and must have capacity limits, encryption, and cleanup policies.
Backends differ materially in operational fit. A commercial observability platform generally provides managed ingestion, search, dashboards, trace exploration, and support, but charges according to spans, metric series, log volume, retention, or a bundled subscription. A vendor-neutral backend can be attractive for organizations that want direct control over storage and query infrastructure, but it transfers collection, upgrades, capacity planning, and incident response to the adopting team. OpenTelemetry itself does not determine these costs; its APIs and Collector allow the destination to change. The decisive questions are data residency, query latency, retention, supported Java and AI attributes, and the total number of systems engineers must operate.
| Feature | Managed observability backend | Self-hosted OpenTelemetry stack |
|---|---|---|
| Time to first useful telemetry | Often days to a few weeks | Often several weeks for production parity |
| Upfront software cost | Usually none or subscription-based | Infrastructure, storage, and engineering costs |
| Variable cost | Based on ingested spans, metrics, logs, retention, or requests | Based on compute, storage, network, and database growth |
| Operational burden | Lower; provider handles much of the platform | Higher; team owns upgrades, capacity, and incidents |
| Best fit | Teams needing fast deployment and vendor support | Regulated or specialized environments requiring control |
Common Java and OpenTelemetry Mistakes
One common mistake is enabling all three signals at maximum detail by default. Traces, metrics, and logs serve different purposes, and duplicating the same information increases cost without improving diagnosis. Another is instrumenting only inbound HTTP calls, leaving asynchronous work and database behavior invisible. In Java, request context can disappear when work is moved to an executor, scheduled task, or queue; propagation must therefore be tested across ExecutorService, thread-pool boundaries, and messaging headers. Auto-instrumentation helps in supported cases, but custom asynchronous pipelines still require deliberate context handling.
A second class of mistakes concerns naming and semantics. Span names should be stable and low-cardinality, such as POST /orders or inventory.reserve, rather than containing IDs or raw URLs. Metric units should follow conventions, and HTTP status should be represented consistently as a bounded attribute. Teams should also avoid mixing OpenTelemetry resource attributes with span attributes in a way that changes global labels unexpectedly. Resource-level fields such as service name, environment, region, and version are useful for filtering, but changing a resource attribute can create a new stream of metric series.
Finally, many projects treat the Collector as a black box. A Collector that restarts, loses queued data, or silently drops spans may leave the application healthy while destroying diagnostic evidence. Monitor Collector CPU, memory, queue length, refused spans, export failures, and receiver counts. Set alerts on sustained export failure rather than on every single dropped item. In AI applications, do not label a model name with a full version string that changes on every deployment unless the resulting cardinality is intentionally accepted. Stable semantic conventions and reviewed attribute schemas are more valuable than collecting every field available in an SDK.
AI Workloads Need Their Own Trace Model
AI services require additional care because a single user action may involve several model calls, retrieval operations, tool calls, and validation steps. Trace the orchestration path separately from vendor SDK internals, using spans for meaningful work such as agent.plan, model.generate, vector_search, tool.execute, and response.validate. Record model family, operation type, input and output token counts, latency, finish reason, and estimated cost when available. Avoid recording raw prompts or completions by default, especially in regulated or multi-tenant environments; they may contain credentials, personal data, source code, or confidential business information.
Token and cost fields should use documented units and clear names. A trace can show that a request used 2,400 input tokens and 380 output tokens, but that number is not enough to infer business value without model pricing and whether retries occurred. Correlate the AI span with the parent request, then expose aggregate token and cost metrics by bounded model, operation, and tenant tier. The architecture should distinguish application latency from provider latency, tool latency, and queue time. Otherwise, a slow response may be misdiagnosed as a Java CPU problem when the actual delay was a 20-second model or retrieval call.
OpenTelemetry’s vendor-neutral model is useful here because model providers and agent frameworks change faster than core Java services. Still, neutrality does not eliminate semantic design work. Define which fields are required for incident response, which are safe to retain, and which should be sampled less aggressively. A sensible policy might retain detailed AI traces for 1% of ordinary requests, 100% of tool failures, and 100% of requests above a defined token or latency threshold. These are policy examples, not universal thresholds; the correct values depend on traffic, incident frequency, and retention cost.
Rollout, Timing, and Cost Decisions
A staged rollout usually takes 4 to 12 weeks for a moderate Java estate: one week for inventory and data classification, one to two weeks for a pilot, two to four weeks for backend and Collector hardening, and several weeks for service-by-service adoption. The exact duration depends on deployment automation, framework compatibility, and whether logs and metrics already exist. Start with a service involved in a real incident path, not a toy application, and compare telemetry quality against known failures. The pilot should answer whether engineers can identify a slow dependency, failed message, and database wait without reading application source code.
Cost should be modeled before broad rollout. If a service produces 1.2 million spans per day, 10% sampling reduces that to 120,000 retained spans before vendor-specific enrichment or platform multipliers. A low-cost open-source SDK does not make the resulting platform free: storage, query compute, network transfer, retention, and human investigation remain. Commercial plans may be priced per ingested span, active series, log volume, host, or monthly platform usage, so obtain current pricing rather than quoting a generic per-million-span figure. A useful approval threshold is a defined telemetry budget per service, such as 5% of the platform’s monthly bill or a maximum monthly storage footprint.
Act now when teams have multiple services, unclear latency ownership, or an AI workflow whose failures cannot be explained from logs alone. Delay instrumentation when a monolith has stable dashboards, low operational complexity, and no need for cross-system context; a smaller logging and metrics investment may provide more value. OpenTelemetry Java is a strong production foundation, but its value depends on disciplined semantics, controlled sampling, redaction, and operational ownership. The right design is not the one exporting the most telemetry; it is the one allowing an engineer to answer specific production questions within minutes while keeping data and cost within agreed limits.