# How Should You Design an OpenTelemetry Trace Architecture for AI Agents?

Savannah Jenkins · September 28, 2026

> Start with the request, not the model call The best OpenTelemetry trace architecture for an AI agent begins with a stable execution boundary around the...

## Start with the request, not the model call

The best OpenTelemetry trace architecture for an AI agent begins with a stable execution boundary around the user’s request, not around an individual prompt sent to a foundation model. A request may trigger planning, multiple model calls, retrieval, tool use, code execution, memory operations, handoffs, retries, and human approval. If the trace starts only when the first model invocation occurs, important orchestration work disappears from the system’s operational history. Create a root span such as agent.execute for the complete request, or use a parent span belonging to the API gateway or job scheduler when the agent runs asynchronously. The root span should carry the request or conversation correlation identifier, tenant or environment identifiers, agent version, and a safe classification of the request. Child spans should represent meaningful units of work rather than every internal function call.

**Also worth reading:** [How Should an Enterprise Design MLOps Governance Architecture in 2026?](https://agustin-otegui.com/knowledge/how_should_an_enterprise_design_mlops_governance_architecture_in_2026.php) · [What Is Secure Agent Architecture and How Should AI Teams Design It in 2026?](https://agustin-otegui.com/knowledge/what_is_secure_agent_architecture_and_how_should_ai_teams_design_it_in_2026.php) · [How Do You Design a Truly Scalable AI Infrastructure Architecture in 2026?](https://agustin-otegui.com/knowledge/how_do_you_design_a_truly_scalable_ai_infrastructure_architecture_in_2026.php)

This distinction matters because agents are not ordinary request-response services. A single user action can produce a graph of dependent operations over several seconds or minutes. A model may select a tool, wait for another agent, revise its plan, and call a vector database before producing a final answer. The trace must make those dependencies visible without implying that every span has the same importance. The root execution span is therefore both a correlation anchor and a place to record the overall outcome, duration, token usage, cost estimate, and error status. It should not contain the complete prompt, chain-of-thought, or raw tool output merely because those values are available in application memory.

OpenTelemetry supplies the instrumentation and propagation model, but it does not determine the correct domain vocabulary. The architecture team must decide which operations deserve spans, which attributes are safe and useful, and how traces relate to evaluations, logs, and business metrics. Treating the agent as a conventional distributed system first, then adding a small number of agent-specific conventions, usually produces a more maintainable design than inventing a completely separate tracing model.

## Define a span hierarchy that reflects agent work

A practical hierarchy separates orchestration, reasoning, model interaction, tools, data access, and external work. The root request span can contain a planning span, one or more model-generation spans, retrieval spans, tool-execution spans, and a final response span. Delegated tasks should receive child spans, even when they run in another process or service. Human approval should normally appear as its own span or event-linked operation so that queue time is not confused with model or tool latency. Retries belong under the operation that failed, with retry count and final outcome represented explicitly, rather than creating disconnected traces for each attempt.

Use span names that are stable, low-cardinality, and meaningful to someone debugging production. Names such as chat.completions, vector.search, browser.navigate, or agent.delegate are more useful than names containing the user’s question, document title, tool result, or model-generated plan. A span name should identify the kind of operation; attributes should identify the instance, subject, and result. For model spans, record the provider, model identifier, request mode, token counts, finish reason, latency, and cost fields supported by the platform. For tools, record the tool name, status, duration, and safe result metadata. Avoid recording free-form reasoning text as a span name or a required attribute.

The hierarchy should also distinguish work performed synchronously from work transferred to a queue. When an agent submits a job to a worker, propagate the OpenTelemetry trace context through the message headers and create a span representing the enqueue operation. The worker should continue the same trace or create a linked span if the messaging system cannot preserve the original context cleanly. This is particularly important for systems built with multiple frameworks, model gateways, and asynchronous evaluators. A trace that stops at the framework boundary gives an incomplete view of user-perceived latency and prevents teams from finding the actual source of a delay.

## Use OpenTelemetry conventions without treating them as a complete agent standard

OpenTelemetry provides a vendor-neutral foundation through traces, metrics, logs, context propagation, and increasingly defined semantic conventions for generative-AI systems. Those conventions can help standardize attributes such as model provider, model name, token usage, and operation type. They should be used as an interoperability baseline, not as a substitute for a domain model. The conventions are still evolving, and support differs across SDKs, collector pipelines, storage backends, and agent frameworks. Pin the versions of your instrumentation libraries, document the conventions you implement, and test whether your backend preserves the attributes and links you expect.

A sound design has three layers. The first is the generic service layer, using normal OpenTelemetry conventions for HTTP, database, messaging, and remote procedure calls. The second is an AI layer, using agreed attributes for models, prompts, retrievers, agents, and tool calls. The third is an application layer, using your own fields for business outcomes such as booking completed, claim approved, or answer accepted. Keep these layers distinguishable. For example, a vector search can be represented with database or client spans for infrastructure debugging and a retrieval span for semantic interpretation, but the two should not create conflicting names or duplicate expensive payloads.

Instrumentation should happen at framework boundaries and at explicit workflow transitions. Automatic instrumentation can cover HTTP, database, and messaging operations, while agent-specific instrumentation should be added around planning, model invocation, tool selection, and delegation. Manual spans are valuable when an agent framework hides the timing or dependency information you need. However, manual instrumentation should be bounded. A span should correspond to an operation that can be independently understood, measured, or failed. Creating a span for every prompt token, internal state update, or helper function increases volume while making the trace harder to read.

## Control sensitive data before the collector receives it

The most important production requirement is to prevent prompts, secrets, credentials, personal information, and proprietary documents from entering telemetry unintentionally. AI agents often pass highly sensitive material through prompts, tool arguments, retrieved passages, and model responses. If the application attaches an entire message payload to a span, the observability backend becomes a secondary data store with different retention, access-control, and residency properties. Redaction must therefore occur before data is exported, ideally in instrumentation code or at the OpenTelemetry SDK processor boundary. Do not rely on a backend search feature to hide sensitive values after ingestion.

Design attributes around purpose rather than convenience. Instead of storing the full user question, store a request identifier, request category, language, tenant, and an approved hash or fingerprint where correlation is necessary. Instead of recording a retrieved document, record the collection, document count, ranking method, top-k value, and document identifiers that are appropriate for your security model. Tool arguments should expose a typed schema and redact fields such as passwords, access tokens, payment details, and personal identifiers. Model responses may contain secrets even when the input does not, so response capture requires an equally strict policy.

A useful data classification can distinguish public metadata, internal operational metadata, customer content, regulated content, and prohibited secrets. Each class should have an explicit capture and retention rule. For example, token counts, model names, latency, and error codes may be retained for 30 days, while full prompts might be disabled by default or retained for seven days in a restricted evaluation environment. These periods are examples, not universal rules; legal, contractual, and regional requirements must determine the actual values. Apply sampling, filtering, and attribute limits in the SDK and collector, then verify the result with automated tests that search exported spans for known secret patterns.

## Propagate context across every boundary

Trace context must cross HTTP calls, model gateways, databases, vector stores, message brokers, browser sessions, and worker processes. Use W3C Trace Context or the propagation mechanism supported by your environment, and ensure that the same trace and span identifiers are available on both sides of every asynchronous boundary. Context propagation is not merely a tracing feature. It allows a backend to reconstruct the full execution path, but it also gives every downstream service the opportunity to attach its own safe telemetry. The parent-child relationship should remain correct even when the agent makes parallel tool calls.

Parallelism requires particular care. When several tools begin at the same time, each should become a child of the relevant planning or execution span. If a tool result triggers another tool, that second operation should be linked or nested beneath the workflow operation that consumed the result. OpenTelemetry span links are useful when one operation is caused by another without a simple parent-child relationship, such as an evaluator inspecting a completed trace or a background job reacting to an agent response. Links should supplement, not replace, proper context propagation.

External vendors often do not return an entire trace context, or they may generate their own trace for a request made through a gateway. Preserve your internal parent span and record the vendor request identifier, status, and duration. If the vendor supports trace headers, propagate them. If it does not, do not fabricate a continuous distributed trace that implies stronger correlation than actually exists. Store a separate vendor trace reference and use links or searchable attributes. This approach gives operators a truthful picture while still allowing them to compare internal behavior with provider-side diagnostics.

## Compare traces, metrics, logs, and evaluations

Traces answer where time was spent and how work flowed. They are best for investigating a specific request, identifying a slow retrieval operation, or reconstructing a multi-agent handoff. Metrics answer aggregate questions such as whether p95 latency increased, how many agents failed, or how many model calls were made per successful task. Logs are appropriate for detailed diagnostic events, policy decisions, and application messages. Evaluations assess whether an answer or action was useful, safe, and consistent with a task-specific standard. A trace architecture that tries to answer all four questions with one span stream usually becomes expensive and difficult to operate.

Measure agent quality and operational performance separately, then connect them. A model call can be fast and cheap but produce an incorrect answer; an apparently successful agent run can also consume excessive tokens or trigger an unauthorized tool. Track operational metrics such as request duration, model latency, tool latency, retry rate, queue time, token usage, estimated cost, and failure category. Track outcome metrics such as task completion, citation validity, policy violation, human correction, and evaluator score. Use trace exemplars or IDs to move from an aggregate regression to the individual executions that caused it.

The following table summarizes the primary purpose of each telemetry type and the AI-agent information it should normally contain.

| Telemetry type | Primary question | Useful AI-agent information |
| --- | --- | --- |
| Traces | Where did this request spend time and how did work flow? | Agent execution, planning, model calls, retrieval, tools, handoffs, approvals, and errors |
| Metrics | Is the system getting faster, cheaper, or less reliable? | Latency, token use, cost, failure rate, tool success, queue time, and throughput |
| Logs | What diagnostic detail explains an event? | Policy decisions, sanitized error context, configuration version, and event messages |
| Evaluations | Was the result correct, useful, and safe? | Task completion, factual support, policy compliance, relevance, and human or automated scores |

A practical baseline is to sample successful traces at a low rate, retain a larger sample of failures, slow requests, high-cost requests, and requests selected by evaluators. Be careful with statistical bias: if only failures are retained, teams cannot establish a reliable view of normal behavior. Conversely, capturing every full interaction may be unnecessary and unsafe. Sampling policies should be documented, tested, and adjusted for regulatory and debugging requirements.

## Plan for volume, cost, and backend behavior

Agent traces can generate substantially more data than conventional service traces because a single request may contain many model, retrieval, and tool spans. Instrumentation libraries reduce collection overhead, but payload size and cardinality remain design decisions. Avoid serializing the entire conversation into every span, embedding large tool results, or attaching dynamic keys such as arbitrary user IDs to metric labels. High-cardinality values are appropriate for trace attributes when necessary, but they are generally unsuitable as metric dimensions because they can overwhelm the metrics backend.

Use the OpenTelemetry Collector as a controlled pipeline rather than a pass-through pipe. Apply processors for redaction, attribute transformation, sampling, batching, and routing. Send operational traces to a general observability backend and sensitive evaluation traces to a restricted environment when the data classification requires separation. Configure tail-based sampling carefully: a policy that retains only errors can hide slow or expensive successful runs, while indiscriminate retention can create unexpected storage costs. Establish volume estimates before production rollout. A reasonable planning exercise is to measure spans per request, average encoded span size, peak requests per second, and retention duration, then multiply those values by expected growth and sampling rates.

Backends also differ in their ability to search trace attributes, preserve links, display long-running spans, and correlate traces with logs and metrics. Validate the complete path, not just the SDK export. Create a test request that traverses a model gateway, vector database, message queue, and external tool, then confirm that identifiers, timing relationships, attributes, errors, and links appear correctly in the chosen backend. If the system must migrate between vendors, keep a backend-neutral export path and avoid making application code depend on one vendor’s query language.

## Avoid common architectural mistakes

The first mistake is treating a framework’s internal callback as the complete agent trace. Frameworks may represent planning, tools, and state transitions differently, and they may hide the time spent waiting for external services. Instrument the application’s orchestration boundaries so the trace remains meaningful even when the framework changes. The second mistake is recording every message and tool result by default. This produces noisy, expensive, and potentially noncompliant telemetry. Capture structured metadata first, and make content capture an explicit, governed feature.

Another mistake is using the same span for both the agent and the foundation model. These operations have different owners, failure modes, and audiences. A provider call should have its own child span with model-specific attributes, while the agent span records the decision to call the model. Similarly, do not use one enormous agent.think span for a multi-minute workflow. That approach hides tool latency, queue delays, and partial failures. Conversely, do not create a span for every internal function; the result is difficult to interpret and can distort the backend’s ingestion statistics.

Finally, assume that a trace proves correctness. A trace can show that an agent called a tool, but not whether the tool was authorized, the retrieved evidence supported the answer, or the final result satisfied the user. Security policy checks, evaluation results, and business outcomes need explicit attributes or linked records. Good architecture makes technical execution observable while keeping semantic judgment in the evaluation and governance layers.

## Decide when to act and how to mature the design

A minimal design is sufficient for a prototype with low volume, non-sensitive inputs, and a single agent process. Create a root request span, add model and tool spans, propagate context through HTTP, and store only operational metadata. Before production, act on the architecture if the system handles personal or regulated data, uses multiple agents, makes external side effects, or needs to explain latency and cost. Those conditions justify a formal data-classification policy, collector-side redaction, controlled sampling, and separate evaluation telemetry.

Introduce richer conventions gradually. In the first release, standardize span names, request correlation, model identifiers, token counts, tool status, and error categories. In the second, add planning and delegation relationships, retrieval quality metadata, policy decisions, and links to evaluator runs. In the third, introduce workflow-level metrics, cohort comparisons, regression detection, and cost-quality optimization. Review the design quarterly against actual trace volume, backend performance, incident patterns, and changes in agent frameworks.

A useful maturity test is whether an engineer can answer four questions from one request without reading source code: What did the agent attempt? Which external operations were involved? Where did latency or cost accumulate? Why was the final result accepted, rejected, or failed? OpenTelemetry can provide the evidence for those answers, but only a deliberate hierarchy, propagation strategy, privacy boundary, and evaluation model will make the evidence actionable. The right goal is not maximal instrumentation; it is a trace that is complete enough for operations, safe enough for governance, and precise enough to improve the agent.

## Quick answers

### What is the best OpenTelemetry pattern for an AI agent?

Use one correlated trace per user request or autonomous task, with spans for planning, model calls, retrieval, tools, handoffs, and final response generation. Propagate W3C Trace Context through synchronous and asynchronous boundaries, and apply the same trace model across every agent service.

### Should prompts and agent reasoning be stored in OpenTelemetry traces?

Usually not by default. Prompts, retrieved documents, tool results, and model reasoning may contain credentials, regulated data, or third-party confidential information. Store bounded, redacted prompt metadata by default, and gate full-content capture with explicit retention, access, and deletion policies.

### How many spans should a typical AI agent generate?

There is no universal number. A simple request may produce 5 to 15 meaningful spans, while a research agent that runs 40 tool calls can produce hundreds unless intermediate results are summarized. Group repeated polling or low-level framework operations and apply sampling before exporting.

### Can OpenTelemetry replace an agent evaluation platform?

No. OpenTelemetry supplies operational telemetry, but an evaluation platform is needed to score task success, answer correctness, tool selection, policy compliance, and sometimes human preference. The best systems use trace IDs to join operational records with evaluation results and business outcomes.

### What is the minimum useful OpenTelemetry setup for a small agent?

Instrument the API boundary, model gateway, tool clients, and external agents; propagate trace context; and export spans to one searchable backend. Begin with latency, token usage, estimated cost, errors, and tool outcomes before adding fine-grained prompt evaluation or automated quality scoring.

Canonical: https://agustin-otegui.com/knowledge/how_should_you_design_an_opentelemetry_trace_architecture_for_ai_agents.php
Markdown: https://agustin-otegui.com/knowledge/how_should_you_design_an_opentelemetry_trace_architecture_for_ai_agents.php/index.md
