OpenTelemetry
OpenTelemetry is a vendor-neutral set of APIs, SDKs and conventions for collecting telemetry such as traces, metrics and logs. For AI applications, it helps connect model and tool calls to the surrounding system, while semantic conventions and content-capture policies require explicit versioning and attention to data sensitivity.
What it is
A trace links operations through spans and propagated context, metrics summarize measurements and logs capture events. OpenTelemetry provides instrumentation and export mechanisms so data can be sent to an observability backend without requiring one vendor-specific application interface. Generative AI conventions describe attributes and operations for model or agent interactions, but this area evolves separately and its stability should be checked. OpenTelemetry is not itself a storage or visualization backend, nor an evaluator of answer quality. It defines and transports operational evidence that other systems can inspect or combine with evaluation labels.
What the work involves
The practitioner instruments meaningful operations, propagates context through asynchronous calls and configures an exporter and collector where appropriate. It chooses attributes that support diagnosis without unnecessary content capture. Sampling, redaction and retention are planned with the backend and application owners. Useful artifacts include span examples, telemetry schema and pipeline configuration. Tests verify that one user request can be followed across retrieval, inference and tools. SDK and convention versions are recorded so dashboard changes do not silently mix differently named or interpreted fields.
Illustrative example
A document assistant uses one backend for retrieval and another for model inference. OpenTelemetry trace context connects both operations to the user's request, showing where time was spent and whether a tool failed. Operational attributes identify the model and usage where supported, while document text is excluded from routine telemetry. A separate evaluation labels an answer unsupported and links that label to the trace. The combined evidence helps locate the failure without treating trace completeness as proof that the answer was correct.
Limits and common mistakes
Telemetry can be sampled, incomplete or inconsistent across libraries. Export failures and instrumentation overhead also need monitoring. Generative AI conventions may change, and capturing prompts or tool results can expose sensitive data. OpenTelemetry improves interoperability when correctly configured, but it does not automatically provide useful dashboards or quality judgments. Good implementation verifies data accuracy and controls content collection, preserving the distinction between observed execution and evaluated task success.
Prerequisites
Related skills
- → is an instance of: LLM Observability
Sources and further reading
- OpenTelemetry observability primer
Official explanation of traces, metrics and logs and their complementary roles.
- OpenTelemetry GenAI semantic conventions
Current official repository for generative AI instrumentation conventions.
Last updated: 2026-10-10