LLM Observability
LLM observability makes the behavior of a language-model application inspectable through traces, metrics and related records. It links prompts, retrieval, model calls and tool execution to a request outcome, helping teams investigate failures and resource use without assuming that a trace alone establishes the quality or truth of an answer.
What it is
An LLM request may contain several nested operations: retrieving passages, calling a model, executing a tool and verifying an output. Distributed traces connect these steps through spans and timing, while metrics summarize patterns such as errors, latency and usage. Application annotations can attach evaluation scores or feedback. Observability differs from evaluation: a trace shows what happened, while an evaluator assesses it against a requirement. Generative AI semantic conventions can improve consistency across instrumentation, but their versions and supported attributes need attention. Capturing content is optional and carries different privacy implications from recording operational metadata.
What the work involves
The practitioner instruments meaningful boundaries and propagates request context across asynchronous calls and tools. It records model and prompt versions, errors and usage in a consistent schema. Sampling, redaction and retention are planned before storing inputs or outputs. Useful artifacts include trace examples, dashboards and an incident investigation path. Evaluation labels are linked to traces so a poor answer can be inspected for missing evidence or wrong actions. Instrumentation overhead and gaps are measured, especially when streaming or retries split one user request into several calls.
Illustrative example
A document assistant becomes slower after a release. Traces show that retrieval time is stable but an added verification step doubles model calls. A separate quality review finds unsupported answers linked to empty retrieval results. The team can address the latency and evidence failures independently because the trace preserves step boundaries and configurations. A dashboard of average response time alone would not reveal either mechanism or distinguish a successful answer from a fast incomplete response.
Limits and common mistakes
Telemetry can be incomplete, sampled or misleading if spans lack consistent boundaries. Token counts and costs also depend on provider reporting. Logging full content can expose sensitive information without improving diagnosis. Observability should preserve enough evidence to investigate behavior under explicit access and retention rules. A well-instrumented system is easier to understand, but factual correctness, task success and permission compliance still require their own checks.
Prerequisites
- mediumLLM Inference Serving
LLM observability tools monitor inference engines — understanding what the engine does helps interpret the traces
- mediumAI Agent Design
Tracing agent workflows (tool calls, reasoning steps) is a primary LLM observability use case
Related skills
- → is subcategory of: MLOps
- → is subcategory of: ML Monitoring
- ← is part of: Stochastic System Debugging
- ← is an instance of: LangSmith
- ← is an instance of: OpenTelemetry
Sources and further reading
- OpenTelemetry observability primer
Explains traces, metrics and logs as complementary operational signals.
- OpenTelemetry GenAI semantic conventions
Official conventions repository for generative AI instrumentation and evolving attribute definitions.
Last updated: 2026-10-10