Atlas · skill

LLM Observability

LLM observability makes the behavior of a language-model application inspectable through traces, metrics and related records. It links prompts, retrieval, model calls and tool execution to a request outcome, helping teams investigate failures and resource use without assuming that a trace alone establishes the quality or truth of an answer.

conceptObservability & Tracing

What it is

An LLM request may contain several nested operations: retrieving passages, calling a model, executing a tool and verifying an output. Distributed traces connect these steps through spans and timing, while metrics summarize patterns such as errors, latency and usage. Application annotations can attach evaluation scores or feedback. Observability differs from evaluation: a trace shows what happened, while an evaluator assesses it against a requirement. Generative AI semantic conventions can improve consistency across instrumentation, but their versions and supported attributes need attention. Capturing content is optional and carries different privacy implications from recording operational metadata.

What the work involves

The practitioner instruments meaningful boundaries and propagates request context across asynchronous calls and tools. It records model and prompt versions, errors and usage in a consistent schema. Sampling, redaction and retention are planned before storing inputs or outputs. Useful artifacts include trace examples, dashboards and an incident investigation path. Evaluation labels are linked to traces so a poor answer can be inspected for missing evidence or wrong actions. Instrumentation overhead and gaps are measured, especially when streaming or retries split one user request into several calls.

Illustrative example

A document assistant becomes slower after a release. Traces show that retrieval time is stable but an added verification step doubles model calls. A separate quality review finds unsupported answers linked to empty retrieval results. The team can address the latency and evidence failures independently because the trace preserves step boundaries and configurations. A dashboard of average response time alone would not reveal either mechanism or distinguish a successful answer from a fast incomplete response.

Limits and common mistakes

Telemetry can be incomplete, sampled or misleading if spans lack consistent boundaries. Token counts and costs also depend on provider reporting. Logging full content can expose sensitive information without improving diagnosis. Observability should preserve enough evidence to investigate behavior under explicit access and retention rules. A well-instrumented system is easier to understand, but factual correctness, task success and permission compliance still require their own checks.

Prerequisites

  • LLM observability tools monitor inference engines — understanding what the engine does helps interpret the traces

  • Tracing agent workflows (tool calls, reasoning steps) is a primary LLM observability use case

Related skills

Sources and further reading

Last updated: 2026-10-10