Atlas · skill

TruLens

TruLens is a toolkit for instrumenting and evaluating language-model applications through recorded executions and feedback functions. It can connect evaluation scores to the application steps and evidence that produced an answer, while leaving the developer responsible for selecting meaningful checks and validating their interpretation.

toolEvaluation Frameworks

What it is

A feedback function evaluates selected parts of a recorded application execution, such as a question, retrieved context or answer. TruLens organizes these evaluations alongside traces or records. Its RAG triad distinguishes context relevance, groundedness and answer relevance, which are related but different properties. Feedback may use model-based judgments or other functions, depending on the configuration. TruLens is an implementation framework rather than proof that these scores are objective or sufficient. The input selectors and evaluator definitions matter because a score computed over the wrong context can be misleading even when execution succeeds.

What the work involves

The practitioner instruments the relevant application stages and verifies that feedback functions receive the intended inputs. It checks scores on labeled examples and records provider, prompt and aggregation settings. Useful artifacts include instrumentation code, selector tests, feedback configuration and examined execution records. Evaluations should preserve failures and missing values rather than quietly aggregate them away. Privacy planning controls stored content, and operational measurement includes evaluator cost. The resulting traces help connect a quality issue to retrieval or generation, but the metric still needs independent calibration.

Illustrative example

A RAG service records the question, selected passages and final answer. TruLens feedback checks whether the passages concern the question, whether claims are supported and whether the answer responds to the request. A test intentionally supplies irrelevant context with a plausible answer, checking that the feedback dimensions diverge as expected. Reviewers inspect each low score with its record. When a selector accidentally uses all retrieved candidates rather than the passages actually sent to the model, the instrumentation is corrected before comparisons continue.

Limits and common mistakes

Model-based feedback can be biased or miss subtle contradictions. Instrumentation gaps and incorrect selectors can also make scores describe a different interaction from the user's actual experience. High groundedness does not establish source truth or answer completeness. TruLens is useful when records and feedback provide interpretable evidence, with explicit validation of the recorded data, evaluator behavior and the quality dimensions the application needs.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10