AI Evaluation & Observability
25 skills · ontology graph below shows relations within this section.
What this domain covers
This edition groups 25 capabilities in AI Evaluation & Observability across 10 named categories. The inventory contains 18 concepts and 7 tools. Open an entry for its mechanism, practical workflow, example, limitations, and primary references.
Current category labels: Benchmarking · Debugging & Diagnostics · Evaluation Design · Evaluation Frameworks · LLM Testing · Monitoring & Drift · Observability & Tracing · Output Quality & Review · and 2 more
Frequent learning foundations
- Model Evaluation supports 5 mapped skills
- Retrieval-Augmented Generation supports 4 mapped skills
- LLM Evaluation Frameworks supports 3 mapped skills
- LLM Observability supports 3 mapped skills
- Statistical Inference supports 3 mapped skills
Skills in this section
Benchmark analysis examines what an evaluation result actually measures and how far it can support a decision. It considers task coverage, data construction, scoring, uncertainty and comparison conditions, preventing a leaderboard number from being mistaken for a general account of model quality or suitability for an application.
LLM benchmarking runs a documented set of tasks to compare language models under a defined protocol. It requires consistent inputs, settings and scoring, with separate attention to knowledge, reasoning, code or interaction capabilities; a benchmark score is evidence about that protocol rather than a universal measure of usefulness.
Stochastic system debugging investigates failures in applications whose outputs or execution paths can vary between runs. It combines reproducible configuration, detailed traces and repeated controlled experiments, so a developer can distinguish a systematic defect from sampling variation, nondeterministic infrastructure or an external dependency that changed.
Agent evaluation measures whether an agent completes tasks correctly while using tools, state and permissions appropriately across an execution trajectory. It goes beyond scoring a final sentence: the evaluator may need to inspect actions, intermediate evidence, environment changes and whether the result actually satisfies the user's goal.
LLM evaluation design defines what acceptable behavior means and how to measure it on representative cases. It connects application goals to test data, scoring and release decisions, with explicit coverage of uncertainty, edge cases and the difference between a correct-looking response and a result supported by independent evidence.
LLM-as-judge uses a language model to score, classify or compare another output according to an evaluation rubric. It can scale some review tasks, but the judge is itself a probabilistic system whose agreement with expert judgments, sensitivity to presentation and susceptibility to misleading content must be measured.
DeepEval is a framework for evaluating language-model applications with test cases, configurable metrics and execution workflows. It helps teams automate checks and inspect results, while leaving the meaning of quality, the relevance of test data and the validity of model-based scoring to the evaluation design.
LLM evaluation frameworks provide software for organizing test datasets, running systems and evaluators, and comparing results across changes. They make measurement easier to repeat, but they do not determine what quality means; selecting cases, validating metrics and interpreting failures remain central parts of evaluation work.
LLM testing checks that a model-powered application meets its functional contracts and handles known failure cases. It combines deterministic software tests with task-level evaluation of probabilistic outputs, so schema correctness, tool behavior, regression risks and answer quality are examined through methods appropriate to each property.
Data drift is a change in the distribution of data observed by a system compared with a reference period or dataset. It can signal that deployment inputs no longer resemble development data, but a detected change does not by itself prove model degradation or explain whether retraining is the right response.
ML monitoring observes a deployed model and its data pipeline over time to detect operational failures, distribution changes and deterioration in task outcomes. It connects measurements to investigation and response, distinguishing a service that is available from a model that still produces useful predictions for the populations it serves.
LLM observability makes the behavior of a language-model application inspectable through traces, metrics and related records. It links prompts, retrieval, model calls and tool execution to a request outcome, helping teams investigate failures and resource use without assuming that a trace alone establishes the quality or truth of an answer.
Langfuse is a platform and toolkit for observing and evaluating language-model applications, with features for tracing and prompt management. The skill involves instrumenting meaningful application steps, connecting runs to their configuration and designing useful evaluation or review workflows, while controlling which sensitive content enters stored telemetry.
AI output verification checks a generated result against evidence, rules or observable behavior before relying on it. It distinguishes plausible language from supported claims and correct actions, choosing a stronger check when available rather than treating model confidence or a second fluent response as sufficient proof.
RAG evaluation measures how retrieval and generation jointly produce answers from an evidence collection. It separates whether relevant material was found, whether the answer used that material faithfully and whether it addressed the question, so a single polished response or overall score does not hide the stage that failed.
Ragas is a framework for evaluating retrieval-augmented generation and related language-model workflows through configurable metrics and test data. It offers ways to assess context and answers, but using its scores responsibly requires inspecting metric definitions, supplying the required evidence and validating automated judgments against the application's real quality criteria.
Hallucination detection identifies generated claims that are unsupported, contradicted or otherwise unreliable under a specified evidence standard. It can use source comparison, consistency checks or trained evaluators, but the detector's scope must be explicit because contextual support, real-world factuality and repeated model agreement are different properties.
BERTScore evaluates generated text by matching contextual token representations with those of reference text. It can recognize semantic similarity beyond exact word overlap, but it measures a relationship to the supplied references rather than directly proving factual correctness, completeness or suitability for a particular user task.
BLEU is a reference-based text generation metric built from modified n-gram precision and a penalty for overly short candidates. It originated in machine translation evaluation and remains a useful reproducible baseline, while its dependence on lexical overlap limits what it can say about open-ended answers, factuality or user usefulness.
ROUGE is a family of reference-based metrics that compare generated text with reference summaries using lexical or sequence overlap. It provides reproducible signals about content overlap, commonly emphasizing recall, but different variants measure different relationships and none alone establishes factual consistency or the overall quality of a summary.
TruLens is a toolkit for instrumenting and evaluating language-model applications through recorded executions and feedback functions. It can connect evaluation scores to the application steps and evidence that produced an answer, while leaving the developer responsible for selecting meaningful checks and validating their interpretation.
Evidently is a framework for evaluating and monitoring data and AI systems through configurable metrics, tests and reports. It can help examine data quality, drift and model behavior, but the practitioner must choose reference data, thresholds and response rules that make the measurements meaningful for the deployed task.
LangSmith provides tracing, datasets and evaluation workflows for language-model applications. The skill is connecting application executions to their inputs, configuration and quality checks, so developers can compare changes and investigate failures while retaining clear boundaries around sensitive content and the assumptions of each evaluator.
OpenTelemetry is a vendor-neutral set of APIs, SDKs and conventions for collecting telemetry such as traces, metrics and logs. For AI applications, it helps connect model and tool calls to the surrounding system, while semantic conventions and content-capture policies require explicit versioning and attention to data sensitivity.
RAG faithfulness evaluation checks whether claims in a generated answer are supported by the retrieved context supplied to the model. It isolates contextual support from answer relevance and real-world truth, requiring explicit rules for claim decomposition, evidence matching and aggregation so the resulting score has a defensible interpretation.