RAG Evaluation
RAG evaluation measures how retrieval and generation jointly produce answers from an evidence collection. It separates whether relevant material was found, whether the answer used that material faithfully and whether it addressed the question, so a single polished response or overall score does not hide the stage that failed.
What it is
A RAG pipeline can retrieve irrelevant context, omit a necessary passage or generate claims unsupported by good context. Evaluation therefore needs several dimensions. Retrieval metrics compare ranked candidates with relevance judgments or known supporting passages. Answer checks examine completeness, relevance and contextual support; references or expert labels may also establish correctness. Faithfulness to retrieved context is distinct from real-world truth because a source can itself be wrong. Model-based evaluators can estimate some dimensions, but their prompts and validation matter. The unit of evaluation is the configured pipeline, including data preparation and context selection.
What the work involves
The practitioner builds questions with supporting evidence and expected behavior, including unanswerable and conflicting-source cases. It retains retrieved passages and source versions during runs. Separate scores and error categories identify retrieval misses, unused evidence and unsupported claims. Useful artifacts include relevance labels, claim–source annotations and a baseline comparison. Evaluators are checked against human judgments, and cost and latency are recorded alongside quality. Changing chunking, embeddings or generation settings is tested end to end because a local metric gain can worsen the final answer.
Illustrative example
A policy assistant answers carry-over questions using a policy library. Evaluation finds that the relevant exception appears among candidates but is removed by context selection, leading to an incomplete answer. Another case retrieves an old policy and produces a perfectly faithful but outdated answer. These failures need different corrections. The report preserves retrieval coverage, source version correctness and answer support separately, allowing the team to improve selection or ingestion rather than simply replacing the generator.
Limits and common mistakes
Automated scores can disagree with expert review and may overlook subtle qualifiers. Reference answers can be incomplete, while evaluation questions generated from the same source may be unrealistically easy. Strong faithfulness does not establish source accuracy or answer completeness. Quality reports should name the dimensions, labels and pipeline version, and retain evidence for inspection. RAG evaluation is useful when it diagnoses concrete failure mechanisms instead of collapsing all behavior into one reassuring number.
Prerequisites
You cannot evaluate a RAG system without understanding its components (retrieval quality, generation faithfulness, grounding)
- mediumModel Evaluation
RAG evaluation uses metrics concepts (precision@k, recall, F1) adapted to retrieval+generation context
Related skills
- → is part of: Retrieval-Augmented Generation
- → is subcategory of: LLM Evaluation Design
- ← is subcategory of: RAG Faithfulness Evaluation
Sources and further reading
- Ragas: Automated Evaluation of Retrieval Augmented Generation
Defines distinct dimensions for retrieval context and generated-answer evaluation.
- Ragas faithfulness metric
Documents claim support as a context-dependent property rather than universal factual correctness.
Last updated: 2026-10-10