Atlas · skill

RAG Faithfulness Evaluation

RAG faithfulness evaluation checks whether claims in a generated answer are supported by the retrieved context supplied to the model. It isolates contextual support from answer relevance and real-world truth, requiring explicit rules for claim decomposition, evidence matching and aggregation so the resulting score has a defensible interpretation.

conceptRAG Evaluation

What it is

An evaluator divides an answer into statements and assesses whether the available context supports each one, often using an entailment model or language-model judge. A score may aggregate supported statements, but details such as compound claims, uncertainty and denominator selection affect the result. The supplied context is the reference, not every fact the evaluator happens to know. A faithful answer can still be wrong if the context is outdated or false, and an unfaithful claim can happen to be true but absent from the permitted evidence. These distinctions make faithfulness one dimension of RAG quality rather than overall correctness.

What the work involves

The practitioner defines atomic claim rules, allowed evidence and how partial or ambiguous support is labeled. They validate decomposition and support judgments against expert annotations, including numerical differences and missing qualifiers. Useful artifacts include claim–passage pairs, judge configuration, aggregation rules and error analysis. The evaluator receives the context actually used during generation, not a broader candidate set that could retroactively support the answer. Separate checks assess source reliability, answer completeness and relevance. Changes in the judge or decomposition prompt require renewed calibration before scores are compared.

Illustrative example

A manual states that a device supports outdoor use when protected from direct rain. An answer says it is suitable outdoors without qualification. Faithfulness evaluation should detect the omitted condition rather than mark the broad sentence supported because the manual contains outdoor use. Another answer faithfully repeats a limit from an obsolete manual; it passes contextual support but fails a separate version check. Testing both cases clarifies why a high faithfulness score is useful evidence about attribution, not a guarantee of a safe or correct instruction.

Limits and common mistakes

A judge can overlook subtle contradictions, split claims inconsistently or reward answers with fewer substantive statements. Aggregation can conceal one critical unsupported claim among many easy supported ones. Missing evidence is also different from explicit contradiction. Quality reports should preserve claim-level judgments and state how uncertainty is handled. Faithfulness is most useful when calibrated and combined with independent checks of source validity and task coverage, without pretending that contextual agreement establishes universal truth.

Prerequisites

  • Metric design requires an end-to-end understanding of RAG inputs, outputs and evaluation units.

Related skills

Sources and further reading

Last updated: 2026-10-10