Atlas · skill

Self-Reflective RAG

Self-reflective RAG adds decisions that assess retrieved evidence or generated answers and use that assessment to retrieve again, revise or stop. Self-RAG and Corrective RAG are specific research approaches within this space; ordinary reflection instructions alone do not reproduce their trained mechanisms or evaluation claims.

conceptAdvanced RAG

What it is

A retrieval pipeline can fail because evidence is irrelevant, incomplete or poorly used. Self-reflective designs introduce signals that assess these stages. Self-RAG trains a model to emit reflection tokens governing retrieval and evaluating passages and generation. Corrective RAG uses a retrieval evaluator and corrective actions when retrieved evidence is unreliable. Other applications use a separate judge or programmed checks. The common idea is feedback controlling the retrieval–generation process, but the implementations differ in training, critique representation and fallback behavior. These differences matter when attributing a claimed result.

What the work involves

The practitioner identifies which failure a check is intended to detect and validates that check against independently labeled examples. A control policy specifies the response to low relevance, unsupported claims or missing evidence, including step limits. Traces retain critique decisions and the passages they evaluated. The artifact should name the actual mechanism rather than label every retry loop Self-RAG. End-to-end evaluation compares the reflective system with the same retriever and generator without correction, including costs and errors introduced by mistaken critiques.

Illustrative example

A product assistant retrieves passages about a similarly named model. A relevance check detects that the identifier differs and requests a new search with the exact model code. After generation, a support check finds that the answer's temperature limit has no matching passage and removes it or retrieves further evidence. In evaluation, reviewers inspect both cases where correction helped and cases where the checker wrongly rejected valid material. The system is judged on supported answers, not the number of reflective steps it performs.

Limits and common mistakes

A model's critique is fallible and may agree with its own unsupported answer. Repeated reflection can add confident wording without new evidence, while overstrict checks discard useful passages. Research mechanisms requiring trained reflection tokens cannot be assumed to emerge from a generic prompt. The correction policy needs an abstention path and resource limit. Independent validation of the evaluator is as important as evaluation of the generator it is intended to improve.

Prerequisites

  • Self-RAG adds a reflection loop ON TOP of a basic RAG pipeline — without understanding RAG, you can't add reflection to it

Related skills

Sources and further reading

Last updated: 2026-10-10