Atlas · skill

Retrieval-Augmented Generation

Retrieval-augmented generation supplies a generative model with external evidence found for the current request. The application indexes or searches a collection, selects relevant material and uses it during answer generation, making knowledge updates and source attribution possible without relying only on information encoded in the model's weights.

Also searchable as: Retrieval-Augmented Generation (RAG), retrieval augmented generation

conceptRAG Architecture

What it is

A typical RAG pipeline prepares documents, indexes searchable representations, retrieves candidates for a query and passes selected evidence to a generator. Retrieval may be lexical, vector-based, graph-assisted or combined. The original RAG research integrated retrieved passages with a generative model; deployed applications often use a simpler explicit retrieve-then-prompt architecture. RAG differs from fine-tuning because adding or changing evidence in the collection does not inherently update the model's parameters. It also differs from merely providing a long document: retrieval chooses relevant evidence from a larger source space for each request.

What the work involves

The practitioner builds ingestion with source versioning, chooses retrieval and context policies and evaluates questions with known supporting evidence. Tests distinguish failure to retrieve a passage from failure to use it correctly. Answer checks inspect coverage, unsupported claims and citations. Useful artifacts include the indexing pipeline, relevance labels, generation contract and traces of selected passages. Operational work covers access controls, document updates and deletion. The application needs an abstention or clarification path when the collection cannot support a reliable answer.

Illustrative example

A product assistant answers a question about an installation requirement. It searches current manuals, retrieves the matching model's procedure and uses those passages to explain the requirement with a source reference. When an older manual differs, source metadata prevents the older revision from silently becoming the answer. A test asks about an undocumented configuration: the assistant should state that the evidence is missing instead of filling the gap with a plausible instruction learned during pretraining.

Limits and common mistakes

RAG can reduce some knowledge gaps but does not eliminate hallucination. Retrieval may miss the evidence, return an unauthorized document or rank a near match above the correct item. The generator may ignore qualifiers or combine conflicting passages incorrectly. Quality depends on the collection and complete pipeline, so adding a vector database alone does not establish a trustworthy system. Grounding, source reliability and access enforcement require their own checks.

Prerequisites

  • hardNLP

    RAG combines retrieval (embeddings, vector search) with generation (LLM) — NLP foundations are the glue

  • The retrieval step in RAG requires a vector database to store and search document embeddings

  • Understanding how the LLM processes retrieved context (attention over concatenated tokens) helps debug RAG quality issues

Related skills

Sources and further reading

Last updated: 2026-10-10