Retrieval-Augmented Generation
Retrieval-augmented generation supplies a generative model with external evidence found for the current request. The application indexes or searches a collection, selects relevant material and uses it during answer generation, making knowledge updates and source attribution possible without relying only on information encoded in the model's weights.
Also searchable as: Retrieval-Augmented Generation (RAG), retrieval augmented generation
What it is
A typical RAG pipeline prepares documents, indexes searchable representations, retrieves candidates for a query and passes selected evidence to a generator. Retrieval may be lexical, vector-based, graph-assisted or combined. The original RAG research integrated retrieved passages with a generative model; deployed applications often use a simpler explicit retrieve-then-prompt architecture. RAG differs from fine-tuning because adding or changing evidence in the collection does not inherently update the model's parameters. It also differs from merely providing a long document: retrieval chooses relevant evidence from a larger source space for each request.
What the work involves
The practitioner builds ingestion with source versioning, chooses retrieval and context policies and evaluates questions with known supporting evidence. Tests distinguish failure to retrieve a passage from failure to use it correctly. Answer checks inspect coverage, unsupported claims and citations. Useful artifacts include the indexing pipeline, relevance labels, generation contract and traces of selected passages. Operational work covers access controls, document updates and deletion. The application needs an abstention or clarification path when the collection cannot support a reliable answer.
Illustrative example
A product assistant answers a question about an installation requirement. It searches current manuals, retrieves the matching model's procedure and uses those passages to explain the requirement with a source reference. When an older manual differs, source metadata prevents the older revision from silently becoming the answer. A test asks about an undocumented configuration: the assistant should state that the evidence is missing instead of filling the gap with a plausible instruction learned during pretraining.
Limits and common mistakes
RAG can reduce some knowledge gaps but does not eliminate hallucination. Retrieval may miss the evidence, return an unauthorized document or rank a near match above the correct item. The generator may ignore qualifiers or combine conflicting passages incorrectly. Quality depends on the collection and complete pipeline, so adding a vector database alone does not establish a trustworthy system. Grounding, source reliability and access enforcement require their own checks.
Prerequisites
- hardNLP
RAG combines retrieval (embeddings, vector search) with generation (LLM) — NLP foundations are the glue
- hardVector Databases
The retrieval step in RAG requires a vector database to store and search document embeddings
- mediumTransformer Architecture
Understanding how the LLM processes retrieved context (attention over concatenated tokens) helps debug RAG quality issues
Related skills
- ← is part of: AI Grounding & Citations
- ← is subcategory of: Agentic RAG
- ← is part of: Document AI
- ← is part of: Document Chunking
- ← is part of: Embedding Models
- ← is subcategory of: GraphRAG
- ← is part of: Hybrid Search
- ← is an instance of: LlamaIndex
- ← is subcategory of: Multimodal RAG
- ← is part of: Query Optimization
- ← is part of: RAG Evaluation
- → is subcategory of: GenAI
- ← is part of: Search Re-Ranking
- ← is subcategory of: Secure RAG
- ← is subcategory of: Self-Reflective RAG
- ← is part of: Semantic Search
- ← is part of: Vector Databases
- ← is part of: Visual Document Retrieval
Sources and further reading
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Primary retrieval-plus-generation formulation and comparison with parameter-only generation.
Last updated: 2026-10-10