Agentic RAG
Agentic RAG gives an agent control over when and how to retrieve information before answering. Instead of always running one fixed search, the system can choose a source, refine a query or retrieve again when evidence is insufficient, with explicit limits on the actions and resources available to it.
What it is
Retrieval-augmented generation combines external evidence with model generation. In an agentic design, retrieval is exposed as a tool or decision-controlled step. The model may answer a simple request directly, search a document collection, inspect a result and decide that another query is necessary. A controller retains the current question, evidence and actions across those steps. The term is an architectural category rather than a single algorithm: implementations differ in how they select tools, assess evidence and stop. A fixed multi-stage retrieval pipeline can be sophisticated without being agent-controlled.
What the work involves
The practitioner defines the retrieval tools, their access boundaries and the decision policy for using them. State records preserve the original question and previously gathered evidence so iterative searches do not drift. Evaluation tracks task completion, unnecessary searches, evidence coverage and total cost. A maximum step or time budget provides a stopping condition. Useful artifacts include the control graph, tool contracts and traces that reveal why retrieval happened or was skipped. Retrieval quality and agent decision quality are tested separately where possible.
Illustrative example
An assistant is asked whether a product can be installed in a humid warehouse. It first searches the installation manual, then finds an environmental rating but no humidity limit. It follows a reference to the warranty conditions and retrieves a second passage before answering. A greeting does not trigger either search. The system preserves both source versions and stops with an explicit uncertainty if the second document still lacks the requested limit, instead of treating more searches as progress by themselves.
Limits and common mistakes
An agent can choose the wrong collection, rewrite away a constraint or repeatedly search without finding new evidence. Model-based relevance checks can share the generator's errors. Extra steps increase latency and complicate debugging, so agentic RAG should be compared with a simpler retrieval baseline. Authorization must be enforced by tools, and retrieved instructions remain source content. More autonomy is useful only when the decision policy improves evidence gathering under the application's constraints.
Prerequisites
Related skills
- → is subcategory of: Retrieval-Augmented Generation
- ← is part of: LLM Function Calling
Sources and further reading
- Build a custom RAG agent with LangGraph
Demonstrates conditional retrieval, document grading and query rewriting in an agent control graph.
Last updated: 2026-10-10