Atlas · skill

Contextual Retrieval

Contextual retrieval adds explanatory context to individual chunks before indexing them so that isolated passages remain meaningful during search. A specific Anthropic approach uses generated chunk context for both embeddings and lexical indexing; the broader practice is useful only when the added context accurately preserves the source's identity and meaning.

conceptIndexing & Chunking

What it is

A document chunk can contain an amount or pronoun without identifying the company, event or section it refers to. Contextual enrichment supplies a short description derived from the surrounding document, such as what entity and reporting period the chunk concerns. The enriched representation is then indexed, while the original passage remains available as evidence. This differs from query rewriting, which changes the search input, and from merely expanding context after retrieval. Implementations vary in how context is produced and whether it is used for dense, lexical or combined search.

What the work involves

The practitioner defines a context-generation procedure, checks sampled enrichments against original documents and compares retrieval with an unenriched baseline. Source identifiers and versions are retained so context can be refreshed when content changes. Evaluation includes short passages with missing antecedents and rare identifiers, where enrichment might help or accidentally mislead. Useful artifacts include the enrichment prompt or rules, indexed text format and relevance results. Generated context should be distinguishable from quoted source content in any downstream evidence view.

Illustrative example

A report chunk states that operating costs increased but does not repeat the business unit named earlier. The index adds a short contextual sentence identifying that unit and the reporting period. A query about the unit's costs can now match the passage more directly. Review confirms that the generated context names the correct unit and does not infer an increase percentage absent from the report. The answer system cites the original passage and surrounding section, rather than presenting the enrichment as an independent fact.

Limits and common mistakes

Generated context can introduce an incorrect entity, date or interpretation and make a wrong chunk easier to retrieve. It also adds preprocessing cost and update work. Improved retrieval in one reported experiment is not a guarantee across document types. A useful assessment compares relevance and downstream support under the actual collection, audits context accuracy and preserves the original source so enrichment never becomes an untraceable replacement for evidence.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

  • Contextual Retrieval

    Describes generating chunk context and incorporating it into embedding and lexical retrieval representations.

Last updated: 2026-10-10