Atlas · skill

Document Chunking

Document chunking divides source material into units that can be indexed, retrieved and supplied to a model. Good boundaries preserve enough local meaning to answer a question while keeping units selective, traceable and small enough for the retrieval and generation system that will use them.

conceptIndexing & Chunking

What it is

A chunk can be a fixed token window, paragraph, section, table or another document-aware unit. Overlap repeats some content across boundaries so a split does not discard continuity, but it also creates redundancy. Chunk size interacts with the embedding model's input limit, the granularity of relevance labels and the amount of context the generator receives. A short chunk may lose a definition's subject; a long one may bury the relevant sentence. Chunking changes the evidence units in the index, not merely the formatting of a document for display.

What the work involves

The practitioner tests several boundary and size policies on annotated questions, preserving source identifiers, headings and locations. Tables and code blocks may need special handling so structure is not split arbitrarily. Evaluation inspects retrieval recall and final-answer support, including the additional cost of overlap and duplicated context. Useful artifacts include the splitter configuration, sample chunks and an error analysis of missing or misleading boundaries. Changing the policy requires rebuilding affected index entries and checking whether existing source links still point to the intended material.

Illustrative example

A troubleshooting manual has a warning at the end of one page and the procedure on the next. Fixed page-sized chunks retrieve the steps without the warning. A document-aware policy groups the warning with its procedure and retains the section title as context. Tests then check whether queries about the procedure retrieve both. The team also avoids combining several unrelated procedures into a large chunk that would make every answer appear supported by a broad but poorly targeted block.

Limits and common mistakes

There is no universally best chunk size. Results depend on document structure, query style, embedding behavior and downstream context selection. Overlap can inflate apparent retrieval success by producing near-duplicate hits, while generated chunk summaries can introduce inaccuracies. Quality checks should preserve original source evidence and compare complete pipelines. The useful unit is one that supports the intended question with the necessary qualifiers, not one selected solely because a framework offers that default.

Prerequisites

  • Chunking only makes sense in the context of a RAG pipeline — it's the data preparation step that determines retrieval quality

  • mediumNLP

    Understanding token counts, embedding window sizes, and semantic boundaries requires NLP foundations

Related skills

Sources and further reading

Last updated: 2026-10-10