Multi-Vector Retrieval
Multi-vector retrieval represents a source item with several vectors and defines how their matches produce an item-level result. The vectors may represent chunks or derived views, or tokens and patches used in late interaction; the aggregation rule is central because multiple vectors alone do not define a retrieval method.
Also searchable as: Multi Vector Retrieval
What it is
A single vector compresses an item into one representation. Multiple vectors can retain separate facets, such as passages, summaries or visual regions. One system may retrieve these views and map hits back to a parent document, while another computes token-level interactions and aggregates their scores. ColBERT is a specific late-interaction approach that compares query-token vectors with document-token vectors. These designs share representational multiplicity but differ in scoring, indexing and storage requirements. Multi-vector retrieval is therefore an umbrella practice, not an automatic synonym for ColBERT or for splitting a document into independently returned chunks.
What the work involves
The practitioner states what each vector represents and how vector hits are deduplicated, combined or scored at source-item level. Evaluation compares with a single-vector baseline and checks whether items with more vectors receive an unintended advantage. Source mappings and update procedures must keep all representations synchronized. Useful artifacts include the representation schema, aggregation function, storage measurements and relevance results. For late interaction, token or patch handling and supported index mechanisms require explicit configuration. Latency and memory are assessed alongside quality gains.
Illustrative example
A technical manual is represented by passage vectors and a short overview vector, all linked to one manual identifier. Search may match a specific procedure or the overview; the application aggregates hits into a ranked manual result and retains the supporting passage. A separate experiment uses token-level late interaction for the same query set. The team compares these mechanisms rather than grouping their scores as if they were equivalent, and checks that long manuals do not win merely because they contain more indexed views.
Limits and common mistakes
More representations can increase storage, indexing and scoring cost without improving relevant retrieval. Poor aggregation can overcount duplicate evidence or bias ranking toward large documents. Generated views may also contain unsupported summaries. Quality checks need original source traceability and a clearly defined scoring rule. Any claimed improvement should identify the multi-vector design and baseline, since gains from one token-level method do not establish the value of every chunk or view aggregation strategy.
Prerequisites
- mediumEmbedding Models
The method depends on producing and interpreting multiple embeddings for one retrievable item.
Related skills
- → is subcategory of: Dense Retrieval
Sources and further reading
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Primary token-level multi-vector late-interaction retrieval design.
Last updated: 2026-10-10