Atlas · skill

Vector Databases

Vector databases store vector representations alongside identifiers and often metadata, providing similarity search and lifecycle operations for an application. They support retrieval infrastructure; deciding what to embed, which similarity means relevance and how retrieved evidence should be used remains a separate modeling and application responsibility.

conceptVector Search

What it is

A vector database receives embeddings from an encoding process and retrieves items near a query vector under a configured distance or similarity measure. It may provide approximate indexes, metadata filters, updates, persistence and distributed operations. Some products also offer lexical or hybrid search. The category includes embedded stores and managed or self-hosted services with different operational trade-offs. A vector database does not inherently know that a nearby item answers a question, and storing embeddings is distinct from training the model that produced them. Access control and consistency semantics depend on the implementation.

What the work involves

The practitioner compares candidate systems against realistic collection size, update rate, filtering needs and latency targets. It evaluates retrieval recall alongside memory, ingestion cost and operational reliability. Records should preserve embedding model version and source identity so incompatible representations are not silently mixed. Useful artifacts include the collection schema, index configuration, backup or recovery plan and relevance benchmark. Deletion and tenant isolation need end-to-end tests, especially where a retrieved vector maps to an external document with its own permissions.

Illustrative example

A company indexes product manuals with vectors, product identifiers and revision dates. A query searches only manuals for the selected product and returns passage identifiers for generation. The team measures whether the database finds annotated relevant passages under those filters and confirms that deleted revisions disappear from both search and source resolution. Selecting the store involves update and recovery requirements as well as raw search speed, because a fast stale result can still produce a wrong answer.

Limits and common mistakes

Nearest neighbors can be irrelevant, and approximate search adds another source of misses. Aggressive filtering may change recall or leave too few candidates. Vendor features and limits evolve, so product selection needs current documentation and workload testing. A database is one stage of a retrieval system rather than a guarantee of answer correctness. Comparing exact search and a lexical baseline helps identify whether errors originate in indexing, representation or the evidence collection itself.

Prerequisites

  • hardNLP

    Vector databases store and index embeddings — you must understand what vectors represent to choose the right index, metric, and parameters

  • Production vector DBs involve sharding, replication, and latency trade-offs — distributed systems literacy helps make informed choices

Related skills

Sources and further reading

  • Qdrant concepts

    Official vector-store concepts covering collections, vectors, payloads and similarity search.

  • Milvus overview

    Provides a second official implementation perspective on vector database functions.

Last updated: 2026-10-10