Atlas · skill

pgvector

pgvector is a PostgreSQL extension that adds vector data types and similarity-search operations to a relational database. It lets an application keep embeddings near its existing records and query them with SQL, while requiring deliberate index selection, filtering and performance evaluation as the collection and workload grow.

toolVector Search

What it is

The extension stores vectors in PostgreSQL columns and exposes operators for supported distance or similarity calculations. Exact search compares eligible records directly, while approximate indexes such as HNSW and IVFFlat can reduce search work with a recall trade-off. Relational joins and ordinary metadata conditions remain available, which is useful when source data already lives in PostgreSQL. pgvector differs from an embedding model and from a separate vector service: it supplies vector operations within the database's existing transaction and operational environment. Index behavior depends on the query shape and configuration.

What the work involves

The practitioner defines vector dimensions and similarity consistently with the encoder, tests SQL queries and inspects query plans. Exact-search results provide a recall reference for approximate indexes. Filter selectivity, result limits and search parameters are tested together because an approximate candidate set may not contain enough matching records. Useful artifacts include migrations, index definitions, relevance tests and capacity measurements. Embedding updates should maintain source version consistency, and normal PostgreSQL access controls still need correct application use rather than reliance on a similarity predicate.

Illustrative example

A document application already stores records and access metadata in PostgreSQL. It adds an embedding column to searchable passages and uses a query that combines a permitted-document condition with vector ordering. The team compares exact results and an approximate index for both broad and highly selective filters. A test account with access to only a small folder verifies that search returns enough relevant authorized passages. If the approximate query misses them, search settings or the query plan are revised.

Limits and common mistakes

Sharing a relational database does not make every vector workload inexpensive. Large indexes consume resources alongside transactional queries, and approximate filtering can affect recall. Supported types and index features depend on the installed extension version. Evaluation should measure the actual deployment rather than assume it matches a dedicated vector service. pgvector is useful when SQL integration and database operations fit the application, with explicit checks for retrieval quality, concurrency and recovery.

Prerequisites

  • hardSQL

    pgvector extends PostgreSQL — you must understand SQL, indexing, and query planning to use it effectively

  • hardNLP

    Without understanding what embeddings are, pgvector is just a mysterious column type

Related skills

Sources and further reading

Last updated: 2026-10-10