Atlas · skill

Semantic Search

Semantic search retrieves material using learned representations intended to capture meaning beyond exact keyword overlap. The competence is selecting embeddings and similarity measures, constructing a useful index and evaluating retrieval against real information needs. Similar wording or a high vector score does not by itself establish that a result answers the query.

conceptText Understanding

What it is

A common approach encodes queries and candidate passages into vectors, then ranks candidates by cosine similarity or another compatible distance. Bi-encoders represent each item separately, enabling precomputed document vectors and efficient search. Cross-encoders jointly score a query and candidate and can rerank a smaller set at greater computational cost. Symmetric similarity tasks and asymmetric question-to-passage retrieval may require different training or encoding conventions. Chunk boundaries, metadata filters and index approximation affect which evidence is available. Semantic search is one retrieval technique; lexical search remains useful for exact identifiers and rare terms, and hybrid systems can combine the two signals.

What the work involves

Define relevant results with representative queries and graded judgments. Choose a model suited to the language, domain and query-document relationship, and preserve its encoding instructions and similarity convention. Build passages that retain enough context without hiding relevant detail in overly long chunks. Evaluate recall and rank quality against lexical and hybrid baselines, holding out source families and query variants. Inspect approximate-index misses separately from embedding errors. The result is an evaluated search configuration with documented corpus version, filters, encoding and ranking stages, plus a clear rule for insufficient or irrelevant matches.

Illustrative example

An illustrative technician searches manuals with a description of a fault rather than the official component name. Embedding retrieval finds relevant explanatory passages, but searches for an exact part number work better lexically. The developer combines the candidates and reranks them. Held-out queries include similar failures with different causes, and review checks whether retrieved passages contain the actual diagnostic step. A topically similar introduction is not counted as an adequate answer source.

Limits and common mistakes

Embeddings can miss negation, numbers and specialized identifiers, or retrieve a related topic without the needed fact. Similarity values are model-specific and generally not calibrated relevance probabilities. Approximate indexes add retrieval errors, and stale embeddings can misrepresent changed documents. Semantic search differs from question answering because retrieval returns candidates rather than a verified answer. Evaluate ranked relevance and evidence coverage, and retain exact-match support when the application's information needs require it.

Prerequisites

  • hardNLP

    Semantic search = query embedding → nearest-neighbor search in vector space. Without understanding embeddings, it's magic

Related skills

Sources and further reading

Last updated: 2026-10-10