Atlas · skill

Hybrid Search

Hybrid search combines retrieval signals, commonly lexical matching and dense-vector similarity, to find candidates that either method might miss alone. The skill includes defining the combination rule and evaluating the resulting ranking, with particular attention to exact terms, semantic paraphrases and the behavior of filters.

conceptRetrieval Techniques

What it is

Lexical search scores shared terms and can preserve rare identifiers, while dense search compares learned representations that may connect different wording. A hybrid system runs both or incorporates both signals into a retrieval engine. Scores can be normalized and weighted, or rankings can be fused with a method such as reciprocal rank fusion. A fusion rule is necessary because raw scores from different retrievers usually have different scales. Hybrid retrieval is distinct from reranking: it combines candidate generation signals, while a reranker may then rescore the smaller candidate set more expensively.

What the work involves

The practitioner creates relevance judgments covering semantic queries and exact-match cases, then measures the hybrid against each component separately. Candidate counts, fusion weights and filters are tuned on development data and checked on held-out queries. Deduplication should preserve source identity. Useful artifacts include index mappings, query definitions, a documented fusion rule and rank-level error analysis. Operational measurement records latency and resource use for both retrievers, since combining them can increase work even when the returned result count stays small.

Illustrative example

A parts catalog contains model codes and descriptions. A query with an exact code benefits from lexical matching, while a query describing a leaking seal can benefit from semantic matching with a passage using different terminology. The hybrid ranking combines these candidates and retains the exact model restriction. Evaluation includes near-identical codes and paraphrased descriptions. A fusion configuration is accepted only if it improves those retrieval needs without allowing a thematically similar part to displace the specified one.

Limits and common mistakes

Combining weak retrievers does not guarantee a strong ranking. Score normalization, candidate truncation and filter placement can suppress relevant items before fusion. Hybrid gains also depend on query distribution and corpus vocabulary. A final similarity score is not a probability of relevance or evidence of truth. The combination should be reported with its concrete method and measured against component baselines, rather than presenting hybrid search as automatically superior for every collection.

Prerequisites

  • Hybrid search is an optimization OF the RAG retrieval step — you must understand basic RAG to know what you're improving

  • Hybrid search combines dense (semantic) and sparse (keyword) retrieval — understanding both sides is required

Related skills

Sources and further reading

Last updated: 2026-10-10