BM25
BM25 is a lexical relevance scoring method that ranks documents using query-term matches, term rarity and document-length normalization. It is a strong baseline for text retrieval and the lexical component of many hybrid systems, especially when exact identifiers or specialist terms matter more than broad semantic similarity.
What it is
BM25 belongs to the probabilistic information retrieval tradition. Its score gives more weight to terms that are informative across the collection, reduces the marginal contribution of repeated occurrences and adjusts for document length. Parameters control aspects of term-frequency saturation and length normalization. Text analysis also matters: tokenization, stemming and field handling determine which terms match before scoring. BM25 differs from dense retrieval because it primarily uses lexical overlap rather than learned vector proximity. Its numeric scores are ranking signals within a query and index configuration, not direct probabilities of relevance.
What the work involves
The practitioner configures text analysis and searchable fields, establishes relevance labels and tunes scoring only where evaluation justifies it. Tests include identifiers, abbreviations, spelling variants and queries whose relevant document uses different wording. Useful artifacts include the index mapping, query strategy and measured ranking baseline. Comparing BM25 with dense or hybrid search helps identify whether errors come from vocabulary mismatch or candidate ranking. Field weights and filters are documented so a result's position can be explained without attributing every effect to the BM25 formula.
Illustrative example
A technical catalog contains short product codes and long descriptions. A query with the exact code should strongly favor the matching item even if another description is topically similar. The team tests field weighting so the code field contributes appropriately and analyzes whether tokenization splits meaningful punctuation. A paraphrased symptom query may require dense or expanded retrieval as well. BM25 remains the reference baseline, making it possible to see what additional semantic machinery improves and what precision it loses.
Limits and common mistakes
Lexical matching can miss paraphrases, synonyms and cross-language relevance. Scores are also affected by corpus statistics and document preparation, so thresholds do not transfer automatically between indexes. Aggressive stemming can merge distinct identifiers, while repeated terms can still distort ranking under unsuitable configuration. Evaluation should include the actual query vocabulary and field structure. BM25's simplicity and usefulness do not make it a universal solution or justify skipping a retrieval quality test.
Prerequisites
Related skills
- → is subcategory of: Hybrid Search
Sources and further reading
- Elasticsearch similarity configuration
Official documentation for BM25 scoring parameters and similarity configuration.
- The Probabilistic Relevance Framework: BM25 and Beyond
Primary author account of BM25's probabilistic basis and scoring design.
Last updated: 2026-10-10