Entity Resolution
Entity resolution determines which records refer to the same real-world entity across imperfect or disconnected sources. It combines identity rules, similarity evidence and uncertainty handling to link or consolidate records without assuming that matching names are unique or that different identifiers always imply different entities.
What it is
Sources may represent one person, organization or product with varying names, addresses and identifiers. Entity resolution compares candidate pairs using exact agreement and approximate similarity, then decides whether the evidence supports a link. Blocking narrows the set of candidate comparisons; probabilistic methods estimate how informative agreements and disagreements are. Pairwise decisions may be combined into clusters, which introduces further consistency questions. The process differs from deduplicating identical rows because the underlying entity can have legitimately different records. A linked identity is a modeled conclusion with error risk, not a fact guaranteed by a high similarity score.
What the work involves
The practitioner defines entity meaning, chooses candidate-generation rules and creates labeled pairs or clerical review procedures. They evaluate false links and missed links separately, including difficult cases such as shared addresses and renamed organizations. Useful artifacts include matching rules, confidence thresholds and a reversible linkage table with provenance. Cluster review checks whether transitive links produce implausible groups. The team decides which links may be automatic and which require review, because merging two unrelated customers can have different consequences from failing to connect two copies.
Illustrative example
A retailer combines online and store loyalty records. Names alone create many plausible matches, so the resolver uses additional permitted evidence such as contact details and address agreement. Two family members share an address but have distinct purchase histories; review prevents an incorrect merge. Accepted links retain their source identifiers and evidence, allowing the team to reverse a decision if a later correction shows the records belong to different people.
Limits and common mistakes
Similarity is not identity, and blocking can silently exclude genuine matches before scoring begins. Rare groups or changing data formats can have different error rates. Cluster formation can amplify one bad pairwise link into a large incorrect merge. Evaluation needs representative ground truth and downstream impact analysis. Privacy constraints also limit which evidence can be used, so higher matching accuracy is not the only criterion for an acceptable solution.
Prerequisites
Related skills
- → is part of: Ontology Engineering
Sources and further reading
- Splink documentation
Official documentation of probabilistic record linkage and entity-resolution workflows.
Last updated: 2026-10-10