Entity resolution and deduplication
Entity resolution asks whether two mentions identify the same real-world thing. Deduplication applies that decision to stored records. Both seek a balance: missed matches fragment knowledge, while false merges combine information that should remain separate.
| Python’s configurable resolvers and manual deduplication review methods are Bolt capabilities. NAMS performs resolution server-side and returns service-specific outcomes. See Backend capabilities and Bolt and NAMS before applying a workflow. |
Since 0.7.0, Bolt message ingestion resolves extracted mentions by default (resolution.resolve_on_ingest=True). Before 0.7.0 only add_entity deduplicated, and ingestion stored one node per distinct surface form. See Where resolution happens.
|
The entity resolution problem
An abbreviated name can identify an existing entity, a different entity, or nothing specific enough to resolve. Context and authoritative identifiers matter more than textual resemblance.
| Mentions | Identity decision | Evidence needed |
|---|---|---|
"Northstar Robotics" and "Northstar Robotics, Inc." |
Possible aliases |
The same verified organization identifier |
"Maya Chen" and "M. Chen" |
Unresolved from names alone |
Matching person identifier and context |
"Apple iPhone 15 Pro" and "iPhone 15 Pro" |
Possible aliases |
Matching model, capacity, color, and catalog identifier |
"iPhone 15 Pro" and "iPhone 15 Pro Max" |
Distinct variants |
Keep separate product records |
"Nike Air Max" and "Nike Air Max 90" |
Product family versus a specific model |
Do not merge without a more specific identity |
"order NR-2048" and "the package" |
Unresolved from names alone |
A source linking that package to that exact order |
A shared family name, brand, or description is evidence of relatedness, not necessarily identity.
Why identity matters
A graph split across duplicate nodes can lose useful context: a message mentions one spelling while a purchase or tool result points to another. Counts can also be misleading when aliases are counted separately.
The opposite error is more damaging in some domains. Merging two product variants can mix inventory or prices; merging people can attach one person’s statements to another. A resolution policy should consider the cost of a false merge as well as the cost of missed matches.
Evidence used by resolution
Exact-name matching detects repeated forms. Fuzzy matching tolerates spelling and word-order differences. Embedding similarity finds semantically related candidates. Each is a signal, not an identity proof: nearby vectors can describe competitors or different variants of the same product.
Type-constrained matching reduces some false positives, such as an organization named "Apple" and an object representing the fruit. It does not distinguish two people with the same name or guarantee that a broad OBJECT category contains only interchangeable products.
On Bolt, the default OntologyResolver (what resolution.strategy="composite" builds) composes these signals rather than replacing them. It puts the candidate search in Cypher, lets deterministic evidence (an exact normalized match, a declared ontology alias, an acronym expansion) settle the score before any similarity computation, and takes its aliases and per-type bands from the ontology, so a controlled vocabulary works even with embeddings switched off.
The consequence of type-constrained blocking is that a cross-typed duplicate, such as one company extracted once as ORGANIZATION and once as OBJECT, is never found. Fix the typing, usually with a sharper ontology description, rather than loosening the blocking.
The resolver’s five stages
The stages are separable, and knowing which one a bad decision came from is most of the work of tuning:
| Stage | What happens |
|---|---|
1. Normalize |
|
2. Block |
Candidates come from Cypher, and blocking never crosses the POLE+O type: "Apple" the company and "Apple" the product embed almost identically. One exact-key query per entity type covers the whole message; the wider buckets (head/tail token prefix, then the entity vector index) are queried only when the exact bucket produced no exact or alias match. An oversized token-prefix bucket (more than 60 rows) carries no signal and is discarded. |
3. Score |
Deterministic evidence short-circuits: an exact normalized match scores |
4. Band |
At or above |
5. Cluster |
A second pass clusters the mentions that matched nothing stored against each other, anchored on the first one seen, without counting their shared context window as corroboration. Why a message needs two passes explains why. |
Why a whole-token prefix needs corroboration
The prefix rule is where most false merges would otherwise come from. A whole-token prefix is not an identity, whatever the type: "Apple" and "Apple Bank" are different companies, "Kansas" and "Kansas City" different places, "Paris" and "Paris Hilton" different people. So the rule scores 0.92 only when independent surrounding context agrees (context similarity 0.60 or above), and 0.88, inside the review band, otherwise.
Name embeddings do not count as agreement: a name and its extension embed alike whether or not they are one entity (all-MiniLM-L6-v2 scores "Acme Bank" against "Acme Corp" at 0.76, above "Northwind" against "Northwind Logistics" at 0.73). The rule replaces the weighted blend rather than being combined with it, because RapidFuzz rates every whole-token prefix pair at exactly 0.90, the default auto-merge line, which would merge exactly the pairs the rule exists to hold back.
Where resolution happens
Two write paths reach the same decision through the same component:
long_term.add_entity(…)-
Deduplicates explicitly and returns an
(Entity, DeduplicationResult)tuple. This path predates 0.7.0 and keeps its shape. - Message ingestion (
short_term.add_message,add_messages_batch,extract_entities_from_session) -
Resolves every extracted mention while storing the message. This 0.7.0 change alters what the graph looks like: fewer entity nodes than mentions, aliases accumulating on survivors, and pending
SAME_ASedges appearing without anyadd_entitycall. One setting,resolution.resolve_on_ingest=False, restores the earlier behavior.
OntologyResolver is what turns two paths into one decision. When it is the configured resolver, add_entity delegates its duplicate check to the resolver, so both paths share one normalization, one blocking strategy and one set of thresholds. Any other resolver keeps the earlier embedding-similarity check for add_entity and leaves ingestion unresolved.
Resolution is an enhancement, never a gate: if a blocking query fails, the failure is logged and the mentions are stored unresolved.
Why a message needs two passes
Resolving each mention against stored entities is not enough. A message that mentions "Acme" and "Acme Corp" for the first time has no stored anchor for either, so a single pass would create two nodes and fragment the graph from the first message. The resolver therefore runs a second pass over the still-unmatched mentions, clustering them against each other and anchoring on the first one seen. Two mentions from the same message share one context window, so context does not count as corroboration between them; otherwise "Apple Bank has no relationship with Apple, the phone maker" would corroborate itself into a merge.
Three bands, not two
A binary merge-or-create decision forces a bad trade: merge aggressively and corrupt entities, or never merge and fragment them. A middle band avoids it. At or above review_threshold but below auto_merge_threshold, the new node is created and a pending SAME_AS edge records the suspicion, so a reviewer can examine the pair without blocking ingestion. A steady stream of rejected reviews at high confidence signals that auto_merge_threshold is about to make the same mistake silently.
The :SAME_AS review pattern
A :SAME_AS edge can record a potential duplicate without collapsing the records immediately. This preserves both candidates while an application or reviewer gathers evidence.
Review needs more than a similarity score: compare descriptions, provenance, domain identifiers, and relationships. Two missing SKU values are not evidence of a shared SKU. A meaningful comparison requires both identifiers to be present and to refer to the same identifier system.
Follow Review duplicates for the implemented Bolt workflow. The pending, confirmed, or rejected state of a candidate is different from a completed merge.
What a merge changes
A merge selects a surviving record and reconciles supported properties and relationships. It can affect every query that previously referenced either node, so the choice of primary record and handling of conflicting values matter.
An alias preserves an alternative name; it should not conceal a product variant or a separate legal entity. Do not assume that adding arbitrary metadata.aliases makes every resolver bypass similarity search. Aliases declared on an ontology entity type are different: the ontology resolver treats them as a controlled vocabulary. Use the actual alias and merge fields in the API reference.
Thresholds express a policy
Bolt’s resolution configuration separates an automatic-merge threshold (resolution.auto_merge_threshold, 0.90 by default) from a review threshold (resolution.review_threshold, 0.85). An ontology can override both for one entity type. Raising the automatic threshold reduces the set of candidates merged without review; moving the review threshold changes which ambiguous pairs reach the queue. The thresholds are score cutoffs, not guaranteed precision percentages.
Choose them using labeled matches and non-matches from the actual domain. A broad, low threshold is risky for similar product variants or same-name people. Per-entity deduplication controls can preserve intentionally distinct records; see Configure deduplication and Tune entity resolution.
Ambiguity, history, and related entities
Names need context. A company can change its name while retaining an identity, but a merger, acquisition, or subsidiary relationship may involve separate legal entities. Preserve dates and provenance when the distinction matters.
A parent company and a subsidiary should normally be modeled as related entities, not merged because they share a brand. Likewise, "Air Max" can denote a family while "Air Max 90" denotes a model. The graph can retain their relationship without asserting that they are the same thing.
Evaluating resolution quality
Evaluate both accepted and rejected matches against a labeled sample. Review queue size alone does not establish quality: a small queue might mean accurate resolution, an overly narrow candidate search, or excessive automatic merging.
Useful observations include false merges, missed aliases, disagreement by entity type, and time spent reviewing candidates. Choose workload-specific targets and record how the sample was labeled. No universal rejection-rate or queue-size percentage establishes that a deployment is healthy. The evaluation harness’s resolution dimension scores the configured resolver against labeled clusters; see Evaluate memory quality.
A conservative identity policy
Use stable domain identifiers where available, retain provenance, and keep distinct variants separate. Treat broad names as ambiguous until a source resolves them. Adjust thresholds after reviewing representative matches and non-matches.
A shared memory graph benefits from consistent naming, but consistency is not a reason to merge unrelated records. Periodic review should include automatic decisions as well as flagged candidates.