Deduplication and provenance
Ingest-time deduplication and extraction-provenance operations from LongTermMemory API reference. Both are bolt-only. Signatures show types, keyword-only arguments (*), and defaults; they are reference declarations, not calls to execute directly.
Deduplication (bolt only)
Since 0.7 the duplicate bands come from ResolutionConfig: at or above auto_merge_threshold (default 0.90) a new name merges onto the matched entity as an alias, and at or above review_threshold (default 0.85) the entity is created with a pending SAME_AS edge for review. Per-type overrides come from the ontology (EntityTypeDef.resolution_threshold, EntityTypeDef.review_threshold).
Two write paths fill the same review queue. Message ingestion resolves extracted mentions with OntologyResolver by default (resolution.resolve_on_ingest=True). add_entity delegates its check to that resolver whenever it is the configured one, the bolt default, so both paths band identically; with any other resolver, add_entity falls back to the embedding-similarity check configured by DeduplicationConfig from neo4j_agent_memory.memory.long_term. add_entity checks only when an embedder is configured. There is no top-level DeduplicationStrategy or NAM_DEDUPLICATION settings group. See Deduplication configuration.
find_potential_duplicates
Find entities that are flagged as potential duplicates.
async def find_potential_duplicates(*, limit: int=100) -> list[tuple[Entity, Entity, float]]: ...
| Parameter | Description |
|---|---|
|
Maximum number of duplicate pairs to return |
Each pending SAME_AS pair is returned once, as the flagged (newly added) entity, then its existing match, then the edge’s confidence, highest confidence first. Pairs flagged during message ingestion are included.
merge_duplicate_entities
Merge two entities, keeping the target and marking source as merged. The source’s name is added to the target’s aliases, and the Bolt query copies the source’s edges of these types onto the target: inbound :MENTIONS (from Message), :RELATED_TO in both directions, :SAME_AS, outbound :EXTRACTED_FROM and :EXTRACTED_BY provenance, inbound :APPLIES_TO (from Preference), and inbound :TOUCHED (from ReasoningStep). Edges are copied, not moved: the merged-away source keeps its own edges — so the merge stays reversible and auditable — and is marked with merged_into/merged_at to exclude it from later deduplication scans. The copies are marked differently by type: :RELATED_TO copies keep the source edge’s id and gain a migrated_from property set to the source entity ID (nothing looks a :RELATED_TO edge up by id, so the shared id is not ambiguous); :EXTRACTED_FROM, :EXTRACTED_BY, :APPLIES_TO, and :TOUCHED copies gain migrated_from only; :MENTIONS copies carry no marker; and :SAME_AS copies are recorded with match_type: 'merged' but no source reference. Relationship types outside this list are not transferred automatically; review any custom links separately.
async def merge_duplicate_entities(source_id: UUID, target_id: UUID) -> tuple[Entity, Entity] | None: ...
| Parameter | Description |
|---|---|
|
ID of entity to merge from (will be marked as merged) |
|
ID of entity to merge into (will be kept) |
review_duplicate
Review a potential duplicate pair.
async def review_duplicate(
source_id: UUID,
target_id: UUID,
*,
confirm: bool,
) -> bool: ...
| Parameter | Description |
|---|---|
|
ID of first entity |
|
ID of second entity |
|
True to confirm as duplicate (merge), False to reject |
get_same_as_cluster
Get all entities in the same :SAME_AS cluster as the given entity.
async def get_same_as_cluster(entity_id: UUID) -> list[Entity]: ...
| Parameter | Description |
|---|---|
|
Entity ID to find cluster for |
get_deduplication_stats
Get statistics about entity deduplication.
async def get_deduplication_stats() -> DeduplicationStats: ...
DeduplicationResult.action is none, merged, or flagged; an ordinary new entity uses none. When add_entity delegates to OntologyResolver, the resolver’s review band maps onto flagged. Thresholds are inclusive. The flag, auto-merge, and fuzzy thresholds must be between 0 and 1, and auto-merge must be at least the flag threshold.
DeduplicationConfig
The thresholds apply to the embedding-similarity check, which runs when the store’s resolver is not an OntologyResolver. DeduplicationConfig.from_resolution_config(config) builds one from a ResolutionConfig, copying auto_merge_threshold, review_threshold (as flag_threshold), fuzzy_threshold, and candidate_limit (as max_candidates); MemoryClient builds its store’s config this way from settings.resolution.
Fields and defaults:
| Field | Type | Default | Description |
|---|---|---|---|
|
|
|
Whether |
|
|
|
Similarity at or above which the new name is merged into the existing entity as an alias. Matches |
|
|
|
Similarity at or above which a pending |
|
|
|
Also compare names with RapidFuzz; needs the |
|
|
|
Fuzzy name ratio at or above which the embedding and fuzzy scores are averaged. Matches |
|
|
|
Maximum number of vector-search candidates checked. Matches |
|
|
|
Only compare against entities of the same type. |
DeduplicationResult
Fields and defaults:
| Field | Type | Default | Description |
|---|---|---|---|
|
|
|
Whether a match reached the review band: |
|
|
|
|
|
|
|
ID of the matched entity, if any. |
|
|
|
Name of the matched entity, if any. |
|
|
|
Score of the best match; on the embedding-similarity path, the average of both scores when the match type is |
|
|
|
|
DuplicateCandidate
Fields and defaults:
| Field | Type | Default | Description |
|---|---|---|---|
|
|
|
ID of the potential duplicate. |
|
|
|
Name of the potential duplicate. |
|
|
|
Canonical name, if set. |
|
|
|
Entity type. |
|
|
|
Embedding similarity score. |
|
|
|
Fuzzy name-match score, if computed. |
|
|
|
Status of the |
DeduplicationStats
Fields and defaults:
| Field | Type | Default | Description |
|---|---|---|---|
|
|
|
Total number of entities. |
|
|
|
Number of entities merged into another. |
|
|
|
Number of |
|
|
|
Number of |
Provenance and extraction records (bolt)
register_extractor
Register an extractor for provenance tracking.
async def register_extractor(
name: str,
*,
version: str | None=None,
config: dict[str, Any] | None=None,
) -> dict[str, Any]: ...
| Parameter | Description |
|---|---|
|
Unique extractor name (e.g., "gliner2", "SpacyNER") |
|
Optional version string |
|
Optional configuration dict (will be JSON serialized) |
link_entity_to_message
Link an entity to its source message (:EXTRACTED_FROM relationship).
async def link_entity_to_message(
entity: Entity | UUID,
message_id: UUID | str,
*,
confidence: float=1.0,
start_pos: int | None=None,
end_pos: int | None=None,
context: str | None=None,
) -> bool: ...
| Parameter | Description |
|---|---|
|
Entity or entity ID |
|
ID of the source message |
|
Extraction confidence score |
|
Start character position in the message |
|
End character position in the message |
|
Surrounding text context |
link_entity_to_extractor
Link an entity to its extractor (:EXTRACTED_BY relationship).
async def link_entity_to_extractor(
entity: Entity | UUID,
extractor_name: str,
*,
confidence: float=1.0,
extraction_time_ms: float | None=None,
) -> bool: ...
| Parameter | Description |
|---|---|
|
Entity or entity ID |
|
Name of the extractor |
|
Extraction confidence score |
|
Time taken for extraction in milliseconds |
get_entity_provenance
Get provenance information for an entity.
async def get_entity_provenance(entity_id: Entity | UUID | str) -> dict[str, Any]: ...
| Parameter | Description |
|---|---|
|
Entity, entity UUID, or entity id string. ( |
get_entities_from_message
Get all entities extracted from a message.
async def get_entities_from_message(message_id: UUID | str) -> list[tuple[Entity, dict[str, Any]]]: ...
| Parameter | Description |
|---|---|
|
ID of the source message |
get_entities_by_extractor
Get all entities extracted by a specific extractor.
async def get_entities_by_extractor(extractor_name: str, *, limit: int=100) -> list[tuple[Entity, dict[str, Any]]]: ...
| Parameter | Description |
|---|---|
|
Name of the extractor |
|
Maximum number of entities to return |
list_extractors
List all registered extractors with entity counts.
async def list_extractors() -> list[dict[str, Any]]: ...
get_extraction_stats
Get overall extraction statistics.
async def get_extraction_stats() -> dict[str, Any]: ...