Tune entity resolution
|
Available on NAMS: No. The procedure on this page needs the Bolt backend. With the hosted NAMS backend its operations are unavailable: depending on the call, the client raises |
Message ingestion resolves extracted mentions against the entities already in the graph, so "Acme", "Acme Corp" and "ACME Corporation" converge on one node instead of three. This page covers the settings, how to measure whether a change helped, and how to turn resolution off.
|
Resolution on ingest is on by default since 0.7.0
( |
Prerequisites
-
A Bolt connection to Neo4j, such as the AuraDB instance from the first memory tutorial’s Aura setup.
-
neo4j-agent-memory0.7.0 with an embedding provider, and optionally thefuzzyextra for RapidFuzz scoring (the resolver falls back todifflibwithout it):python -m pip install 'neo4j-agent-memory[openai,fuzzy]==0.7.0' -
An extractor that produces entities on ingestion. Resolution only sees mentions that extraction produced.
How it works
OntologyResolver normalizes each mention, fetches candidates of the same type
in Cypher, scores them, sorts the best score into a band, and finally clusters
the mentions that matched nothing stored. Knowing which stage a bad decision
came from is most of the work of tuning;
The resolver’s five stages
explains each one and the scoring rules, including why a whole-token prefix
such as "Acme" and "Acme Bank" lands in the review band (0.88) unless the
surrounding context corroborates it.
What lands in the graph per band:
| Band | Written |
|---|---|
|
No new node. The message’s |
|
The node is created as usual, plus
|
|
A node, exactly as before 0.7.0. |
Every created node records the decision in its metadata blob
(resolution.action, resolution.score, resolution.match_type), alongside
the character offsets and the context window. That stored context is what lets
a later mention be compared on context rather than on name alone.
Configure the thresholds
Build a ResolutionConfig and pass it as resolution= to MemorySettings,
alongside your connection settings. The values below are the defaults:
from neo4j_agent_memory.config.settings import ResolutionConfig, ResolverStrategy
resolution = ResolutionConfig(
strategy=ResolverStrategy.COMPOSITE, # OntologyResolver on Bolt
resolve_on_ingest=True,
auto_merge_threshold=0.90,
review_threshold=0.85,
candidate_limit=12,
use_alias_gazetteer=True,
use_embedding_blocking=True,
context_window_chars=90,
scope="global",
fuzzy_threshold=0.85,
)
# MemorySettings(neo4j=..., resolution=resolution)
| Field | Default | Effect |
|---|---|---|
|
|
Resolve mentions while storing a message. The single opt-out. |
|
|
At or above this score the mention merges onto the match. Raise it if you see wrong merges; the review queue grows instead. |
|
|
At or above this score (but below auto-merge) the pair lands in the review
queue. Must be less than or equal to |
|
|
Blocking candidates per mention and per bucket. The exact-key query for a
message is capped at |
|
|
Use |
|
|
Also propose candidates from the entity vector index. Needs an embedder; with none configured the embedding component drops out and the remaining score weights are renormalized. |
|
|
Half-width of the mention’s context window, compared against a candidate’s stored context or description. |
|
|
|
|
|
Used by the |
strategy still selects the resolver: NONE, EXACT, FUZZY and SEMANTIC
build the simple resolvers, and COMPOSITE (the default) builds
OntologyResolver on Bolt. Only OntologyResolver resolves on the ingestion
path; any other strategy leaves ingestion as it behaved before 0.7.0.
Per-type overrides
Global thresholds are a blunt instrument: person names are far more ambiguous than product ids. Override them per entity type in the ontology, and the resolver uses them for any mention that maps onto that label:
entity_types:
- label: Customer
pole_type: PERSON
subtype: CUSTOMER
resolution_threshold: 0.95 # merge people only on strong evidence
review_threshold: 0.88
- label: Product
pole_type: OBJECT
subtype: PRODUCT
resolution_threshold: 0.88 # catalog names are distinctive; merge sooner
threshold on the same object is a different setting: it is the extractor’s
per-label confidence floor, not a resolution band.
Alias gazetteer
An aliases map is a controlled vocabulary, and it beats any similarity
computation: it produces a 1.0 match, it works with embeddings switched off,
and each member is indexed under both its raw and normalized forms, so a stored
"Acme Corp" and a mention "Acme" land in the same group.
entity_types:
- label: Organization
pole_type: ORGANIZATION
aliases:
Acme Corp: ["Acme", "Acme Corporation", "ACME", "Acme Corp."]
Northwind Ltd: ["NWL", "Northwind"]
This is the fix for anything the scorer cannot be expected to infer: internal project code names, ticker symbols, product SKUs, or an acronym whose expansion never co-occurs with it in your text.
The gazetteer is keyed per canonical name, not per surface form. If two
canonicals declare the same short form ({"Apple Inc": ["Apple"],
"Apple Records": ["Apple"]}), then "Apple" identifies neither of them, and the
pair falls through to the blend instead of scoring 1.0.
Review the middle band
The review band is the point of having three outcomes instead of two. Work the
queue with the same calls used for pairs that add_entity flags:
pairs = await client.long_term.find_potential_duplicates(limit=50)
for entity1, entity2, confidence in pairs:
print(f"{entity1.name} ({entity1.type}) <-> {entity2.name} {confidence:.2%}")
print(f" {entity1.description or '(no description)'}")
print(f" {entity2.description or '(no description)'}")
if same_entity(entity1, entity2): # your judgement, or a rule
await client.long_term.review_duplicate(entity1.id, entity2.id, confirm=True)
else:
await client.long_term.review_duplicate(entity1.id, entity2.id, confirm=False)
Each pair appears once, highest confidence first, with the newly flagged entity
first. confirm=True merges the pair: the first entity’s name becomes an alias
of the second, its relationships and provenance are copied onto the second, and
the SAME_AS edge is marked confirmed. confirm=False marks the edge
rejected and leaves both nodes alone. A rejected pair is a signal, not just a
chore: a steady stream of them at high confidence means auto_merge_threshold
is about to start making the same mistake silently.
get_deduplication_stats() reports the queue depth and merge counts:
stats = await client.long_term.get_deduplication_stats()
print(stats.total_entities, stats.merged_entities, stats.pending_reviews)
Scope candidates to a tenant
scope="user" restricts blocking to entities the tenant has actually mentioned,
reached through
(:Conversation {user_identifier})-[:HAS_MESSAGE]->(:Message)-[:MENTIONS]->(:Entity).
Pass user_identifier= on the ingesting call for it to apply. All three buckets
honor it: the exact-key query, the token-prefix query and the entity vector
index. The vector index in particular is global, so the tenant filter is applied
after the index returns its candidate_limit nearest neighbors. A tenant
whose entities are a small slice of the graph may see fewer candidates from that
bucket; the other two still run.
resolution = ResolutionConfig(scope="user")
Use it when different tenants talk about different "Apollo"s and you would rather have two nodes than one wrong merge.
|
|
Turn it off
resolution = ResolutionConfig(resolve_on_ingest=False)
# MemorySettings(neo4j=..., resolution=resolution)
Ingestion then behaves as it did before 0.7.0: one node per distinct surface
form, no blocking queries, no SAME_AS edges from ingestion. Everything else
keeps working: long_term.add_entity still deduplicates through the resolver,
and the resolver is still available for explicit calls.
Reasons to switch it off: you deduplicate downstream in your own pipeline, you are bulk-loading a source you already trust, or you want the extra time per message back. Measured overhead is about +1.7 ms for a two-mention message, roughly 8% of that write path; the exact-key blocking query is one round trip per message and entity type, not per mention.
Resolution never gates a write. If a blocking query fails, the failure is logged and the mentions are stored unresolved.
Measure a change
Do not tune by eye. Two tools, depending on what you are changing.
Resolution cases in the evaluation harness
ResolutionCase scores predicted clusters against gold clusters with B-cubed
F1, the standard entity-resolution metric, which rewards both splitting
distinct entities and joining variants:
from neo4j_agent_memory.memory.eval import EvalSuite, ResolutionCase
suite = EvalSuite(
resolution=[
ResolutionCase(
mentions=[
("Acme", "ORGANIZATION"),
("Acme Corp", "ORGANIZATION"),
("ACME Corporation", "ORGANIZATION"),
("Apple Bank", "ORGANIZATION"),
],
gold_clusters={
"Acme": "acme",
"Acme Corp": "acme",
"ACME Corporation": "acme",
"Apple Bank": "apple-bank",
},
),
],
)
report = await client.eval.run(suite, dimensions=["resolution"])
print(report.resolution.score) # B-cubed F1
print(report.resolution.details[0]) # predicted vs gold clusters, P/R/F1
Run it before and after a threshold change and keep both numbers. The dimension
is skipped, and named in report.skipped, when no resolver is configured.
|
With |
See Evaluate memory quality for wiring a suite into CI.
Benchmark script
benchmarks/compare_extractors.py is a repository-only sweep tool. It reports
entity and relation precision, recall and F1, latency, throughput and ontology
endpoint-type violations per checkpoint. For the resolution side,
benchmarks/data/poleo_resolution_gold.json and core.metrics.bcubed use the
same normalization as the resolver. Run it from a repository checkout:
uv run python benchmarks/compare_extractors.py \
--models fastino/gliner2.5-base-v1 \
--ontology ontology.yaml \
--relation-thresholds 0.3,0.5 \
--output results.json
Known limits
- Type-scoped blocking misses cross-typed duplicates
-
A mention typed
ORGANIZATIONis never compared against an entity typedOBJECT, so an extractor that typed the same company both ways produces two nodes. Fix the typing (a sharper ontology description usually does it) rather than loosening the blocking. - Acronym expansion needs a candidate to score
-
The acronym rule is reliable once the pair is in the candidate pool, but "NWL" and "Northwind Ltd" share no normalized key and no head token, so exact and prefix blocking never propose the pair. It only surfaces when embedding blocking finds it (which needs an embedder, and short acronyms embed poorly) or when the ontology declares the alias. For acronyms that matter, declare them.
- Low-entropy names are penalized, deliberately
-
Short repetitive names ("N600", "Ada") lose 0.25, because fuzzy similarity is near-meaningless on them. The penalty is keyed on the lower-entropy side of the pair, so it applies whichever of the two is the mention. If those are your real identifiers, an alias entry or a per-type
resolution_thresholdis the right lever; the penalty itself is not configurable. - A shared alias surface form identifies nothing
-
The gazetteer is per canonical name. Declare
"Apple"under two canonicals and it stops being evidence for either; the pair falls through to the blend. Give each canonical a surface form only it uses. - The stored graph is global
-
See the warning under Scope candidates to a tenant.
See also
-
Drive extraction from an ontology — where
aliasesand the per-type thresholds live. -
Merge a reviewed entity pair — the
add_entitypath and the review workflow. -
Entity resolution and deduplication — why resolution matters and what it costs to get wrong.
-
Resolution configuration — every field and environment variable.