Tune entity resolution

Available on NAMS: No. The procedure on this page needs the Bolt backend. With the hosted NAMS backend its operations are unavailable: depending on the call, the client raises NotSupportedError, AttributeError or TypeError, or ignores a Bolt-only setting or argument (NAMS manages embedding and extraction server-side). See the backend capabilities reference for what NAMS provides instead.

Message ingestion resolves extracted mentions against the entities already in the graph, so "Acme", "Acme Corp" and "ACME Corporation" converge on one node instead of three. This page covers the settings, how to measure whether a change helped, and how to turn resolution off.

Resolution on ingest is on by default since 0.7.0 (resolution.resolve_on_ingest=True). Before 0.7.0, ingestion stored one node per distinct surface form. One setting restores that behavior; see Turn it off.

Prerequisites

  • A Bolt connection to Neo4j, such as the AuraDB instance from the first memory tutorial’s Aura setup.

  • neo4j-agent-memory 0.7.0 with an embedding provider, and optionally the fuzzy extra for RapidFuzz scoring (the resolver falls back to difflib without it):

    python -m pip install 'neo4j-agent-memory[openai,fuzzy]==0.7.0'
  • An extractor that produces entities on ingestion. Resolution only sees mentions that extraction produced.

How it works

OntologyResolver normalizes each mention, fetches candidates of the same type in Cypher, scores them, sorts the best score into a band, and finally clusters the mentions that matched nothing stored. Knowing which stage a bad decision came from is most of the work of tuning; The resolver’s five stages explains each one and the scoring rules, including why a whole-token prefix such as "Acme" and "Acme Bank" lands in the review band (0.88) unless the surrounding context corroborates it.

What lands in the graph per band:

Band Written

merged

No new node. The message’s MENTIONS edge points at the matched entity and the mention’s surface form is appended to its aliases list.

review

The node is created as usual, plus (new)-[:SAME_AS {status: "pending", confidence, match_type}]→(matched), so find_potential_duplicates() returns the pair.

created

A node, exactly as before 0.7.0.

Every created node records the decision in its metadata blob (resolution.action, resolution.score, resolution.match_type), alongside the character offsets and the context window. That stored context is what lets a later mention be compared on context rather than on name alone.

Configure the thresholds

Build a ResolutionConfig and pass it as resolution= to MemorySettings, alongside your connection settings. The values below are the defaults:

from neo4j_agent_memory.config.settings import ResolutionConfig, ResolverStrategy

resolution = ResolutionConfig(
    strategy=ResolverStrategy.COMPOSITE,  # OntologyResolver on Bolt
    resolve_on_ingest=True,
    auto_merge_threshold=0.90,
    review_threshold=0.85,
    candidate_limit=12,
    use_alias_gazetteer=True,
    use_embedding_blocking=True,
    context_window_chars=90,
    scope="global",
    fuzzy_threshold=0.85,
)
# MemorySettings(neo4j=..., resolution=resolution)
Field Default Effect

resolve_on_ingest

True

Resolve mentions while storing a message. The single opt-out.

auto_merge_threshold

0.90

At or above this score the mention merges onto the match. Raise it if you see wrong merges; the review queue grows instead.

review_threshold

0.85

At or above this score (but below auto-merge) the pair lands in the review queue. Must be less than or equal to auto_merge_threshold; settings validation rejects the inversion rather than silently reordering.

candidate_limit

12

Blocking candidates per mention and per bucket. The exact-key query for a message is capped at candidate_limit × mentions, and never above 200 rows.

use_alias_gazetteer

True

Use EntityTypeDef.aliases as blocking keys and as a 1.0 match rule.

use_embedding_blocking

True

Also propose candidates from the entity vector index. Needs an embedder; with none configured the embedding component drops out and the remaining score weights are renormalized.

context_window_chars

90

Half-width of the mention’s context window, compared against a candidate’s stored context or description.

scope

global

user restricts candidates to entities the tenant has mentioned; see Scope candidates to a tenant.

fuzzy_threshold

0.85

Used by the FUZZY strategy and by DeduplicationConfig.from_resolution_config(). The COMPOSITE blend does not gate on it; the bands do the gating.

strategy still selects the resolver: NONE, EXACT, FUZZY and SEMANTIC build the simple resolvers, and COMPOSITE (the default) builds OntologyResolver on Bolt. Only OntologyResolver resolves on the ingestion path; any other strategy leaves ingestion as it behaved before 0.7.0.

Per-type overrides

Global thresholds are a blunt instrument: person names are far more ambiguous than product ids. Override them per entity type in the ontology, and the resolver uses them for any mention that maps onto that label:

entity_types:
  - label: Customer
    pole_type: PERSON
    subtype: CUSTOMER
    resolution_threshold: 0.95   # merge people only on strong evidence
    review_threshold: 0.88

  - label: Product
    pole_type: OBJECT
    subtype: PRODUCT
    resolution_threshold: 0.88   # catalog names are distinctive; merge sooner

threshold on the same object is a different setting: it is the extractor’s per-label confidence floor, not a resolution band.

Alias gazetteer

An aliases map is a controlled vocabulary, and it beats any similarity computation: it produces a 1.0 match, it works with embeddings switched off, and each member is indexed under both its raw and normalized forms, so a stored "Acme Corp" and a mention "Acme" land in the same group.

entity_types:
  - label: Organization
    pole_type: ORGANIZATION
    aliases:
      Acme Corp: ["Acme", "Acme Corporation", "ACME", "Acme Corp."]
      Northwind Ltd: ["NWL", "Northwind"]

This is the fix for anything the scorer cannot be expected to infer: internal project code names, ticker symbols, product SKUs, or an acronym whose expansion never co-occurs with it in your text.

The gazetteer is keyed per canonical name, not per surface form. If two canonicals declare the same short form ({"Apple Inc": ["Apple"], "Apple Records": ["Apple"]}), then "Apple" identifies neither of them, and the pair falls through to the blend instead of scoring 1.0.

Review the middle band

The review band is the point of having three outcomes instead of two. Work the queue with the same calls used for pairs that add_entity flags:

pairs = await client.long_term.find_potential_duplicates(limit=50)

for entity1, entity2, confidence in pairs:
    print(f"{entity1.name} ({entity1.type}) <-> {entity2.name}  {confidence:.2%}")
    print(f"  {entity1.description or '(no description)'}")
    print(f"  {entity2.description or '(no description)'}")

    if same_entity(entity1, entity2):      # your judgement, or a rule
        await client.long_term.review_duplicate(entity1.id, entity2.id, confirm=True)
    else:
        await client.long_term.review_duplicate(entity1.id, entity2.id, confirm=False)

Each pair appears once, highest confidence first, with the newly flagged entity first. confirm=True merges the pair: the first entity’s name becomes an alias of the second, its relationships and provenance are copied onto the second, and the SAME_AS edge is marked confirmed. confirm=False marks the edge rejected and leaves both nodes alone. A rejected pair is a signal, not just a chore: a steady stream of them at high confidence means auto_merge_threshold is about to start making the same mistake silently.

get_deduplication_stats() reports the queue depth and merge counts:

stats = await client.long_term.get_deduplication_stats()
print(stats.total_entities, stats.merged_entities, stats.pending_reviews)

Scope candidates to a tenant

scope="user" restricts blocking to entities the tenant has actually mentioned, reached through (:Conversation {user_identifier})-[:HAS_MESSAGE]->(:Message)-[:MENTIONS]->(:Entity). Pass user_identifier= on the ingesting call for it to apply. All three buckets honor it: the exact-key query, the token-prefix query and the entity vector index. The vector index in particular is global, so the tenant filter is applied after the index returns its candidate_limit nearest neighbors. A tenant whose entities are a small slice of the graph may see fewer candidates from that bucket; the other two still run.

resolution = ResolutionConfig(scope="user")

Use it when different tenants talk about different "Apollo"s and you would rather have two nodes than one wrong merge.

scope="user" narrows candidate generation, not storage. :Entity nodes are global and the create path merges on (name, type), so two tenants mentioning the identical surface form still share one node. Scoping stops fuzzy, embedding and alias matches from crossing tenants; it is not tenant isolation. For that, see scoped reads and writes.

Turn it off

resolution = ResolutionConfig(resolve_on_ingest=False)
# MemorySettings(neo4j=..., resolution=resolution)

Ingestion then behaves as it did before 0.7.0: one node per distinct surface form, no blocking queries, no SAME_AS edges from ingestion. Everything else keeps working: long_term.add_entity still deduplicates through the resolver, and the resolver is still available for explicit calls.

Reasons to switch it off: you deduplicate downstream in your own pipeline, you are bulk-loading a source you already trust, or you want the extra time per message back. Measured overhead is about +1.7 ms for a two-mention message, roughly 8% of that write path; the exact-key blocking query is one round trip per message and entity type, not per mention.

Resolution never gates a write. If a blocking query fails, the failure is logged and the mentions are stored unresolved.

Measure a change

Do not tune by eye. Two tools, depending on what you are changing.

Resolution cases in the evaluation harness

ResolutionCase scores predicted clusters against gold clusters with B-cubed F1, the standard entity-resolution metric, which rewards both splitting distinct entities and joining variants:

from neo4j_agent_memory.memory.eval import EvalSuite, ResolutionCase

suite = EvalSuite(
    resolution=[
        ResolutionCase(
            mentions=[
                ("Acme", "ORGANIZATION"),
                ("Acme Corp", "ORGANIZATION"),
                ("ACME Corporation", "ORGANIZATION"),
                ("Apple Bank", "ORGANIZATION"),
            ],
            gold_clusters={
                "Acme": "acme",
                "Acme Corp": "acme",
                "ACME Corporation": "acme",
                "Apple Bank": "apple-bank",
            },
        ),
    ],
)

report = await client.eval.run(suite, dimensions=["resolution"])
print(report.resolution.score)          # B-cubed F1
print(report.resolution.details[0])     # predicted vs gold clusters, P/R/F1

Run it before and after a threshold change and keep both numbers. The dimension is skipped, and named in report.skipped, when no resolver is configured.

With OntologyResolver, the harness hands each case’s mentions to resolve_episode() in one call, so the second pass clusters variants within the case even against an empty database, and each case’s detail reports clustered_by: "resolve_episode". Seed the stored entities a case should resolve onto when you want to measure matching against existing data, not just within the case. Other resolvers are scored through resolve_batch(), which resolves each mention independently against the stored graph; for them, seed the canonical entities first, or every mention resolves to itself.

See Evaluate memory quality for wiring a suite into CI.

Benchmark script

benchmarks/compare_extractors.py is a repository-only sweep tool. It reports entity and relation precision, recall and F1, latency, throughput and ontology endpoint-type violations per checkpoint. For the resolution side, benchmarks/data/poleo_resolution_gold.json and core.metrics.bcubed use the same normalization as the resolver. Run it from a repository checkout:

uv run python benchmarks/compare_extractors.py \
    --models fastino/gliner2.5-base-v1 \
    --ontology ontology.yaml \
    --relation-thresholds 0.3,0.5 \
    --output results.json

Known limits

Type-scoped blocking misses cross-typed duplicates

A mention typed ORGANIZATION is never compared against an entity typed OBJECT, so an extractor that typed the same company both ways produces two nodes. Fix the typing (a sharper ontology description usually does it) rather than loosening the blocking.

Acronym expansion needs a candidate to score

The acronym rule is reliable once the pair is in the candidate pool, but "NWL" and "Northwind Ltd" share no normalized key and no head token, so exact and prefix blocking never propose the pair. It only surfaces when embedding blocking finds it (which needs an embedder, and short acronyms embed poorly) or when the ontology declares the alias. For acronyms that matter, declare them.

Low-entropy names are penalized, deliberately

Short repetitive names ("N600", "Ada") lose 0.25, because fuzzy similarity is near-meaningless on them. The penalty is keyed on the lower-entropy side of the pair, so it applies whichever of the two is the mention. If those are your real identifiers, an alias entry or a per-type resolution_threshold is the right lever; the penalty itself is not configurable.

A shared alias surface form identifies nothing

The gazetteer is per canonical name. Declare "Apple" under two canonicals and it stops being evidence for either; the pair falls through to the blend. Give each canonical a surface form only it uses.

The stored graph is global

See the warning under Scope candidates to a tenant.

See also