Adopt an existing domain graph
|
Available on NAMS: No. The procedure on this page needs the Bolt backend. With the hosted NAMS backend its operations are unavailable: depending on the call, the client raises |
How to layer neo4j-agent-memory on top of a Neo4j graph that already
exists in production, so that library writes (MENTIONS edges, relation
writes, entity upserts) link to your existing nodes instead of creating
duplicates.
By default, the library MERGEs entities on (:Entity {name, type}). If
your existing graph has nodes labelled :Person, :Movie, :Client,
and so on — none of which carry :Entity — those merges will create
duplicates. The client.schema.adopt_existing_graph() helper attaches
the :Entity super-label, the library’s id/type/name properties,
and is idempotent.
Prerequisites
-
A running Neo4j 5.x with your existing domain graph already loaded.
-
The Python SDK plus the extra for your configured embedding provider, for example
pip install 'neo4j-agent-memory[openai]==0.7.0', which providesclient.schema.adopt_existing_graph. -
A
MemorySettingsconfigured for the database (see Configuration reference), and an already-openasync with MemoryClient(settings) as client:block — every snippet below shares that one connectedclient.
If your domain types differ from the POLE+O default ontology, also configure SchemaModel.CUSTOM:
from neo4j_agent_memory import MemorySettings
from neo4j_agent_memory.config.settings import SchemaConfig, SchemaModel
settings = MemorySettings(
schema_config=SchemaConfig(
model=SchemaModel.CUSTOM,
entity_types=["PERSON", "MOVIE", "GENRE"],
strict_types=True,
),
# ...
)
Goal
Inspect the intended label/type mapping, preview changes without writes, then adopt the selected nodes and verify their identity.
Limitations
Limitations of the Bolt implementation:
-
Automatic extraction does not link to adopted nodes by default. The extractors map their labels through POLE+O, so a
MOVIEmention arrives typedOBJECT(subtypeMOVIE) and MERGEs a second:Entity:Objectnode. Ingest-time resolution only compares entities of the sametype, so it does not merge that node into the adopted one. A GLiNER2.5label_mappingthat preservesMOVIEmakes the MERGE find the adopted node, and the ingestion path links theMENTIONSedge to the id the MERGE returns. For a deterministic link, useextraction_mode="explicit"withEntityRef. -
SchemaConfig.strict_types=Trueselects strict validation. Whenvalidation_modeis unset, it derivesvalidation_mode="strict":add_entityraisesValidationErrorfor an entity type thatentity_typesdoes not declare, and message ingestion drops extracted entities of such types. Declare every adopted type inentity_types. -
Adopted ids must be UUID-shaped for the read helpers. Nodes adopted without a pre-existing
idget a deterministic<label_lc>:<name>id; helpers that hydrateEntity.idas aUUID(such assearch_entities()) raiseValueErroron those rows. Read them withclient.query.cypher. -
Adoption writes no embedding. Semantic search such as
search_entities()does not find adopted nodes until you backfill their embeddings. Theretrieve.pyscript in the existing graph example shows a backfill that skips nodes with non-UUID ids.
Steps
1. Identify the labels and name properties to adopt
For each Neo4j label in your existing graph, decide:
-
The library entity type to assign (a string — POLE+O members like
PERSONif you’re keeping the default ontology, or your own type names if you’re usingSchemaModel.CUSTOM). -
The property to use as the
nameof the resulting:Entitynode. Defaults tonameper label, but you can override per label (movies often usetitle, people sometimes usefull_name, etc.).
2. Preview the full mapping
report = await client.schema.adopt_existing_graph(
label_to_type={"Person": "PERSON", "Movie": "MOVIE", "Genre": "GENRE"},
name_property_per_label={"Movie": "title"},
dry_run=True,
)
print(f"Would migrate {report.total_migrated} nodes.")
Review the counts and skipped nodes before applying the same mapping.
3. Apply the reviewed mapping
report = await client.schema.adopt_existing_graph(
label_to_type={
"Person": "PERSON",
"Movie": "MOVIE",
"Genre": "GENRE",
},
name_property_per_label={"Movie": "title"},
)
print(f"Migrated {report.total_migrated} nodes "
f"({report.total_already_adopted} already adopted, "
f"{report.total_skipped} skipped).")
for label_report in report.by_label:
print(
f" {label_report.label} -> {label_report.type}: "
f"+{label_report.migrated_count} new, "
f"={label_report.already_adopted_count} already, "
f"~{label_report.skipped_count} skipped"
)
The helper:
-
Adds
:Entityto every matching node that doesn’t already carry it. -
Sets
typefrom the input map. -
Sets
namefrom the configured property (defaulting to existingname). -
Generates a deterministic
idof the form<label_lc>:<name>for nodes that don’t have one. Existingidproperties are preserved. -
Skips nodes whose configured name property is null.
Verification
After adoption, library writes that MERGE on (:Entity {name, type})
land on your existing nodes. Name the entities explicitly so the write is
deterministic — automatic NER extraction does not link to adopted nodes
(see Limitations):
from neo4j_agent_memory.schema.models import EntityRef
# Add a message that names people and movies in the existing graph.
await client.short_term.add_message(
"demo",
"user",
"Have you seen Inception? Bob Singh recommended it.",
extraction_mode="explicit",
explicit_mentions=[
EntityRef(name="Inception", type="MOVIE"),
EntityRef(name="Bob Singh", type="PERSON"),
],
)
# Count every node the library could have created, across *all* labels —
# a duplicate shows up as a second label set such as ["Entity", "Object"].
# `client.query.cypher` is the portable read accessor (it works on NAMS
# too); `client.graph` remains for write Cypher on bolt.
rows = await client.query.cypher(
"""
UNWIND ['Bob Singh', 'Inception'] AS target
MATCH (n) WHERE n.name = target
RETURN target, count(n) AS total, collect(DISTINCT labels(n)) AS label_sets
ORDER BY target
"""
)
for row in rows:
print(f"{row['target']}: {row['total']} (expect 1) {row['label_sets']}")
Edge cases
| Case | Behavior |
|---|---|
Node missing the configured name property |
Skipped. Reported in |
Node already carries |
No-op. Reported in |
Node has an |
Preserved. The helper only generates an |
Label or name property contains characters outside |
|
See also
-
The POLE+O data model — the default ontology you’re either keeping or overriding with
SchemaModel.CUSTOM. -
Existing graph example — a runnable end-to-end example using a small Movies-style domain.