Adopt an existing domain graph

Available on NAMS: No. The procedure on this page needs the Bolt backend. With the hosted NAMS backend its operations are unavailable: depending on the call, the client raises NotSupportedError, AttributeError or TypeError, or ignores a Bolt-only setting or argument (NAMS manages embedding and extraction server-side). See the backend capabilities reference for what NAMS provides instead.

How to layer neo4j-agent-memory on top of a Neo4j graph that already exists in production, so that library writes (MENTIONS edges, relation writes, entity upserts) link to your existing nodes instead of creating duplicates.

By default, the library MERGEs entities on (:Entity {name, type}). If your existing graph has nodes labelled :Person, :Movie, :Client, and so on — none of which carry :Entity — those merges will create duplicates. The client.schema.adopt_existing_graph() helper attaches the :Entity super-label, the library’s id/type/name properties, and is idempotent.

Prerequisites

  • A running Neo4j 5.x with your existing domain graph already loaded.

  • The Python SDK plus the extra for your configured embedding provider, for example pip install 'neo4j-agent-memory[openai]==0.7.0', which provides client.schema.adopt_existing_graph.

  • A MemorySettings configured for the database (see Configuration reference), and an already-open async with MemoryClient(settings) as client: block — every snippet below shares that one connected client.

If your domain types differ from the POLE+O default ontology, also configure SchemaModel.CUSTOM:

from neo4j_agent_memory import MemorySettings
from neo4j_agent_memory.config.settings import SchemaConfig, SchemaModel

settings = MemorySettings(
    schema_config=SchemaConfig(
        model=SchemaModel.CUSTOM,
        entity_types=["PERSON", "MOVIE", "GENRE"],
        strict_types=True,
    ),
    # ...
)

Goal

Inspect the intended label/type mapping, preview changes without writes, then adopt the selected nodes and verify their identity.

Limitations

Limitations of the Bolt implementation:

  • Automatic extraction does not link to adopted nodes by default. The extractors map their labels through POLE+O, so a MOVIE mention arrives typed OBJECT (subtype MOVIE) and MERGEs a second :Entity:Object node. Ingest-time resolution only compares entities of the same type, so it does not merge that node into the adopted one. A GLiNER2.5 label_mapping that preserves MOVIE makes the MERGE find the adopted node, and the ingestion path links the MENTIONS edge to the id the MERGE returns. For a deterministic link, use extraction_mode="explicit" with EntityRef.

  • SchemaConfig.strict_types=True selects strict validation. When validation_mode is unset, it derives validation_mode="strict": add_entity raises ValidationError for an entity type that entity_types does not declare, and message ingestion drops extracted entities of such types. Declare every adopted type in entity_types.

  • Adopted ids must be UUID-shaped for the read helpers. Nodes adopted without a pre-existing id get a deterministic <label_lc>:<name> id; helpers that hydrate Entity.id as a UUID (such as search_entities()) raise ValueError on those rows. Read them with client.query.cypher.

  • Adoption writes no embedding. Semantic search such as search_entities() does not find adopted nodes until you backfill their embeddings. The retrieve.py script in the existing graph example shows a backfill that skips nodes with non-UUID ids.

Steps

1. Identify the labels and name properties to adopt

For each Neo4j label in your existing graph, decide:

  • The library entity type to assign (a string — POLE+O members like PERSON if you’re keeping the default ontology, or your own type names if you’re using SchemaModel.CUSTOM).

  • The property to use as the name of the resulting :Entity node. Defaults to name per label, but you can override per label (movies often use title, people sometimes use full_name, etc.).

2. Preview the full mapping

Pass dry_run=True to see what would change without mutating the graph:

report = await client.schema.adopt_existing_graph(
    label_to_type={"Person": "PERSON", "Movie": "MOVIE", "Genre": "GENRE"},
    name_property_per_label={"Movie": "title"},
    dry_run=True,
)
print(f"Would migrate {report.total_migrated} nodes.")

Review the counts and skipped nodes before applying the same mapping.

3. Apply the reviewed mapping

Run the same mapping without dry_run to write the changes:

report = await client.schema.adopt_existing_graph(
    label_to_type={
        "Person": "PERSON",
        "Movie": "MOVIE",
        "Genre": "GENRE",
    },
    name_property_per_label={"Movie": "title"},
)

print(f"Migrated {report.total_migrated} nodes "
      f"({report.total_already_adopted} already adopted, "
      f"{report.total_skipped} skipped).")
for label_report in report.by_label:
    print(
        f"  {label_report.label} -> {label_report.type}: "
        f"+{label_report.migrated_count} new, "
        f"={label_report.already_adopted_count} already, "
        f"~{label_report.skipped_count} skipped"
    )

The helper:

  1. Adds :Entity to every matching node that doesn’t already carry it.

  2. Sets type from the input map.

  3. Sets name from the configured property (defaulting to existing name).

  4. Generates a deterministic id of the form <label_lc>:<name> for nodes that don’t have one. Existing id properties are preserved.

  5. Skips nodes whose configured name property is null.

4. Re-run as needed

The helper is idempotent. Re-running attaches :Entity to any new nodes of the configured labels added since the last run, and reports the already-adopted count for the rest.

Verification

After adoption, library writes that MERGE on (:Entity {name, type}) land on your existing nodes. Name the entities explicitly so the write is deterministic — automatic NER extraction does not link to adopted nodes (see Limitations):

from neo4j_agent_memory.schema.models import EntityRef

# Add a message that names people and movies in the existing graph.
await client.short_term.add_message(
    "demo",
    "user",
    "Have you seen Inception? Bob Singh recommended it.",
    extraction_mode="explicit",
    explicit_mentions=[
        EntityRef(name="Inception", type="MOVIE"),
        EntityRef(name="Bob Singh", type="PERSON"),
    ],
)

# Count every node the library could have created, across *all* labels —
# a duplicate shows up as a second label set such as ["Entity", "Object"].
# `client.query.cypher` is the portable read accessor (it works on NAMS
# too); `client.graph` remains for write Cypher on bolt.
rows = await client.query.cypher(
    """
    UNWIND ['Bob Singh', 'Inception'] AS target
    MATCH (n) WHERE n.name = target
    RETURN target, count(n) AS total, collect(DISTINCT labels(n)) AS label_sets
    ORDER BY target
    """
)
for row in rows:
    print(f"{row['target']}: {row['total']} (expect 1) {row['label_sets']}")

Edge cases

Case Behavior

Node missing the configured name property

Skipped. Reported in AdoptionLabelReport.skipped_count.

Node already carries :Entity from a previous run

No-op. Reported in AdoptionLabelReport.already_adopted_count.

Node has an id property already

Preserved. The helper only generates an id when one is absent.

Label or name property contains characters outside [A-Za-z_][A-Za-z0-9_]*

SchemaError raised before any mutation. The helper refuses to interpolate unsafe identifiers into Cypher.

See also