Drive extraction from an ontology

Available on NAMS: No. The procedure on this page needs the Bolt backend. With the hosted NAMS backend its operations are unavailable: depending on the call, the client raises NotSupportedError, AttributeError or TypeError, or ignores a Bolt-only setting or argument (NAMS manages embedding and extraction server-side). See the backend capabilities reference for what NAMS provides instead.

To make extraction produce a predictable, typed graph instead of a pile of loosely labeled nodes, write an ontology and hand it to the client. One document drives all of it: the GLiNER2.5 JointIE schema, the LLM extractor’s prompt, relation validation on ingest, entity resolution, and the strict write paths.

Ontology management (client.ontology) works on both backends. This page is about the bolt backend, where extraction runs in your process and the ontology shapes it locally. On NAMS, extraction runs server-side against the activated version — see Use ontologies.

Prerequisites

  • Install the local extractor, the CLI and YAML support. Quote package extras in shell commands:

    python -m pip install 'neo4j-agent-memory[gliner2,openai,cli]==0.7.0' pyyaml

    The first extraction downloads the GLiNER2.5 checkpoint fastino/gliner2.5-base-v1 (about 407 MB).

  • Use a dedicated AuraDB instance and export NEO4J_URI, NEO4J_USERNAME, and NEO4J_PASSWORD as shown in the Aura connection setup. These examples use the bolt backend; NEO4J_DATABASE defaults to neo4j when omitted.

  • Set OPENAI_API_KEY for the default embedding provider, or configure another one as shown in Configure an embedding provider.

  • Run snippets containing await in an async function. The connected client in later fragments comes from step 3.

1. Write the ontology

An ontology is a YAML (or JSON) document with three parts: a domain, the entity_types, and the relationships between them. Save this one as ontology.yaml in your working directory:

domain:
  id: support-desk
  name: support-desk
  description: Customer support incidents, the people involved, and the products they affect.

entity_types:
  - label: Customer
    pole_type: PERSON
    subtype: CUSTOMER
    description: >-
      A named individual who reports a problem or asks for help. Not the
      support engineer who answers, and not the company the person works for.
    aliases:
      Dana Whitfield: ["Dana W.", "D. Whitfield"]

  - label: Engineer
    pole_type: PERSON
    subtype: ENGINEER
    description: >-
      A named support or engineering staff member who investigates or resolves
      an incident. Not the customer, and not a job title on its own
      ("the on-call engineer" is not an Engineer).

  - label: Organization
    pole_type: ORGANIZATION
    subtype: COMPANY
    description: >-
      A named company or institution. Not a product it sells, and not an
      industry or market segment.
    aliases:
      Acme Corp: ["Acme", "Acme Corporation", "ACME"]

  - label: Product
    pole_type: OBJECT
    subtype: PRODUCT
    description: >-
      A named software product or service the company ships. Not a component
      inside it, and not a generic category ("the database").
    threshold: 0.45

  - label: Component
    pole_type: OBJECT
    subtype: COMPONENT
    description: >-
      A named subsystem or module of a product — a service, library or
      endpoint. Not the product as a whole.

  - label: Incident
    pole_type: EVENT
    subtype: INCIDENT
    description: >-
      A specific reported failure or outage, identified by a ticket id or a
      distinct description. Not a recurring class of problem, and not a bare
      date.

relationships:
  - type: EMPLOYED_BY
    source: Customer
    target: Organization
    description: The customer is paid to work for this organization. Their current employer only.
    unique_source: true
    threshold: 0.35

  - type: REPORTED
    source: Customer
    target: Incident
    description: The customer raised this incident.

  - type: USES
    source: Customer
    target: Product
    description: The customer runs or has deployed this product.

  - type: AFFECTS
    source: Incident
    target: Product
    description: The incident degrades or breaks this product.

  - type: AFFECTS
    source: Incident
    target: Component
    description: The incident degrades or breaks this component.

  - type: PART_OF
    source: Component
    target: Product
    description: The component ships as part of this product.
    acyclic: true

Three things about this document are load-bearing:

description is an annotation guideline, not documentation

GLiNER2.5 serializes each description into the encoder input, so it behaves the way instructions to a human annotator do. Write the positive case in one sentence and then rule things out — the negative cases ("Not the support engineer who answers") are what stop a label from absorbing everything nearby. The same text is rendered into the LLM extractor’s prompt.

source / target are enforced during decoding

JointIE checks endpoint types inside its beam search rather than filtering afterwards, so a REPORTED edge from an Organization is never produced in the first place. Several entries may share one type (AFFECTS above): the compiler groups them into a single relation with head/tail lists.

Constraints shape the decoder, not just the storage layer

unique_source: true becomes unique_head (at most one EMPLOYED_BY per customer), unique_target: true becomes unique_tail, and acyclic: true forbids cycles. Self-loops are forbidden per relation unless a relationship sets allow_self: true; the document-level no_self_loops (default true) is applied relation by relation so an explicit allow_self still wins. To mirror a relationship, use inverse: OTHER_TYPE and declare the reverse pattern too (validate_structure() checks that the pair actually mirrors). Never reach for a "symmetric" flag — there is none, because it is broken upstream.

How labels reach Neo4j

Each entity type carries a POLE+O pole_type and an optional subtype, and the storage layer writes both onto the node as the type and subtype properties and as labels: the pole_type as its PascalCase label, a built-in POLE+O subtype (such as INDIVIDUAL) as its label, and the ontology’s own label for that exact pole_type/subtype pair. So Customer above lands as:

(:Entity:Person:Customer {name: "Dana Whitfield", type: "PERSON", subtype: "CUSTOMER"})

The declared label is written in PascalCase that keeps its own capitals: SupportCase stays SupportCase, and a GLiNER-style tv_show becomes TvShow. Declare the subtype. Without it, Customer and Engineer both map to (PERSON, None); two labels declaring one pair is ambiguous, so neither label is written and the nodes are plain :Entity:Person. The fine-grained distinction then survives in the extractor’s output but not in the graph.

2. Validate and compile it

Check the document before wiring it in. validate reports every structural problem at once (duplicate labels, an undeclared relationship endpoint, a pole_type outside POLE+O, an inverse that does not mirror) and exits 1 if it finds any:

neo4j-agent-memory ontology validate ontology.yaml
✓ Ontology 'support-desk' is valid.
  Entity types: Customer, Engineer, Organization, Product, Component, Incident
  Relationship types: EMPLOYED_BY, REPORTED, USES, AFFECTS, PART_OF
  Relationship patterns: 6

compile goes one step further and prints the GLiNER2.5 JointIE schema the model will actually decode against, which is the fastest way to see how your entries grouped:

neo4j-agent-memory ontology compile ontology.yaml

compile needs the gliner2 extra from the prerequisites; validate does not. Both commands accept a legacy EntitySchemaConfig file too — load_ontology() detects the shape and converts it on the way in.

3. Point the client at it

Two equivalent ways in, depending on whether the path belongs in your configuration or in your code:

import os

from neo4j_agent_memory import MemoryClient, MemorySettings
from neo4j_agent_memory.config.settings import SchemaConfig

settings = MemorySettings(
    backend="bolt",
    neo4j={
        "uri": os.environ["NEO4J_URI"],
        "username": os.environ["NEO4J_USERNAME"],
        "password": os.environ["NEO4J_PASSWORD"],
        "database": os.getenv("NEO4J_DATABASE", "neo4j"),
    },
    schema_config=SchemaConfig(ontology_path="ontology.yaml"),
)

async with MemoryClient(settings) as client:
    doc = client.ontology_document
    print(doc.domain.name, doc.labels())
from neo4j_agent_memory import MemoryClient, load_ontology

doc = load_ontology("ontology.yaml")

async with MemoryClient(settings, ontology=doc) as client:
    ...

The ontology= keyword accepts an OntologyDocument or a path, and it wins over everything in SchemaConfig. The full precedence — resolved once per connection, before the extractor is built — is:

  1. MemoryClient(ontology=…​)

  2. schema_config.ontology_path

  3. schema_config.custom_schema_path

  4. the active stored :OntologyVersion, when schema_config.use_active_ontology is True and one is bound

  5. schema_config.model=SchemaModel.CUSTOM plus entity_types (an ad-hoc document, one label per name)

  6. the built-in template named by schema_config.ontology_template (default poleo)

One setting overrides the list: extraction.gliner_schema, when the client builds a gliner or pipeline extractor itself. The extractor decodes against that template, so the client resolves to it too; otherwise ingest-time validation would drop every relation the template declares.

client.ontology_document reports what was resolved and client.validation_mode reports how strictly it is enforced. See Graph schema configuration for every field.

4. Store it in the database instead

A file is fine for one process. To share one ontology across every client that connects to a database, store and activate it — after which use_active_ontology (on by default) picks it up with no configuration at all:

from neo4j_agent_memory import MemoryClient, load_ontology

doc = load_ontology("ontology.yaml")

async with MemoryClient(settings) as client:
    version = await client.ontology.create("support-desk", doc, validation_mode="permissive")
    await client.ontology.activate(version.id)      # exactly one active version

The effective ontology is resolved at connect time. Activating a version does not re-wire the client that activated it — reconnect (or start the next process) for it to take effect. And because a file beats the stored version in the precedence list, leave ontology_path unset on clients that should follow the database.

Later revisions are immutable and additive: update(ontology_id, doc) appends revision n+1, activate(version.id) swaps which one is bound, and diff(ontology_id, 1, 2) shows what changed. See Ontology API reference for the whole surface and the bolt/NAMS differences.

5. Choose permissive or strict

Mode Behavior

permissive (default)

Extracted entities whose type the ontology does not declare are stored anyway, mapped through the label map. Relations the ontology forbids are always dropped — in either mode — because storing an edge the schema forbids is never what you asked for.

strict

Extracted entities whose (type, subtype) the ontology does not declare are dropped with a warning during ingestion, and long_term.add_entity / add_relationship raise ValidationError for an undeclared type, an undeclared relationship name, or an endpoint pair the ontology does not permit.

import os

from neo4j_agent_memory import MemorySettings
from neo4j_agent_memory.config.settings import SchemaConfig

settings = MemorySettings(
    backend="bolt",
    neo4j={
        "uri": os.environ["NEO4J_URI"],
        "username": os.environ["NEO4J_USERNAME"],
        "password": os.environ["NEO4J_PASSWORD"],
        "database": os.getenv("NEO4J_DATABASE", "neo4j"),
    },
    schema_config=SchemaConfig(
        ontology_path="ontology.yaml",
        validation_mode="strict",
    ),
)

validation_mode=None (the default) defers to the active stored version’s mode, and then to strict if strict_types=True else permissive. Strict mode looks an entity up by its (type, subtype) pair. An undeclared subtype falls back to a base declaration of the same pole_type (one declared without a subtype), and a bare pole_type with no subtype resolves to the first label declared for it. The ontology above declares no base PERSON label, so an entity typed PERSON validates (as Customer), while PERSON:CONTRACTOR and any LOCATION are rejected.

Start permissive, watch the warnings to find what your ontology is missing, then tighten to strict once ingestion is quiet.

6. Verify the typed edges

Store a couple of messages and read the graph back. Relations are stored as RELATED_TO with the relationship name in r.type, and the merge key includes that name, so one pair of entities can carry several differently-typed edges:

rows = await client.query.cypher("""
    MATCH (s:Entity)-[r:RELATED_TO]->(t:Entity)
    RETURN s.name AS source, r.type AS type, t.name AS target,
           r.confidence AS confidence, r.support AS support,
           r.derived AS derived, r.extractor AS extractor,
           r.source_message_ids AS message_ids
    ORDER BY r.support DESC, r.confidence DESC
""")
for row in rows:
    print(f"{row['source']} -[{row['type']}]-> {row['target']} "
          f"support={row['support']} conf={row['confidence']:.2f}")
Property Meaning

type

The relationship name (EMPLOYED_BY). Canonical — read this one. relation_type is a mirror kept for one release.

support

How many times this exact (source, type, target) has been observed. An edge seen in two messages has support = 2.

confidence

The highest confidence any observation reported (it never decreases).

derived

true only while every observation was inferred rather than stated. One asserted sighting clears it permanently.

extractor

The stage that first wrote it (gliner2, llm, …).

source_message_ids

The messages it came from, capped at 25. evidence holds up to three sentence windows.

Entity labels are worth a second query — they are how you confirm that pole_type/subtype landed:

MATCH (e:Entity)
RETURN labels(e) AS labels, e.type AS type, e.subtype AS subtype, count(*) AS n
ORDER BY n DESC

An empty relation result with entities present almost always means the ontology declares no relationships. A bare DomainSchema is a label catalog with no endpoint typing, so relation decoding is skipped entirely. Of the eight built-in templates (GLiNER2Extractor.for_schema, --gliner-schema), poleo, podcast and news declare relationships; the other five do not until you attach some.

7. Watch for the two quiet failure modes

Both look like "the text had nothing in it" if you are not expecting them.

feasible=False — warns

The decoder could not satisfy the ontology’s hard constraints for this input and returned an empty assignment. This is not "no facts in this text". It usually means the constraints are too tight for real language — a unique_source on a relationship that genuinely repeats, or an acyclic flag on something that is not a hierarchy. The extractor raises a RuntimeWarning and logs a warning; relax the constraint the text keeps violating and re-run.

Input longer than about 400 words — silent

Relation recall collapses past roughly 400 words. Input above gliner_max_words (default 384) is windowed automatically through extract_long, and relations whose endpoints land in different windows are dropped with no warning at all. Chunk long documents yourself — ideally on paragraph or turn boundaries, so both ends of a relation stay in one chunk — rather than relying on the automatic window. Nothing in the result distinguishes "this document had no relations" from "its relations were split across windows", so check your own input length:

import re
import warnings

from neo4j_agent_memory.extraction import GLiNER2Extractor

extractor = GLiNER2Extractor.for_ontology(client.ontology_document)
text = "Dana Whitfield at Acme Corp reported that the Billing API is down."

# Turn the feasibility warning into an error so a mis-compiled ontology fails
# loudly rather than returning an empty result.
with warnings.catch_warnings():
    warnings.simplefilter("error", RuntimeWarning)
    result = await extractor.extract(text)

# The long-input case has no warning to catch - measure it. The extractor
# counts every punctuation mark as a word, so str.split() undercounts.
if len(re.findall(r"\w+(?:[-_]\w+)*|\S", text)) > extractor.max_words:
    print("windowed: relations spanning window boundaries will be missing")

Next steps