Apply a custom entity extraction schema

Runs client-side. The extractors on this page run in your Python process and do not call a memory backend, so they work next to either backend. Configuring them as a MemoryClient extraction pipeline applies only to Bolt: NAMS extracts entities server-side and ignores the client’s extraction settings with a UserWarning. See the backend capabilities reference.

To identify customer and product mentions, provide domain descriptions and an explicit mapping from extraction labels to persisted entity types. Test the mapping against a representative sentence before writing graph data.

For the rationale, see Understanding the extraction pipeline.

1. Install the local extractor

Use Python 3.10 or newer and a POSIX shell. Prepare the example files and install the required extra:

Create a local folder and virtual environment for the examples on this page. The commands reuse an existing environment without changing its files:

mkdir -p ~/agent-memory-tutorials
cd ~/agent-memory-tutorials
if [ -e .venv ]; then
  printf '%s\n' 'Using the existing virtual environment.'
else
  python3 -m venv .venv
fi
source .venv/bin/activate

Expected: ~/agent-memory-tutorials is your working directory and its virtual environment is active. Install the published SDK with the command below.

Each complete code block labelled Save as names a file to create in this folder using your editor. Copy the entire block, including imports and the entry point. Expand each helper disclosure and use Copy code to copy its full source. Keep all files together so their imports resolve.

When continuing from another tutorial or guide, retain the existing environment, configuration, session files, and .tutorial-state/. Reuse unchanged helper files; compare an existing file before replacing it, and finish any pending cleanup or recovery before changing the code that owns its state.

python -m pip install 'neo4j-agent-memory[gliner2]==0.7.0'

The maintained program extraction_recipes.py supplies its own imports, source texts, and event-loop entry point. It downloads the GLiNER2.5 checkpoint fastino/gliner2.5-base-v1 (about 407 MB) on first inference; allow disk space and network access for the download. These commands do not connect to a database or require an LLM API key.

Save the complete program below in agent-memory-tutorials/. It has no local helper imports. The next section explains the selected function.

Complete extraction_recipes.py
Save as extraction_recipes.py
"""Complete local-model extraction recipes without a database or LLM API."""

import argparse
import asyncio

from neo4j_agent_memory.extraction.domain_schemas import DomainSchema, get_schema
from neo4j_agent_memory.extraction.gliner2_extractor import GLiNER2Extractor
from neo4j_agent_memory.extraction.pipeline import ExtractionPipeline, MergeStrategy
from neo4j_agent_memory.extraction.streaming import StreamingExtractor
from neo4j_agent_memory.ontology import RelationshipDef

TEXTS = [
    "Maya Chen works at Northstar Robotics in Denver.",
    "Ravi Shah joined Summit Research in Boulder.",
]


# tag::schema[]
def custom_extractor():
    schema = DomainSchema(
        name="support_catalog",
        entity_types={"customer": "A named customer", "product": "A named purchased item"},
    )
    return GLiNER2Extractor(
        ontology=schema,
        label_mapping={"customer": ("PERSON", None), "product": ("OBJECT", "PRODUCT")},
        threshold=0.5,
    )


# end::schema[]


# tag::extract[]
async def extract(selected, text):
    result = await selected.extract(text, extract_relations=False, extract_preferences=False)
    if not result.entities:
        raise RuntimeError("No candidate entities; inspect the input/schema/model threshold")
    for entity in result.entities:
        print(entity.name, entity.type, entity.subtype, entity.confidence)
    print(f"Verified: {result.entity_count} candidate entities returned; inspect their accuracy")
    return result


# end::extract[]


# tag::relations[]
def relation_extractor():
    # The business catalog declares labels only; attach typed relationships.
    ontology = get_schema("business").to_ontology(
        relationships=[
            RelationshipDef(
                type="EMPLOYED_BY",
                source="person",
                target="company",
                description="The person works for or has joined this company",
            ),
            RelationshipDef(
                type="LOCATED_IN",
                source="company",
                target="location",
                description="The company is based or operates in this place",
            ),
        ]
    )
    return GLiNER2Extractor.for_ontology(ontology, threshold=0.5)


async def relations(selected, text):
    result = await selected.extract(text, extract_preferences=False)
    if not result.relations:
        raise RuntimeError("No relations; check the declared relationships and thresholds")
    for relation in result.relations:
        print(relation.source, relation.relation_type, relation.target, relation.confidence)
    print(f"Verified: {len(result.relations)} candidate relations returned; inspect them")
    return result


# end::relations[]


# tag::batch[]
async def batch(selected, texts):
    # Propagate stage errors so the batch can distinguish a failed item from an empty extraction.
    pipeline = ExtractionPipeline(
        stages=[selected], merge_strategy=MergeStrategy.CONFIDENCE, fallback_on_error=False
    )
    result = await pipeline.extract_batch(
        texts,
        batch_size=2,
        max_concurrency=1,
        fail_fast=False,
        extract_relations=False,
        extract_preferences=False,
        on_progress=lambda done, total: print(f"Progress: {done}/{total}"),
    )
    assert result.total_items == len(texts)
    for item in result.results:
        if item.success:
            print(f"Input {item.index}: {item.result.entity_count} entities")
        else:
            print(f"Input {item.index} failed: {item.error}")
    if result.failed_items:
        raise RuntimeError(
            f"Retry failed source indexes after fixing errors: {result.get_errors()}"
        )
    print(f"Verified: all {result.total_items} inputs accounted for")
    return result


# end::batch[]


# tag::streaming[]
async def streaming(selected, text):
    streamer = StreamingExtractor(selected, chunk_size=400, overlap=40, chunk_by_tokens=False)
    result = await streamer.extract(text, extract_relations=False)
    errors = [
        (chunk.chunk.index, chunk.error) for chunk in result.chunk_results if not chunk.success
    ]
    if errors:
        raise RuntimeError(f"Chunk extraction failed: {errors}")
    assert result.chunk_results
    combined = result.to_extraction_result(source_text=text)
    print(
        f"Verified: {len(result.chunk_results)} chunks completed; {combined.entity_count} merged entities"
    )
    return result


# end::streaming[]


async def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("command", choices=["extract", "schema", "relations", "batch", "streaming"])
    args = parser.parse_args()
    if args.command == "schema":
        await extract(custom_extractor(), "Maya Chen bought a Trail Starter shoe.")
    elif args.command == "relations":
        await relations(relation_extractor(), TEXTS[0])
    else:
        selected = GLiNER2Extractor.for_schema("business")
        if args.command == "extract":
            await extract(selected, TEXTS[0])
        elif args.command == "batch":
            await batch(selected, TEXTS)
        else:
            await streaming(selected, "\n".join(TEXTS * 12))


if __name__ == "__main__":
    asyncio.run(main())

2. Configure and run the extraction

The custom schema names customer and product. Its mapping uses (TYPE, SUBTYPE) tuples: a customer maps to PERSON and a product to OBJECT/PRODUCT. Without label_mapping, each label goes through the default mapping table, where product already maps to OBJECT/PRODUCT but the unlisted customer falls back to OBJECT/CUSTOMER. for_schema is reserved for registered template names; a custom DomainSchema is passed to the constructor’s ontology= argument, which also accepts an OntologyDocument.

The following function is included from the complete program. Run the maintained file with the command below it.

def custom_extractor():
    schema = DomainSchema(
        name="support_catalog",
        entity_types={"customer": "A named customer", "product": "A named purchased item"},
    )
    return GLiNER2Extractor(
        ontology=schema,
        label_mapping={"customer": ("PERSON", None), "product": ("OBJECT", "PRODUCT")},
        threshold=0.5,
    )


async def extract(selected, text):
    result = await selected.extract(text, extract_relations=False, extract_preferences=False)
    if not result.entities:
        raise RuntimeError("No candidate entities; inspect the input/schema/model threshold")
    for entity in result.entities:
        print(entity.name, entity.type, entity.subtype, entity.confidence)
    print(f"Verified: {result.entity_count} candidate entities returned; inspect their accuracy")
    return result
python extraction_recipes.py schema

3. Verify the result

Expected: candidates from “Maya Chen bought a Trail Starter shoe,” with PERSON and/or OBJECT/PRODUCT types, then the verified candidate count. Model output can omit or misidentify a mention; inspect it before using the schema on a corpus.

The command raises on the documented failure conditions. Inspect candidate names, types, and confidence against source text before storing them. Counts alone do not measure extraction accuracy.

Select a registered schema when it fits

Registered names are poleo, podcast, news, scientific, business, entertainment, medical, and legal. The exact label inventories and descriptions are in built-in schemas. There are no registered financial or ecommerce names. Use a custom schema for those domains. GLiNER2Extractor.for_schema(name) loads the ontology template of that name, so poleo, podcast, and news also bring typed relationships; get_schema(name) returns the label catalog alone.

DomainSchema is a GLiNER2.5 label catalog: label descriptions and optional relation descriptions, with no endpoint typing. EntitySchemaConfig is the separate SDK schema configuration model, keyed on POLE+O type names; do not interchange their fields or pass unsupported examples fields. Both convert to an OntologyDocument through to_ontology(), which is what GLiNER2Extractor(ontology=…​) and ExtractorBuilder.with_schema compile. A converted EntitySchemaConfig keeps its subtypes as labels and expands its relation types over the declared endpoint types, but a type name outside POLE+O fails the ontology’s structural validation, and the extractor raises ValueError on its first extraction. See schema models and persistence before relying on validation.

Extract relationships and persist schema configuration

GLiNER2.5 decodes relations in the same pass as the entities; there is no separate relation model to install. A DomainSchema declares no endpoint typing, so the custom extractor above returns entities only. To decode relations, attach typed relationships when you convert the catalog to an ontology. Each RelationshipDef names a source and a target label that the catalog declares, and GLiNER2.5 enforces those endpoint types while decoding:

def relation_extractor():
    # The business catalog declares labels only; attach typed relationships.
    ontology = get_schema("business").to_ontology(
        relationships=[
            RelationshipDef(
                type="EMPLOYED_BY",
                source="person",
                target="company",
                description="The person works for or has joined this company",
            ),
            RelationshipDef(
                type="LOCATED_IN",
                source="company",
                target="location",
                description="The company is based or operates in this place",
            ),
        ]
    )
    return GLiNER2Extractor.for_ontology(ontology, threshold=0.5)


async def relations(selected, text):
    result = await selected.extract(text, extract_preferences=False)
    if not result.relations:
        raise RuntimeError("No relations; check the declared relationships and thresholds")
    for relation in result.relations:
        print(relation.source, relation.relation_type, relation.target, relation.confidence)
    print(f"Verified: {len(result.relations)} candidate relations returned; inspect them")
    return result
python extraction_recipes.py relations

Expected: candidate relations such as Maya Chen EMPLOYED_BY Northstar Robotics with confidence values, then Verified: …​ candidate relations returned; inspect them. No relations causes an explicit error. Relations are decoded only when the call and the extractor both keep extract_relations=True (the default) and the ontology declares at least one relationship type. Zero-shot relation recall is weaker than entity recall; tune each relationship’s threshold on representative text. The ontology’s constraint flags (unique_source, unique_target, acyclic, allow_self, and inverse) are covered in Drive extraction from an ontology.

Persisting an EntitySchemaConfig through the Bolt schema manager stores configuration; it does not automatically validate every existing graph node or deploy a hosted ontology. Select/load the intended configuration explicitly when constructing extraction. To share one typed schema with every Bolt client of a database, store and activate an ontology with client.ontology instead; clients adopt the active version when they next connect, as described in store the ontology in the database. See the persistence API and hosted ontology tasks for the separate contracts.

Write effective schema descriptions

GLiNER2.5 reads each description as an annotation guideline. Improve its zero-shot accuracy by tuning the description text passed to DomainSchema:

  1. Use descriptive names — "A real estate property" scores better than "property".

  2. Include synonyms in the description — "A company, startup, or business organization" catches more of the mentions a plain "company" label misses.

  3. Be domain-specific — tailor descriptions to your corpus’s vocabulary rather than reusing a generic label.

  4. Rule out near misses — a clause such as "Not the support engineer who answers" stops a label from absorbing nearby mentions.

  5. Test thresholds on a sample document before committing to one — a lower threshold catches more entities but can reduce precision.