Built-in extractor classes

The concrete single-stage extractors from Extractor classes reference: statistical (spaCy), joint entity and relation extraction (GLiNER2.5), and LLM-based. Import from neo4j_agent_memory.extraction.

SpacyEntityExtractor

Statistical NER using a separately installed spaCy model. Import from neo4j_agent_memory.extraction. Install the matching extraction extra and model before running inference.

def SpacyEntityExtractor(
    model: str='en_core_web_sm',
    type_mapping: dict[str, str] | None=None,
    subtype_mapping: dict[str, str] | None=None,
    default_confidence: float=0.85,
    context_window: int=50,
    *,
    drop_unmapped: bool=False,
): ...
Parameter Description

model

spaCy model name (e.g., "en_core_web_sm", "en_core_web_lg")

type_mapping

Custom mapping from spaCy labels to POLE+O types

subtype_mapping

Custom mapping from spaCy labels to subtypes

default_confidence

Default confidence score (spaCy doesn’t provide per-entity scores)

context_window

Number of characters of context to include around entities

drop_unmapped

Skip entities whose spaCy label type_mapping does not name, instead of typing them OBJECT

The default mappings include PERSON → PERSON, ORG → ORGANIZATION, GPE/LOC → LOCATION, EVENT → EVENT, and PRODUCT/WORK_OF_ART → OBJECT. Subtype mappings are retained on results. Override the mapping explicitly for your domain. spaCy extracts no relations.

When ExtractorBuilder or the configuration factory builds the spaCy stage against an ontology that declares subtypes, each spaCy label is remapped onto the first of these pairs the ontology declares: spaCy’s own subtype for the label (GPE → LOCATION:GEOPOLITICAL), the label itself as the subtype (LAW → OBJECT:LAW), then the bare POLE+O type. A label none of them matches is not extracted (drop_unmapped=True). spaCy never lands on a subtype it did not decide, so under the medical template it extracts nothing and under scientific a person is not made an author. An ontology with base POLE+O labels only keeps the defaults.

from neo4j_agent_memory.extraction import SpacyEntityExtractor

extractor = SpacyEntityExtractor(model="en_core_web_sm")
result = await extractor.extract("John works at Apple in California")

GLiNER2Extractor

Joint entity and relation extraction with a GLiNER2.5 checkpoint, locally and without an LLM call. Install the gliner2 extra, for example pip install 'neo4j-agent-memory[gliner2]==0.7.0', and allow the checkpoint to download on first use. The gliner extra is a deprecated alias of gliner2 in 0.7.

GLiNER2.5 decodes entities and relations in one pass through its JointIE engine. The ontology’s endpoint types are enforced during beam search rather than filtered afterwards, so every relation endpoint is one of the entities in the same result and relations carry mention ids. Preferences are never extracted; pair this extractor with LLMEntityExtractor when you need them. The checkpoint loads lazily when inference is first requested.

def GLiNER2Extractor(
    model: str=DEFAULT_GLINER2_5_MODEL,
    *,
    ontology: OntologyDocument | DomainSchema | EntitySchemaConfig | None=None,
    entity_labels: list[str] | dict[str, str] | None=None,
    threshold: float=0.5,
    relation_threshold: float | None=None,
    device: str='cpu',
    quantize: bool=False,
    compile: bool=False,
    max_words: int=384,
    chunk_overlap: int=64,
    overlap_policy: str | None=None,
    extract_relations: bool=True,
    extract_attributes: bool=False,
    attribute_tolerance: int=2,
    context_window: int=50,
    label_mapping: dict[str, tuple[str, str | None]] | None=None,
    joint_config: Any | None=None,
): ...
Parameter Description

model

GLiNER2.5 checkpoint id or local path. DEFAULT_GLINER2_5_MODEL is fastino/gliner2.5-base-v1. A GLiNER v1 id (urchade/…​, gliner-community/…​, numind/…​) raises ValueError before anything is downloaded.

ontology

The schema to extract against: an OntologyDocument, anything with a to_ontology() method (a DomainSchema, an EntitySchemaConfig), or None for the built-in POLE+O ontology.

entity_labels

Alternative to ontology: labels as a list, or as a {label: description} mapping. Produces a relationship-free ad-hoc ontology, so only entities are decoded. Ignored when ontology is given.

threshold

Entity confidence floor passed to the decoder.

relation_threshold

Confidence floor applied to decoded relations. None keeps what the ontology’s per-relation thresholds selected.

device

Inference device (cpu, cuda, mps).

quantize

Load the weights in fp16.

compile

Run the weights through torch.compile.

max_words

Longest input decoded in one pass; longer text is windowed. Must be greater than 0.

chunk_overlap

Word overlap between windows. Must be at least 0 and smaller than max_words; otherwise the constructor raises ValueError.

overlap_policy

Span-overlap policy for the attribute pass (flat, nested, allow, longest). None keeps the checkpoint’s default.

extract_relations

Set False to decode entities only.

extract_attributes

Run the opt-in second forward pass that recovers the ontology’s enum properties as span attributes.

attribute_tolerance

Character slack when joining the attribute pass back onto entity spans (the two passes tokenize independently).

context_window

Characters of surrounding text captured per entity.

label_mapping

Replacement for the label → (TYPE, SUBTYPE) table. By default the ontology’s own labels are used, over DEFAULT_LABEL_MAPPING.

joint_config

A ready-made JointIEConfig. When given it is used verbatim and threshold no longer applies.

Checkpoint Size on disk Use

fastino/gliner2.5-small-v1

~296 MB

Faster, CPU and edge

fastino/gliner2.5-base-v1

~407 MB

Default

fastino/gliner2.5-multi-v1

~594 MB

Multilingual

The extractor exposes name ("gliner2", recorded on every entity as ExtractedEntity.extractor), version (the installed gliner2 version, or "unknown"), model_id, ontology (the OntologyDocument it compiles), and the lazily loaded model and joint (JointIE engine) handles.

extract

Extract entities and relations from one text.

async def extract(
    text: str,
    *,
    entity_types: list[str] | None=None,
    extract_relations: bool=True,
    extract_preferences: bool=True,
) -> ExtractionResult: ...
Parameter Description

text

The text to extract from

entity_types

Restrict extraction to these ontology labels or POLE+O types (case-insensitive)

extract_relations

Set False to skip relation decoding for this call

extract_preferences

Ignored (GLiNER2.5 extracts no preferences)

Relations are decoded only when the call passes extract_relations=True, the extractor was built with extract_relations=True, and the ontology declares at least one relationship type. Relation types are normalized to UPPER_SNAKE. See Relation extractors.

Input longer than max_words (384 words, counted the way the gliner2 chunker counts them) is windowed through the engine’s extract_long with chunk_overlap words of overlap, and relations whose endpoints fall in different windows are lost. Relation recall degrades sharply past roughly 400 words, so prefer per-message extraction or explicit chunking with StreamingExtractor over one long concatenation.

extract_batch

Extract from several texts using the engine’s batched decoding.

async def extract_batch(
    texts: list[str],
    *,
    entity_types: list[str] | None=None,
    extract_relations: bool=True,
    batch_size: int=8,
    on_progress: Callable[[int, int], None] | None=None,
) -> list[ExtractionResult]: ...
Parameter Description

texts

The texts to extract from, in order

entity_types

Restrict extraction to these labels or POLE+O types

extract_relations

Set False to skip relation decoding for this call

batch_size

Texts decoded per forward pass

on_progress

Optional callback called after each slice. Receives (completed_count, total_count).

Returns one ExtractionResult per input text, in input order. Texts within max_words go through batched decoding; longer texts are windowed individually. There is no BatchExtractionResult wrapper and no max_concurrency; wrap the extractor in ExtractionPipeline for those.

for_schema

Build an extractor from a built-in domain template.

def for_schema(schema_name: str, **kwargs: Any) -> GLiNER2Extractor: ...
Parameter Description

schema_name

Template name (poleo, podcast, news, scientific, business, entertainment, medical, legal), resolved through neo4j_agent_memory.ontology.get_template. An unknown name raises ValueError.

**kwargs

Forwarded to the constructor, for example threshold or device

The extractor receives the template’s relationships as well as its labels: poleo, podcast and news declare relationships; the other five templates are entity-only. See Built-in domain schemas.

for_poleo

Build an extractor on the built-in POLE+O ontology (neo4j_agent_memory.ontology.POLEO_ONTOLOGY). Constructing GLiNER2Extractor() with no ontology or entity_labels gives the same ontology.

def for_poleo(**kwargs: Any) -> GLiNER2Extractor: ...

for_ontology

Build an extractor on an explicit ontology document.

def for_ontology(ontology: OntologyDocument, **kwargs: Any) -> GLiNER2Extractor: ...
from neo4j_agent_memory.extraction import GLiNER2Extractor

extractor = GLiNER2Extractor.for_schema("podcast", threshold=0.6)
result = await extractor.extract("Brian Chesky founded Airbnb in San Francisco.")
for entity in result.entities:
    print(entity.name, entity.full_type, entity.id)
for relation in result.relations:
    print(relation.source, relation.relation_type, relation.target)

Warnings and errors

Condition Behavior

The decoder returns feasible=False

A RuntimeWarning and a log warning. The decoder could not satisfy the ontology’s hard constraints and returned an empty assignment; this is not the same as "no facts in this text". Relax unique_source, unique_target or acyclic constraints, or lower the thresholds. Only single-pass decoding reports it: a windowed extraction cannot.

A GLiNER v1 checkpoint id

ValueError from the constructor, naming the GLiNER2.5 checkpoints. Raised before any download.

gliner2 is not installed

ImportError naming pip install "neo4j-agent-memory[gliner2]", raised when the model is first loaded, not at construction.

The checkpoint cannot be loaded

RuntimeError naming the model id.

The ontology is structurally unsound

ValueError when the JointIE schema is compiled, listing the problems that OntologyDocument.validate_structure() reports.

Use is_gliner2_available() to check for the gliner2 package before you build an extractor.

Attributes

Every entity carries these keys in attributes:

Key Value

gliner2_label

The label as the model emitted it, before POLE+O mapping

gliner2_score

The decoded confidence

rescued

Whether the span came from the decoder’s rescue pass

sentence_id

The sentence the span was found in, when the engine reports one

With extract_attributes=True, a second forward pass built from the ontology’s enum properties adds the decoded values to attributes, joined onto entity spans by label and start offset within attribute_tolerance characters.

LLMEntityExtractor

LLM-based entity, relation, and preference extraction. Inject an LLMProvider or StructuredExtractor through provider; without one, the model/API-key path uses the provider factory. This class is not limited to OpenAI.

def LLMEntityExtractor(
    provider: LLMProvider | StructuredExtractor | None=None,
    *,
    model: str | None=None,
    api_key: str | None=None,
    entity_types: list[str] | None=None,
    subtypes: dict[str, list[str]] | None=None,
    extraction_prompt: str | None=None,
    temperature: float=0.0,
    extract_relations: bool=True,
    extract_preferences: bool=True,
    ontology: OntologyDocument | None=None,
) -> None: ...
Parameter Description

provider

An LLMProvider or StructuredExtractor. When omitted, one is built from model/api_key (default model openai/gpt-4o-mini).

model

Legacy model id used to build a default provider

api_key

Legacy API key used to build a default provider

entity_types

Flat type names the model should emit. Defaults to the POLE+O types, or to the ontology’s pole_type values when ontology is set.

subtypes

Per-type subtype hints. Defaults to the POLE+O table, or to the subtypes the ontology declares.

extraction_prompt

Override for the prompt template

temperature

Sampling temperature

extract_relations

Whether to request relations

extract_preferences

Whether to request preferences

ontology

When set, the prompt’s type guidance comes from the ontology: its annotation guidelines plus the legal SOURCE -[REL]→ TARGET catalogue, so the model emits declared relationship names instead of inventing its own.

extract

Extract entities, relations, and preferences from text.

async def extract(
    text: str,
    *,
    entity_types: list[str] | None=None,
    extract_relations: bool | None=None,
    extract_preferences: bool | None=None,
) -> ExtractionResult: ...

for_poleo

Create extractor configured for POLE+O model.

def for_poleo(
    provider: LLMProvider | StructuredExtractor | None=None,
    *,
    model: str='openai/gpt-4o-mini',
    api_key: str | None=None,
) -> LLMEntityExtractor: ...

for_custom_types

Create extractor for custom entity types.

def for_custom_types(
    entity_types: list[str],
    provider: LLMProvider | StructuredExtractor | None=None,
    *,
    model: str='openai/gpt-4o-mini',
    api_key: str | None=None,
) -> LLMEntityExtractor: ...
from neo4j_agent_memory.extraction import LLMEntityExtractor
from neo4j_agent_memory.llm import from_provider

provider = from_provider("openai/gpt-4o-mini")
extractor = LLMEntityExtractor(provider=provider)
result = await extractor.extract("John Smith, CEO of Acme, spoke at the conference")