Built-in extractor classes
The concrete single-stage extractors from Extractor classes reference: statistical (spaCy), joint entity and relation extraction (GLiNER2.5), and LLM-based. Import from neo4j_agent_memory.extraction.
SpacyEntityExtractor
Statistical NER using a separately installed spaCy model. Import from neo4j_agent_memory.extraction. Install the matching extraction extra and model before running inference.
def SpacyEntityExtractor(
model: str='en_core_web_sm',
type_mapping: dict[str, str] | None=None,
subtype_mapping: dict[str, str] | None=None,
default_confidence: float=0.85,
context_window: int=50,
*,
drop_unmapped: bool=False,
): ...
| Parameter | Description |
|---|---|
|
spaCy model name (e.g., "en_core_web_sm", "en_core_web_lg") |
|
Custom mapping from spaCy labels to POLE+O types |
|
Custom mapping from spaCy labels to subtypes |
|
Default confidence score (spaCy doesn’t provide per-entity scores) |
|
Number of characters of context to include around entities |
|
Skip entities whose spaCy label |
The default mappings include PERSON → PERSON, ORG → ORGANIZATION, GPE/LOC → LOCATION, EVENT → EVENT, and PRODUCT/WORK_OF_ART → OBJECT. Subtype mappings are retained on results. Override the mapping explicitly for your domain. spaCy extracts no relations.
When ExtractorBuilder or the configuration factory builds the spaCy stage against an ontology that declares subtypes, each spaCy label is remapped onto the first of these pairs the ontology declares: spaCy’s own subtype for the label (GPE → LOCATION:GEOPOLITICAL), the label itself as the subtype (LAW → OBJECT:LAW), then the bare POLE+O type. A label none of them matches is not extracted (drop_unmapped=True). spaCy never lands on a subtype it did not decide, so under the medical template it extracts nothing and under scientific a person is not made an author. An ontology with base POLE+O labels only keeps the defaults.
from neo4j_agent_memory.extraction import SpacyEntityExtractor
extractor = SpacyEntityExtractor(model="en_core_web_sm")
result = await extractor.extract("John works at Apple in California")
GLiNER2Extractor
Joint entity and relation extraction with a GLiNER2.5 checkpoint, locally and without an LLM call. Install the gliner2 extra, for example pip install 'neo4j-agent-memory[gliner2]==0.7.0', and allow the checkpoint to download on first use. The gliner extra is a deprecated alias of gliner2 in 0.7.
GLiNER2.5 decodes entities and relations in one pass through its JointIE engine. The ontology’s endpoint types are enforced during beam search rather than filtered afterwards, so every relation endpoint is one of the entities in the same result and relations carry mention ids. Preferences are never extracted; pair this extractor with LLMEntityExtractor when you need them. The checkpoint loads lazily when inference is first requested.
def GLiNER2Extractor(
model: str=DEFAULT_GLINER2_5_MODEL,
*,
ontology: OntologyDocument | DomainSchema | EntitySchemaConfig | None=None,
entity_labels: list[str] | dict[str, str] | None=None,
threshold: float=0.5,
relation_threshold: float | None=None,
device: str='cpu',
quantize: bool=False,
compile: bool=False,
max_words: int=384,
chunk_overlap: int=64,
overlap_policy: str | None=None,
extract_relations: bool=True,
extract_attributes: bool=False,
attribute_tolerance: int=2,
context_window: int=50,
label_mapping: dict[str, tuple[str, str | None]] | None=None,
joint_config: Any | None=None,
): ...
| Parameter | Description |
|---|---|
|
GLiNER2.5 checkpoint id or local path. |
|
The schema to extract against: an |
|
Alternative to |
|
Entity confidence floor passed to the decoder. |
|
Confidence floor applied to decoded relations. |
|
Inference device ( |
|
Load the weights in fp16. |
|
Run the weights through |
|
Longest input decoded in one pass; longer text is windowed. Must be greater than 0. |
|
Word overlap between windows. Must be at least 0 and smaller than |
|
Span-overlap policy for the attribute pass ( |
|
Set |
|
Run the opt-in second forward pass that recovers the ontology’s enum properties as span attributes. |
|
Character slack when joining the attribute pass back onto entity spans (the two passes tokenize independently). |
|
Characters of surrounding text captured per entity. |
|
Replacement for the label → |
|
A ready-made |
| Checkpoint | Size on disk | Use |
|---|---|---|
|
~296 MB |
Faster, CPU and edge |
|
~407 MB |
Default |
|
~594 MB |
Multilingual |
The extractor exposes name ("gliner2", recorded on every entity as ExtractedEntity.extractor), version (the installed gliner2 version, or "unknown"), model_id, ontology (the OntologyDocument it compiles), and the lazily loaded model and joint (JointIE engine) handles.
extract
Extract entities and relations from one text.
async def extract(
text: str,
*,
entity_types: list[str] | None=None,
extract_relations: bool=True,
extract_preferences: bool=True,
) -> ExtractionResult: ...
| Parameter | Description |
|---|---|
|
The text to extract from |
|
Restrict extraction to these ontology labels or POLE+O types (case-insensitive) |
|
Set |
|
Ignored (GLiNER2.5 extracts no preferences) |
Relations are decoded only when the call passes extract_relations=True, the extractor was built with extract_relations=True, and the ontology declares at least one relationship type. Relation types are normalized to UPPER_SNAKE. See Relation extractors.
Input longer than max_words (384 words, counted the way the gliner2 chunker counts them) is windowed through the engine’s extract_long with chunk_overlap words of overlap, and relations whose endpoints fall in different windows are lost. Relation recall degrades sharply past roughly 400 words, so prefer per-message extraction or explicit chunking with StreamingExtractor over one long concatenation.
extract_batch
Extract from several texts using the engine’s batched decoding.
async def extract_batch(
texts: list[str],
*,
entity_types: list[str] | None=None,
extract_relations: bool=True,
batch_size: int=8,
on_progress: Callable[[int, int], None] | None=None,
) -> list[ExtractionResult]: ...
| Parameter | Description |
|---|---|
|
The texts to extract from, in order |
|
Restrict extraction to these labels or POLE+O types |
|
Set |
|
Texts decoded per forward pass |
|
Optional callback called after each slice. Receives (completed_count, total_count). |
Returns one ExtractionResult per input text, in input order. Texts within max_words go through batched decoding; longer texts are windowed individually. There is no BatchExtractionResult wrapper and no max_concurrency; wrap the extractor in ExtractionPipeline for those.
for_schema
Build an extractor from a built-in domain template.
def for_schema(schema_name: str, **kwargs: Any) -> GLiNER2Extractor: ...
| Parameter | Description |
|---|---|
|
Template name (poleo, podcast, news, scientific, business, entertainment, medical, legal), resolved through |
|
Forwarded to the constructor, for example |
The extractor receives the template’s relationships as well as its labels: poleo, podcast and news declare relationships; the other five templates are entity-only. See Built-in domain schemas.
for_poleo
Build an extractor on the built-in POLE+O ontology (neo4j_agent_memory.ontology.POLEO_ONTOLOGY). Constructing GLiNER2Extractor() with no ontology or entity_labels gives the same ontology.
def for_poleo(**kwargs: Any) -> GLiNER2Extractor: ...
for_ontology
Build an extractor on an explicit ontology document.
def for_ontology(ontology: OntologyDocument, **kwargs: Any) -> GLiNER2Extractor: ...
from neo4j_agent_memory.extraction import GLiNER2Extractor
extractor = GLiNER2Extractor.for_schema("podcast", threshold=0.6)
result = await extractor.extract("Brian Chesky founded Airbnb in San Francisco.")
for entity in result.entities:
print(entity.name, entity.full_type, entity.id)
for relation in result.relations:
print(relation.source, relation.relation_type, relation.target)
Warnings and errors
| Condition | Behavior |
|---|---|
The decoder returns |
A |
A GLiNER v1 checkpoint id |
|
|
|
The checkpoint cannot be loaded |
|
The ontology is structurally unsound |
|
Use is_gliner2_available() to check for the gliner2 package before you build an extractor.
Attributes
Every entity carries these keys in attributes:
| Key | Value |
|---|---|
|
The label as the model emitted it, before POLE+O mapping |
|
The decoded confidence |
|
Whether the span came from the decoder’s rescue pass |
|
The sentence the span was found in, when the engine reports one |
With extract_attributes=True, a second forward pass built from the ontology’s enum properties adds the decoded values to attributes, joined onto entity spans by label and start offset within attribute_tolerance characters.
LLMEntityExtractor
LLM-based entity, relation, and preference extraction. Inject an LLMProvider or StructuredExtractor through provider; without one, the model/API-key path uses the provider factory. This class is not limited to OpenAI.
def LLMEntityExtractor(
provider: LLMProvider | StructuredExtractor | None=None,
*,
model: str | None=None,
api_key: str | None=None,
entity_types: list[str] | None=None,
subtypes: dict[str, list[str]] | None=None,
extraction_prompt: str | None=None,
temperature: float=0.0,
extract_relations: bool=True,
extract_preferences: bool=True,
ontology: OntologyDocument | None=None,
) -> None: ...
| Parameter | Description |
|---|---|
|
An |
|
Legacy model id used to build a default provider |
|
Legacy API key used to build a default provider |
|
Flat type names the model should emit. Defaults to the POLE+O types, or to the ontology’s |
|
Per-type subtype hints. Defaults to the POLE+O table, or to the subtypes the ontology declares. |
|
Override for the prompt template |
|
Sampling temperature |
|
Whether to request relations |
|
Whether to request preferences |
|
When set, the prompt’s type guidance comes from the ontology: its annotation guidelines plus the legal |
extract
Extract entities, relations, and preferences from text.
async def extract(
text: str,
*,
entity_types: list[str] | None=None,
extract_relations: bool | None=None,
extract_preferences: bool | None=None,
) -> ExtractionResult: ...
for_poleo
Create extractor configured for POLE+O model.
def for_poleo(
provider: LLMProvider | StructuredExtractor | None=None,
*,
model: str='openai/gpt-4o-mini',
api_key: str | None=None,
) -> LLMEntityExtractor: ...
for_custom_types
Create extractor for custom entity types.
def for_custom_types(
entity_types: list[str],
provider: LLMProvider | StructuredExtractor | None=None,
*,
model: str='openai/gpt-4o-mini',
api_key: str | None=None,
) -> LLMEntityExtractor: ...
from neo4j_agent_memory.extraction import LLMEntityExtractor
from neo4j_agent_memory.llm import from_provider
provider = from_provider("openai/gpt-4o-mini")
extractor = LLMEntityExtractor(provider=provider)
result = await extractor.extract("John Smith, CEO of Acme, spoke at the conference")