Extractor classes reference

Local Python extraction classes and result models. They can be used independently or injected into a bolt MemoryClient; they do not configure the hosted NAMS pipeline.

Signatures below show types, keyword-only arguments (*), and defaults. They are reference declarations; the examples show calls to execute.

See Backend capabilities for availability and data scope. This page is a short overview and index; each subsystem has its own page.

GLiNER v1 and GLiREL were removed in 0.7. GLiNEREntityExtractor, GLiNERConfig, GLiNERWithRelationsExtractor, GLiRELExtractor, GLiRELConfig, DEFAULT_RELATION_TYPES, is_gliner_available and is_glirel_available now raise ImportError (RemovedExtractorError) with a migration hint, and ExtractorBuilder.with_gliner_relations() no longer exists. GLiNER2Extractor replaces them: it decodes entities and relations in one GLiNER2.5 pass. See Migrate to GLiNER2.5.

On this subsystem

  • Built-in extractor classes — SpacyEntityExtractor, GLiNER2Extractor (GLiNER2.5 joint entity and relation extraction), and LLMEntityExtractor.

  • Relation extractors — typed relations from the GLiNER2.5 joint pass and the LLM extractor, the ontology constraints that shape them, and ExtractionResult.validate_relations.

  • Extraction pipeline — ExtractionPipeline, merge strategies, the confidence floor and ontology relation validation, and the pipeline/batch result models.

  • Extractor builder — ExtractorBuilder fluent construction, including with_ontology.

  • Streaming extraction — StreamingExtractor for chunked long-document extraction.

EntityExtractor protocol

The protocol requires extract, not extract_batch. Batch methods are concrete capabilities with different signatures on GLiNER2Extractor and ExtractionPipeline; check the selected class.

async def extract(
    text: str,
    *,
    entity_types: list[str] | None=None,
    extract_relations: bool=True,
    extract_preferences: bool=True,
) -> ExtractionResult: ...

Result models

These are Pydantic models imported from neo4j_agent_memory.extraction. ExtractedEntity uses type and attributes, not entity_type and metadata. ExtractionResult has entities, relations, preferences, and source_text; it has no metadata envelope.

ExtractionResult

Fields and defaults:

Field Type Default Description

entities

list[ExtractedEntity]

[]

relations

list[ExtractedRelation]

[]

preferences

list[ExtractedPreference]

[]

source_text

str | None

None

Original source text

ExtractedEntity

Fields and defaults:

Field Type Default Description

name

str

required

Entity name

type

str

required

Entity type (PERSON, OBJECT, LOCATION, EVENT, ORGANIZATION)

subtype

str | None

None

Entity subtype (e.g., VEHICLE for OBJECT, ADDRESS for LOCATION)

start_pos

int | None

None

Start position in text

end_pos

int | None

None

End position in text

confidence

float

1.0

Confidence score

context

str | None

None

Surrounding context

attributes

dict[str, Any]

{}

Additional attributes extracted for this entity

extractor

str | None

None

Name of the extractor that produced this entity

id

str | None

None

Mention id assigned by the extractor, scoped to one extraction call. GLiNER2.5 sets it so relations can name the exact mention when one name occurs twice in a text.

ExtractedRelation

Fields and defaults:

Field Type Default Description

source

str

required

Source entity name

target

str

required

Target entity name

relation_type

str

required

Type of relationship

confidence

float

1.0

Confidence score

source_id

str | None

None

ExtractedEntity.id of the source mention from the same extraction call. Storage prefers it over a name lookup.

target_id

str | None

None

ExtractedEntity.id of the target mention from the same extraction call.

derived

bool

False

True when the relation was inferred (for example an inverse mirror) rather than observed. Storage folds it with AND across observations, so one asserted sighting clears it.

ExtractedPreference

Fields and defaults:

Field Type Default Description

category

str

required

Preference category

preference

str

required

The preference statement

context

str | None

None

Context where preference applies

confidence

float

1.0

Confidence score

ExtractionResult.entity_count, relation_count, and preference_count are properties. entities_by_type() groups entities into a dictionary; get_entities_of_type(type) returns a list using case-insensitive type matching. filter_invalid_entities(ontology=None) returns a new result with invalid names and affected relations removed. Invalid names are stopwords, numbers and punctuation, and mentions that only name their own type, such as "tickets" typed Ticket. Type names are the mention’s POLE+O type, subtype and extractor label, plus the label ontology declares for its pair. Message ingestion applies it with the client’s ontology. validate_relations(ontology, *, mode="warn") checks every relation against an ontology’s endpoint typing and returns (result, violations); see Relation extractors. There are no filter_by_type or filter_by_confidence methods.

people = result.get_entities_of_type("PERSON")
high_confidence = [entity for entity in result.entities if entity.confidence >= 0.7]
print(result.entities_by_type())

ExtractedEntity.normalized_name lowercases/normalizes its name and full_type includes its subtype. ExtractedRelation.as_triple returns source, relation type, and target.