Extractor classes reference
Local Python extraction classes and result models. They can be used independently or injected into a bolt MemoryClient; they do not configure the hosted NAMS pipeline.
Signatures below show types, keyword-only arguments (*), and defaults. They are reference declarations; the examples show calls to execute.
See Backend capabilities for availability and data scope. This page is a short overview and index; each subsystem has its own page.
|
GLiNER v1 and GLiREL were removed in 0.7. |
On this subsystem
-
Built-in extractor classes —
SpacyEntityExtractor,GLiNER2Extractor(GLiNER2.5 joint entity and relation extraction), andLLMEntityExtractor. -
Relation extractors — typed relations from the GLiNER2.5 joint pass and the LLM extractor, the ontology constraints that shape them, and
ExtractionResult.validate_relations. -
Extraction pipeline —
ExtractionPipeline, merge strategies, the confidence floor and ontology relation validation, and the pipeline/batch result models. -
Extractor builder —
ExtractorBuilderfluent construction, includingwith_ontology. -
Streaming extraction —
StreamingExtractorfor chunked long-document extraction.
EntityExtractor protocol
The protocol requires extract, not extract_batch. Batch methods are concrete capabilities with different signatures on GLiNER2Extractor and ExtractionPipeline; check the selected class.
async def extract(
text: str,
*,
entity_types: list[str] | None=None,
extract_relations: bool=True,
extract_preferences: bool=True,
) -> ExtractionResult: ...
Result models
These are Pydantic models imported from neo4j_agent_memory.extraction. ExtractedEntity uses type and attributes, not entity_type and metadata. ExtractionResult has entities, relations, preferences, and source_text; it has no metadata envelope.
ExtractionResult
Fields and defaults:
| Field | Type | Default | Description |
|---|---|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Original source text |
ExtractedEntity
Fields and defaults:
| Field | Type | Default | Description |
|---|---|---|---|
|
|
|
Entity name |
|
|
|
Entity type (PERSON, OBJECT, LOCATION, EVENT, ORGANIZATION) |
|
|
|
Entity subtype (e.g., VEHICLE for OBJECT, ADDRESS for LOCATION) |
|
|
|
Start position in text |
|
|
|
End position in text |
|
|
|
Confidence score |
|
|
|
Surrounding context |
|
|
|
Additional attributes extracted for this entity |
|
|
|
Name of the extractor that produced this entity |
|
|
|
Mention id assigned by the extractor, scoped to one extraction call. GLiNER2.5 sets it so relations can name the exact mention when one name occurs twice in a text. |
ExtractedRelation
Fields and defaults:
| Field | Type | Default | Description |
|---|---|---|---|
|
|
|
Source entity name |
|
|
|
Target entity name |
|
|
|
Type of relationship |
|
|
|
Confidence score |
|
|
|
|
|
|
|
|
|
|
|
|
ExtractedPreference
Fields and defaults:
| Field | Type | Default | Description |
|---|---|---|---|
|
|
|
Preference category |
|
|
|
The preference statement |
|
|
|
Context where preference applies |
|
|
|
Confidence score |
ExtractionResult.entity_count, relation_count, and preference_count are properties. entities_by_type() groups entities into a dictionary; get_entities_of_type(type) returns a list using case-insensitive type matching. filter_invalid_entities(ontology=None) returns a new result with invalid names and affected relations removed. Invalid names are stopwords, numbers and punctuation, and mentions that only name their own type, such as "tickets" typed Ticket. Type names are the mention’s POLE+O type, subtype and extractor label, plus the label ontology declares for its pair. Message ingestion applies it with the client’s ontology. validate_relations(ontology, *, mode="warn") checks every relation against an ontology’s endpoint typing and returns (result, violations); see Relation extractors. There are no filter_by_type or filter_by_confidence methods.
people = result.get_entities_of_type("PERSON")
high_confidence = [entity for entity in result.entities if entity.confidence >= 0.7]
print(result.entities_by_type())
ExtractedEntity.normalized_name lowercases/normalizes its name and full_type includes its subtype. ExtractedRelation.as_triple returns source, relation type, and target.
See also
-
Batch & Streaming — handling
BatchExtractionResultandStreamingChunkResulterrors