Relation extractors
How the extractors from Extractor classes reference produce typed relations, and how relations are checked against an ontology. Import from neo4j_agent_memory.extraction.
|
|
Relation sources
| Extractor | Relations | Endpoint typing |
|---|---|---|
|
Decoded jointly with the entities (JointIE), for the relationship types the ontology declares |
Enforced during decoding. Relations carry |
|
Returned by the same model call when |
Not enforced. With |
|
None |
— |
Joint decoding with GLiNER2.5
GLiNER2Extractor compiles the ontology into a JointIE schema and decodes entities and relations in one pass. Relations are decoded only when all three of these hold: the call passes extract_relations=True (the default), the extractor was built with extract_relations=True, and the ontology declares at least one relationship type. An entity_labels list, and a DomainSchema converted without relationships, declare none, so they decode entities only.
-
Every endpoint is one of the entities in the same result, and the declared
(source, type, target)patterns are enforced during beam search. -
source_idandtarget_idhold the mention ids of the endpoints (ExtractedEntity.id), so a relation names the exact mention when one name occurs twice. -
relation_typeis normalized to UPPER_SNAKE withnormalize_relation_type, soworks atandWorks-Atboth becomeWORKS_AT. -
derivedis set when the engine reports a relation it inferred, such as the mirror of a declaredinverse. -
The ontology’s per-relation
thresholdis compiled into the schema; the extractor’srelation_threshold, when set, then drops anything below it. -
With an
entity_typesfilter, a relationship type is compiled only when every endpoint label it declares survives the filter. -
Input longer than
max_wordsis windowed; relations whose endpoints fall in different windows are lost.
from neo4j_agent_memory.extraction import GLiNER2Extractor
extractor = GLiNER2Extractor.for_poleo(relation_threshold=0.6)
result = await extractor.extract("John Smith works for Acme Corp in Boston.")
for relation in result.relations:
print(relation.as_triple, relation.source_id, relation.target_id, relation.confidence)
Relationship constraints
The RelationshipDef fields of an OntologyDocument that shape decoding. Several defs may share one type to declare several legal endpoint pairs; they compile into one JointIE relation.
| Field | Effect |
|---|---|
|
Relationship type name (UPPER_SNAKE) |
|
Endpoint labels; each must be a declared entity label |
|
Annotation guideline for the relation; the first non-empty one among defs of a type is used |
|
Per-relation confidence floor; the lowest floor among defs of a type is used |
|
The inverse relationship type; compiled only when that type is also in the schema |
|
At most one relation of this type per source |
|
At most one relation of this type per target |
|
No cycles of this type |
|
Permit self-loops for this type; one def that allows them is enough |
The document-level no_self_loops (default True) forbids self-loops on every relationship type that does not set allow_self; it is applied per relation, never as a global flag. Defs sharing a type must agree on inverse, unique_source, unique_target and acyclic; OntologyDocument.validate_structure() reports a disagreement, and compiling such a document raises ValueError. RelationshipDef has no symmetric field, and the compiler never passes symmetric to JointIE: it rejects every candidate edge in gliner2 2.0.0. Use inverse instead.
Hard constraints can make a decode infeasible. The decoder then falls back to an empty assignment and GLiNER2Extractor emits a RuntimeWarning; that is not the same as "no facts in this text". See Warnings and errors.
ExtractionResult.validate_relations
Check every relation against the ontology’s endpoint typing.
def validate_relations(
ontology: 'OntologyDocument',
*,
mode: Literal['warn', 'drop', 'raise']='warn',
) -> 'tuple[ExtractionResult, list[ExtractedRelation]]': ...
A relation violates the ontology when its relation_type is not a declared relationship type, when either endpoint cannot be resolved to an entity in the same result, or when the resolved endpoint labels are not a declared (source, type, target) pattern (OntologyDocument.permits). Endpoints resolve by source_id/target_id first, then by normalized name. Relationship type names are compared through normalize_relation_type on both sides.
mode |
Behavior |
|---|---|
|
Log the violations and keep them (the default) |
|
Return a copy without them |
|
Raise |
The return value is (result, violations): the result to use (the same object unless drop removed something) and the violating relations. An ontology that declares no relationships cannot express a violation, so the result passes through unchanged.
from neo4j_agent_memory.ontology import POLEO_ONTOLOGY
checked, violations = result.validate_relations(POLEO_ONTOLOGY, mode="drop")
for relation in violations:
print("not permitted:", relation.as_triple)
Validation also runs without a call from you:
-
ExtractionPipelinebuilt withontologyvalidates the merged result withmode="drop"; see Extraction pipeline. -
Message ingestion on bolt drops relations that the client’s ontology does not permit, in both validation modes.
strictmode also drops entities whose type the ontology does not declare, with the relations that referenced them.
Stored relations
Message ingestion with relation extraction enabled stores each relation as a RELATED_TO edge whose canonical name is r.type. The merge key is {type: …}, so one pair of entities can carry several differently typed edges. Repeated observations increment r.support, keep the highest r.confidence, fold r.derived with AND, and append to r.source_message_ids (capped at 25) and r.evidence (capped at 3). Writes prefer source_id/target_id over a name lookup. r.relation_type is written as a mirror of r.type for one release.