How entity extraction works
Extraction converts text into candidate entities, relationships, and preferences. It is an inference step: a well-formed result can still omit information or misidentify what the text means.
| This page describes Python’s configurable client-side extraction on Bolt. NAMS performs extraction server-side; the same client settings do not configure that pipeline. See Backend capabilities and Bolt and NAMS. |
The challenge
Consider this fictional source sentence: "Maya Chen discussed Northstar Robotics at the Denver engineering meetup."
An extractor must distinguish a person, an organization, a place, and an event. A relationship such as "discussed" is different from "works for": a mention alone does not establish employment. Domain definitions and validation help constrain the output, but neither replaces checking the source.
Three extraction approaches
-
spaCy recognizes the labels learned by its installed language model. It can run locally and is useful when that label vocabulary matches the text. Its entity output does not provide the same confidence signal as every other extractor, and it extracts no relations or preferences.
-
GLiNER2.5 (
GLiNER2Extractoron thegliner2package, with thefastino/gliner2.5-{small,base,multi}-v1checkpoints) accepts labels and their descriptions at inference time, and decodes entities and typed relations together in one pass. This allows a domain vocabulary without training a new spaCy model, at the cost of loading and running a transformer model. The model’s training still affects what it can recognize. -
LLM extraction uses a configured model provider and a structured response schema. It can use broader context, but provider calls introduce latency, resource costs, and failure modes. It is not guaranteed to be the most accurate choice for every document. It is the only stage that extracts preferences.
Because GLiNER2.5 decodes relations locally, an LLM is not the only way to produce them. Compare extractors on labeled examples from the intended domain.
GLiNER v1 and the separate GLiREL relation model were removed in 0.7. "GLiNER" on its own refers to that removed stack; the current extractor is GLiNER2Extractor on GLiNER2.5. See Migrate to GLiNER2.5.
|
Why combine extractors?
A pipeline can collect candidates from complementary extractors. A broad first stage may find common names, while another stage covers domain-specific labels. Agreement can be useful evidence, although extractors may share the same errors.
More stages also mean more work and more conflicting candidates. A fallback strategy can avoid later work after a usable result, but may miss information that another stage would have found. The extractor reference defines the stage and result objects; Configure extraction provides the setup procedure.
A pipeline built from configuration runs its enabled stages in a fixed order: spaCy, then GLiNER2.5, then the LLM fallback. A stage whose dependency is not installed logs a warning and is skipped. Every stage, and the pipeline’s final relation validation, receives the same resolved ontology.
How pipeline results combine
The implemented strategy names are confidence, union, intersection, cascade, and first_success. These combine candidate extraction results; they are separate from resolution against entities already stored in the graph.
confidence chooses the highest-scored candidate for an entity key. union retains unique keys; intersection retains repeated keys. cascade keeps the earlier candidate for an existing key and lets later stages add new keys. first_success uses the first stage with entities, limiting later work.
The entity key normalizes the name and includes its type. Therefore "Maya" and "Maya Chen" do not become the same entity solely because one has a higher confidence score. That requires separate identity evidence and resolution. Confidence values from different models are not necessarily calibrated probabilities.
Earlier versions of this page described FIRST and LAST; these are not members of MergeStrategy. Use the actual names in the reference, and see Resolution and deduplication for stored-entity identity decisions.
Domain schemas
A DomainSchema groups entity labels and their descriptions. In 0.7 it is one way to write an ontology: the eight built-in domain templates are OntologyDocument values, and DomainSchema.to_ontology() converts a custom catalog. The resolved ontology configures each Python stage differently. GLiNER2.5 compiles it into the schema it decodes against. The LLM prompt carries its descriptions and relationship catalog. spaCy retains its model’s label set; an ontology that declares subtypes only renames spaCy’s labels onto them.
Descriptions clarify ambiguous labels: "release" might mean a software release or a public statement. This gives the extractor useful context, but an improvement must be measured on representative examples. Domain schemas are neither database constraints nor a guarantee of extraction accuracy.
See Domain schemas for the catalog, Use extraction schemas for setup, Drive extraction from an ontology for a custom ontology, and Ontologies for how ontologies are stored and versioned.
Relationship extraction with GLiNER2.5
GLiNER2.5 decodes typed relations in the same pass as the entities, through its joint decoder (JointIE). The ontology declares which entity types each relationship connects, and the decoder enforces those endpoint types while it decodes rather than filtering afterwards. Each relation therefore carries the ids of its endpoint mentions (source_id and target_id), and both endpoints are entities in the same result. Decoding runs locally: it avoids remote LLM calls for relation extraction, while still requiring a model download and compute resources.
For example, the explicit statement "Maya Chen works at Northstar Robotics" provides evidence for an EMPLOYED_BY relation under the POLE+O ontology. Its extraction remains model output to validate, not a guaranteed return value. Entity-only extraction cannot establish all the connections needed for a graph.
Relations are decoded only when the ontology declares relationship types. Of the eight built-in templates, poleo, podcast, and news declare relationships; the other five are entity catalogs until relationships are attached with DomainSchema.to_ontology(relationships=[…]).
An ontology-aware pipeline finishes with ExtractionResult.validate_relations(ontology, mode="drop"). Merging can only add relations the ontology does not permit: the GLiNER2.5 stage constrains endpoints while decoding, but the spaCy and LLM stages do not. This step drops those relations in every validation mode; strict mode also drops undeclared entities at ingestion.
Two GLiNER2.5 outcomes can look like "nothing to find" when they are not:
-
feasible=Falsemeans the decoder could not satisfy the ontology’s hard constraints (unique_source,unique_target,acyclic) and returned an empty assignment. The extractor raises aRuntimeWarning. Relax those constraints or lower the thresholds instead of concluding that the text was empty. -
Long input loses relations without a warning. Relation recall drops sharply past roughly 400 words, so the extractor windows any input longer than
max_words(default 384;gliner_max_wordsinExtractionConfig). A relation whose endpoints fall in different windows cannot be decoded, and a windowed extraction never reportsfeasible=False.
Compared with the GLiNER v1 stack
GLiNER2.5 is an architecture change, not an accuracy jump. Measured on the POLE+O gold set in benchmarks/data/ with benchmarks/compare_extractors.py (CPU):
| Checkpoint | Entity F1 | Relation F1 (lenient / strict) | Throughput |
|---|---|---|---|
|
0.80 |
0.37 / 0.27 |
~1.6 docs/s |
|
0.75 |
not measured |
faster |
Both runs reported zero endpoint-type violations, which is the joint decoder’s constraint at work. The other gains are structural. One model replaces two, so there is no separate relation model to install, version or keep in memory. The model trains on sequences up to 4,096 words and decodes spans of any length instead of a fixed span-width grid, although relation recall still degrades past roughly 400 words, which is why the extractor windows its input. Entity descriptions are read as annotation guidelines (see Domain schemas).
What did not change: entity F1 is roughly level with the v1 stack, and end-to-end zero-shot relation extraction is weak for every model in this class. Treat relation F1 as a floor to improve with per-relation thresholds.
From extraction to stored relationships
On Bolt, supported message-ingestion operations can store extracted entities, link messages through :MENTIONS, and persist candidate relations as :RELATED_TO edges whose type property names the relation. This is distinct from an application-owned graph using physical relationship types such as :WORKS_AT.
The type property is part of the merge key, so one pair of entities can carry several differently typed edges. r.type is the canonical property; r.relation_type is still written as a mirror for one release. Each edge accumulates provenance: support counts observations, derived stays true only while every observation was derived, and extractor, source_message_ids (up to 25), and evidence (up to 3) record where it came from. Edges written before 0.7 carry only relation_type; a one-time backfill copies it into type on connect, unless schema_config.backfill_relation_types=False.
Message ingestion also resolves extracted entities against stored entities by default (resolution.resolve_on_ingest=True). A confident match reuses the existing node; a near match creates a node and a pending :SAME_AS edge for review. Relation endpoints therefore land on the resolved nodes.
A relation needs resolvable endpoints. The writer prefers the mention ids from the same extraction call, then entities from the same message, then a case-insensitive name lookup in the graph, which lets a relation reach an entity from an earlier message. Ambiguous names and missing entities can prevent a relation from being stored correctly, and a relation whose endpoints resolve to the same node is skipped. Requesting relation extraction is not proof that an edge was created. Inspect the stored result as part of the ingestion task.
The short-term API reference owns method-specific flags and return values. Follow Store messages or Configure extraction for the executable procedure. NAMS uses its own server-side extraction and storage contract.
Latency, batching, and document boundaries
There is no universal extraction time or document-size cutoff. Runtime depends on the model, hardware, input length, label set, number of stages, provider limits, and concurrent work. Evaluate latency together with missed entities and false positives.
Batching controls how many documents are processed together; concurrency controls simultaneous work. Larger values can increase memory use or encounter provider limits. They are not interchangeable controls. ExtractionPipeline.extract_batch() exposes both (batch_size and max_concurrency). GLiNER2Extractor.extract_batch() exposes only batch_size, because it decodes each slice of texts in one forward pass.
Streaming extraction divides a document into overlapping chunks. StreamingExtractor measures chunk_size and overlap in characters by default; chunk_by_tokens=True selects token-based chunking. Choose boundaries according to the actual extractor’s context window and the document structure, rather than an arbitrary 100K-token rule. Overlap preserves context at boundaries but adds repeated work and duplicate candidates. GLiNER2.5 also windows its own input at max_words, whatever chunk size the streamer uses.
See Batch and streaming extraction for configuration and verification.
Choosing an approach
Start with the information the application needs: entity types, relationships, and acceptable uncertainty. Then compare the smallest suitable extraction path on representative text, including ambiguous names, absent entities, long documents, and conflicting statements.
For locally processed text, account for model initialization and memory use. For remote models, account for provider latency and failures. For a combined pipeline, compare the extra coverage against the extra work. No fixed ranking substitutes for this evaluation.
When the graph needs typed relations without an LLM call, GLiNER2.5 on an ontology that declares relationship types is the local option. Zero-shot relation extraction is weaker than entity extraction for models in this class, so measure relation quality separately from entity quality.