Migrate to GLiNER2.5
|
Runs client-side. The extractors on this page run in your Python process and do not call a memory backend, so they work next to either backend. Configuring them as a |
Release 0.7 replaced the local extraction stack. The gliner v1 package and the
separate GLiREL relation model are gone; GLiNER2Extractor on GLiNER2.5 does
both jobs in one pass. This guide covers what to change in your install, your
code, your configuration and your stored data.
|
Naming. "GLiNER" refers to the removed v1 package and its community
checkpoints. "GLiNER2.5" means the |
1. Update the install
python -m pip install 'neo4j-agent-memory[gliner2]==0.7.0'
| Extra | Status |
|---|---|
|
The extra to use. Installs |
|
Deprecated alias of |
|
Still spaCy + the local model; now carries |
|
Unchanged spelling; carries |
You can uninstall gliner and glirel: nothing in the library imports them any
more.
2. Change the model id
| Checkpoint | Parameters | On disk | Use |
|---|---|---|---|
|
74M |
~296 MB |
Fast CPU / edge |
|
194M |
~407 MB |
Default, English |
|
287M |
~594 MB |
Multilingual |
Legacy ids — anything under urchade/, gliner-community/ or numind/ —
raise a ValueError from the constructor, before any download:
ValueError: 'urchade/gliner_medium-v2.1' is a GLiNER v1 checkpoint and cannot be
loaded by GLiNER2Extractor (the 2.5 loader rejects the architecture). Use a
GLiNER2.5 checkpoint such as 'fastino/gliner2.5-base-v1',
'fastino/gliner2.5-small-v1' or 'fastino/gliner2.5-multi-v1'.
Settings validation does not reject them, so a stale
NAM_EXTRACTION__GLINER_MODEL surfaces when the client connects and builds its
extractor, not when the settings load.
3. Update your code
Entity extraction
The diffs in this section show the 0.6 code being removed (-) and its 0.7
replacement (+).
-from neo4j_agent_memory.extraction import GLiNEREntityExtractor
+from neo4j_agent_memory.extraction import GLiNER2Extractor
-extractor = GLiNEREntityExtractor.for_schema("podcast")
-extractor = GLiNEREntityExtractor.for_poleo()
-extractor = GLiNEREntityExtractor(entity_labels=["person", "company"])
+extractor = GLiNER2Extractor.for_schema("podcast")
+extractor = GLiNER2Extractor.for_poleo()
+extractor = GLiNER2Extractor(entity_labels=["person", "company"])
schema= became ontology=, and it accepts more: an OntologyDocument, a
DomainSchema, an EntitySchemaConfig, or anything else with a to_ontology()
method. An EntitySchemaConfig converts cleanly only when its type names are
POLE+O types; any other name fails the ontology’s structural validation on the
first extraction.
-extractor = GLiNEREntityExtractor(schema=my_domain_schema)
+extractor = GLiNER2Extractor(ontology=my_domain_schema)
+extractor = GLiNER2Extractor.for_ontology(my_ontology_document)
Relation extraction
Relations now come from the same pass as the entities. There is no second model, no second install and no second call.
-from neo4j_agent_memory.extraction import (
- GLiNEREntityExtractor,
- GLiRELExtractor,
- GLiNERWithRelationsExtractor,
-)
+from neo4j_agent_memory.extraction import GLiNER2Extractor
-entity_result = await GLiNEREntityExtractor.for_poleo().extract(text)
-relations = await GLiRELExtractor().extract_relations(text, entities=entity_result.entities)
-# or
-extractor = GLiNERWithRelationsExtractor.for_poleo()
+extractor = GLiNER2Extractor.for_poleo()
+result = await extractor.extract(text)
+result.entities
+result.relations
Relations are decoded when three things hold: the call passes
extract_relations=True (the default), the extractor was built with
extract_relations=True, and the ontology declares relationship types. That
last condition is the one that bites: a bare DomainSchema is a label catalog
with no endpoint typing, so GLiNER2Extractor(ontology=some_domain_schema)
returns entities only. The poleo, podcast and news templates declare
relationships; the other five do not until you attach some.
The builder
The builder keeps its spelling. with_gliner() now means GLiNER2.5 and takes
new keyword arguments:
from neo4j_agent_memory.extraction import ExtractorBuilder
extractor = (
ExtractorBuilder()
.with_spacy()
.with_gliner(threshold=0.5, relation_threshold=0.35)
.with_llm_fallback()
.build()
)
# Domain template, unchanged
ExtractorBuilder().with_gliner_schema("podcast", threshold=0.5).build()
# New: extract against an ontology document
ExtractorBuilder().with_ontology(doc).with_gliner().build()
with_gliner_relations() is gone — it had no replacement because it has no job
left to do.
Removed names
Importing any of these from neo4j_agent_memory.extraction raises an
ImportError naming its replacement:
| Removed in 0.7 | Replacement |
|---|---|
|
|
|
|
|
|
|
|
|
|
|
Declare relationships on an |
|
|
|
|
|
|
4. Output shape changes
| What | Change |
|---|---|
|
|
|
New. A per-result mention id, scoped to one extraction call. |
|
Keys are now |
|
New. The mention ids of the endpoints. Storage prefers them over a name lookup, so a name mentioned twice no longer collapses into one edge. |
|
New. |
Relations present at all |
Only when the ontology declares relationship types (see above). |
Code that read entity.attributes["gliner_label"] or matched
extractor == "gliner" needs updating.
5. Configuration changes
The ExtractionConfig field names keep the gliner_ prefix, so existing
NAM_EXTRACTION__GLINER_* environment variables keep working. What changed:
| Field | Change | Note |
|---|---|---|
|
Default is now |
Was a community v1 checkpoint |
|
New, default |
Confidence floor for decoded relations |
|
New, default |
Longer input is windowed |
|
New, default |
Word overlap between windows |
|
New, default |
Span-overlap policy for the attribute pass |
|
New, default |
Opt-in second pass for enum properties |
|
New, default |
fp16 weights |
|
New, default |
|
ExtractionConfig uses extra="forbid", so a misspelled field raises at
construction rather than being silently dropped.
6. Stored data: the RELATED_TO backfill
Extracted relationships are still RELATED_TO edges, but the merge key now
includes the relationship name:
MERGE (source)-[r:RELATED_TO {type: $relation_type}]->(target)
r.type is the canonical property and every reader uses it; r.relation_type
is written as a mirror for one release. Edges also carry provenance: support
(observation count), derived, extractor, source_message_ids (capped at 25)
and evidence (capped at 3).
Edges written before 0.7 only carried relation_type. A one-shot backfill runs
during schema setup on the first connect:
MATCH ()-[r:RELATED_TO]->()
WHERE r.type IS NULL AND r.relation_type IS NOT NULL
SET r.type = r.relation_type, r.support = coalesce(r.support, 1)
The query is idempotent — it matches nothing once every edge has been migrated,
including on an empty database — but it is still a write transaction that scans
every RELATED_TO edge, which is not something to pay for on every connection.
So it runs once per database: completion is recorded on a marker node and
later connections skip the scan.
(:SchemaMigration {name: "relation_type_backfill", completed_at: datetime()})
To run it again — after bulk-loading pre-0.7 data, say — delete the marker and reconnect:
MATCH (m:SchemaMigration {name: "relation_type_backfill"}) DELETE m
To skip it entirely and run the equivalent yourself in batches
(apoc.periodic.iterate), turn it off. The settings read the NEO4J_*
variables exported for your AuraDB instance, as shown in
the Aura connection setup:
import os
from neo4j_agent_memory import MemorySettings
settings = MemorySettings(
backend="bolt",
neo4j={
"uri": os.environ["NEO4J_URI"],
"username": os.environ["NEO4J_USERNAME"],
"password": os.environ["NEO4J_PASSWORD"],
"database": os.getenv("NEO4J_DATABASE", "neo4j"),
},
schema_config={"backfill_relation_types": False},
)
Or set NAM_SCHEMA_CONFIG__BACKFILL_RELATION_TYPES=false. On a very large
graph, prefer this and migrate during a maintenance window.
The same switch gates a second one-time backfill, entity_keys_backfill, which
fills the name_key and surface_keys properties entity resolution looks
entities up by. Run it yourself too when you turn the switch off; the
schema objects reference describes both
properties.
Because the merge key changed, one pair of entities can now carry several differently-typed edges where before a second type would overwrite the first. Expect the edge count to grow on re-ingestion.
7. Adjust for the new model’s behavior
Expect entity quality roughly level with the v1 stack, and relation quality that needs per-relation thresholds; Compared with the GLiNER v1 stack has the measurements and what changed. Then check for the two outcomes that look like an empty extraction when they are not:
-
Long input loses relations without a warning. Relation recall degrades sharply past roughly 400 words while entity recall stays fine. Input longer than
max_wordsis windowed throughextract_long, and a relation whose endpoints land in different windows is lost. Extract per message or per section rather than per long concatenation. -
A
RuntimeWarningaboutfeasible=False. The decoder could not satisfy the ontology’s hard constraints (uniqueness, acyclicity) and fell back to an empty assignment. This is not "no facts in this text": relax the constraints or lower the thresholds.
8. Verify
from neo4j_agent_memory.extraction import GLiNER2Extractor, is_gliner2_available
assert is_gliner2_available()
extractor = GLiNER2Extractor.for_poleo()
print(extractor.name, extractor.version, extractor.model_id)
# gliner2 2.0.0 fastino/gliner2.5-base-v1
result = await extractor.extract("Brian Chesky founded Airbnb in San Francisco.")
assert result.entities
assert all(e.id for e in result.entities)
assert all(r.source_id and r.target_id for r in result.relations)
From the command line, with the cli extra installed:
neo4j-agent-memory extract "Brian Chesky founded Airbnb in San Francisco." --format json
See also
-
Built-in extractor classes — the full
GLiNER2Extractorparameter list -
Extract and inspect entity candidates — a complete local extraction program
-
Drive extraction from an ontology — typed entities and relations from one document
-
Domain schemas reference — the eight catalogs and their relationships
-
Understanding the extraction pipeline — stage design and merge strategies
-
Extraction configuration — every extraction field