Frequently asked questions
This FAQ answers common questions and links to the task or reference page that owns the details. Backend capabilities and installed release versions affect which instructions apply.
Choosing a backend
What is neo4j-agent-memory?
Neo4j Agent Memory provides Python and TypeScript SDKs for persisted conversations, reusable knowledge, and recorded task execution. It is an experimental, community-supported Neo4j Labs project. See the overview and the three memory types.
What is the POLE+O data model?
POLE+O means Person, Object, Location, Event, and Organization. These broad entity categories can be refined with subtypes or domain schemas. See the POLE+O model.
Do I need my own Neo4j to use this package?
You can use hosted NAMS with an authorized workspace and API key, or connect Python directly to a Neo4j database through Bolt. The backends share concepts but support different operations and return shapes. See Backend capabilities and Getting started. Python’s local extraction components can also operate on text independently of storage.
What’s not available on the NAMS backend?
NAMS manages extraction, embeddings, and resolution server-side. Python client-side settings do not configure those services, and several Bolt operations are unsupported on the hosted transport. See the per-operation capability table rather than assuming that an accessor or similarly named method is supported on both backends.
Which SDK should I use — Python or TypeScript?
Choose by language and backend. Python supports Bolt and NAMS; TypeScript targets NAMS REST. Agents can share supported records in the same authorized workspace, but the APIs and capabilities are not identical. See Python, TypeScript, and Backend capabilities.
Which Neo4j versions are supported?
For Bolt, use Neo4j 5.26 or later (including 2025.x releases), or Neo4j AuraDB. The test suite runs against 5.26. Semantic search needs vector indexes (Neo4j 5.11 and later), and merging confirmed duplicate entities (merge_duplicate_entities(), review_duplicate(confirm=True)) uses the CALL (source, target) { … } subquery syntax that Neo4j added in 5.23; older servers reject it as a syntax error. The Python Neo4j driver version is not the database server version. NAMS users do not manage the service’s Neo4j version.
Installing and versions
How do I install the CLI?
Install the package with the CLI extra and any extraction/provider extras needed by the chosen command. Follow CLI installation and commands for the exact package and invocation syntax.
How do I use the CLI?
The CLI supports extraction and related utilities through its implemented commands. Use the CLI reference for arguments, flags, input modes, and output formats; command names from older examples may not exist.
How do I configure the CLI?
The current CLI constructs MemorySettings, which supports the documented environment variables and .env loading. It does not automatically load a .neo4j-memory.yaml configuration file. See CLI configuration and Environment variables.
ImportError: GLiNER2.5 not installed
GLiNER2.5 is optional. Install the gliner2 extra, pip install "neo4j-agent-memory[gliner2]==0.7.0", in the environment running the program; the extraction and full extras include it too, and gliner remains a deprecated alias until 0.8. GLiNER2Extractor loads the gliner2 package on its first extraction, so the ImportError appears there rather than when the extractor is constructed. A broad extra does not necessarily include local models, and installing a Python package does not download every model. See Configure extraction.
Importing a name that 0.7.0 removed, such as GLiNEREntityExtractor, GLiRELExtractor, GLiNERWithRelationsExtractor or is_gliner_available, raises an ImportError that names its replacement. See Migrate to GLiNER2.5.
spaCy model not found
Installing the spaCy package does not necessarily install the configured language model. Install the exact model named by the extraction configuration, then rerun the extraction guide’s verification. See Configure extraction.
GLiNER2.5 model download fails
Check the selected model identifier, access permissions, network access, and any configured offline mode. If using a local model path, confirm the required files are present. See the model setup procedure.
A ValueError that names a GLiNER v1 checkpoint is a different problem: identifiers under urchade/, gliner-community/ or numind/ belong to the removed GLiNER package, and GLiNER2Extractor rejects them before downloading anything. Use a GLiNER2.5 checkpoint such as the default fastino/gliner2.5-base-v1.
Why does MemorySettings(…) raise extra_forbidden for neo4j_uri / openai_api_key?
Releases before 0.2.1 raised this error when a .env file held unprefixed keys such as NEO4J_URI or OPENAI_API_KEY. From 0.2.1, including 0.7.0, MemorySettings ignores .env keys that are not its own settings, so upgrading fixes that case:
pip install --upgrade 'neo4j-agent-memory==0.7.0'
If the error persists on 0.7.0, a keyword argument is misspelled or flattened: MemorySettings(neo4j_uri=…) still raises extra_forbidden, because the field is MemorySettings(neo4j={"uri": …}), or NAM_NEO4J__URI in the environment. Validation stays strict so that typos fail; do not disable it to hide an unknown key. See Settings and Environment variables.
Storing and retrieving memory
What indexes does the package create?
Bolt manages constraints and indexes needed by its configured schema and search features. Index names, dimensions, and requirements belong in Schema objects and Configuration. NAMS manages its own storage indexes.
How are entities labeled in Neo4j?
On Bolt, managed entities carry :Entity, a type property, supported type/subtype labels such as :Person or :Vehicle, and the label the client’s ontology declares for the entity’s exact type and subtype, such as :Customer. See POLE+O and Schema objects for the distinction between stored labels and conceptual categories.
Can I query entities by type directly?
Supported type filters and, on permitted Cypher surfaces, node labels can select entity categories. The query must match the actual backend schema and scope. See Work with entities and Backend capabilities.
What is entity resolution?
Resolution decides whether different mentions refer to the same entity. Similarity is evidence to investigate, not proof of identity. See Resolution and deduplication.
What resolution strategies are available?
Python’s Bolt client selects exact, fuzzy, semantic, or composite matching with resolution.strategy. In 0.7.0 the default, composite, builds an ontology-aware OntologyResolver, which resolves extracted mentions against stored entities on ingest unless resolution.resolve_on_ingest is False. NAMS owns its server-side behavior. See Tune entity resolution, Resolution settings and Backend capabilities.
Is resolution type-aware?
Type constraints can prevent some false matches, but the behavior depends on the resolver and its configuration. OntologyResolver only fetches candidates of the same type, so it never merges a PERSON with a LOCATION, and it also never finds the same company extracted once as ORGANIZATION and once as OBJECT; correct the typing instead. Even same-type entities can be different people or product variants. See Identity evidence and Configure deduplication.
How does entity deduplication on ingest work?
On Bolt in 0.7.0, message ingestion and long_term.add_entity() make the same decision through OntologyResolver when it is the configured resolver, which is the default. ResolutionConfig sets the bands: a score at or above auto_merge_threshold (0.90) reuses the stored entity and keeps the new name as an alias, and a score at or above review_threshold (0.85) creates the entity with a pending SAME_AS edge for review. An ontology can override both per entity type. These are similarity cutoffs, not calibrated correctness probabilities. Set resolution.resolve_on_ingest=False to have message ingestion store one node per distinct extracted name and type instead. There is no deduplication settings section. See Tune entity resolution and the review workflow. NAMS uses its service-side contract.
What are SAME_AS relationships?
A :SAME_AS edge can record a candidate duplicate for review while preserving both records. Its presence is not a completed merge or a guarantee of identity. See the review pattern and Review duplicates.
How do I enable geocoding for LOCATION entities?
Geocoding is a configurable Bolt capability. Use Geocoding configuration and the location API for the actual settings and supported operations. Client-side geocoding settings do not configure NAMS.
What geocoding providers are supported?
The Bolt configuration exposes Nominatim and Google geocoding providers. Their credentials, service policies, quotas, and availability must be checked for the selected deployment; neither provider is universally more accurate. See Geocoding settings.
How do I geocode existing locations?
The Bolt long-term API includes an operation for geocoding stored locations. Its parameters and returned statistics are documented in Long-term memory. Check those fields rather than assuming it shares the batch-extraction signature.
What geospatial queries are supported?
The Bolt long-term API supports location searches such as proximity and bounds, subject to stored coordinates and the method’s scope. See Location operations and Backend support.
Can I disable geocoding for specific entities?
Use the supported per-entity geocoding controls when adding an entity on Bolt. Explicit coordinates and automatic geocoding are separate inputs; use the exact types in the entity API.
Extraction and entities
How does the extraction pipeline work?
Python’s Bolt pipeline runs configured extraction stages and combines their candidates. When an ontology is configured, the pipeline then drops relations the ontology does not permit. Stages differ in vocabulary, resource requirements, and failure modes. See How extraction works for the design and Configure extraction for the procedure.
Can I use just one extractor instead of the full pipeline?
Yes. An extractor can be used independently of a multi-stage pipeline. Choose a component with the entity and relation capabilities your task needs. See Extractor classes and Configure extraction.
What are merge strategies?
Pipeline merge strategies combine candidates from extraction stages. The implemented names are union, intersection, confidence, cascade, and first_success; this differs from deduplicating entities already stored in the graph. See How results combine and the exact API.
Which extractor should I use?
Compare the required labels and relationships, local versus remote processing, resource limits, and quality on representative text. No extractor is universally fastest or most accurate. GLiNER2.5, when the ontology declares relationship types, and LLM extraction can both produce relationships. See Extraction tradeoffs.
Are domain schemas only used with GLiNER2.5?
Python’s DomainSchema gives GLiNER2.5 its labels and descriptions. Since 0.7.0 a domain schema converts into an ontology document, and every stage built from the same ontology reads it: GLiNER2.5 compiles it into its extraction schema, the LLM prompt includes its descriptions and permitted relationships, and spaCy can only map its fixed, model-defined labels onto the ontology’s types. Ontologies can also be stored and activated, on NAMS or in your own database through client.ontology on Bolt. See Domain schemas and Ontologies.
What’s the difference between entity labels and domain schemas?
Labels name the desired entity categories; a domain schema also gives those labels descriptions and can describe relation types. Descriptions clarify the intended vocabulary, but their effect on quality depends on the model and text. See Use extraction schemas.
Can I use domain schemas with the full pipeline?
Yes. With ExtractorBuilder.with_gliner_schema(), the named template becomes the ontology for every stage the builder creates, and the pipeline drops relations that ontology does not permit. It cannot add labels to spaCy’s trained model. See Configure the pipeline and Extractor builder reference.
Which schema should I use?
Choose a vocabulary matching what the application needs to retain, then inspect representative outputs. The domain schema catalog lists available schemas; the schema guide explains how to apply or customize them.
Can I create custom schemas?
Yes. Define the entity labels and descriptions for the GLiNER2.5 task, using the actual DomainSchema fields. See Use custom extraction schemas. To declare typed relationships or validate what is stored, write an ontology document: see ontology-driven extraction on Bolt and NAMS ontologies for a hosted workspace.
Do schema entity types map to POLE+O types?
Extractor label mappings translate selected labels into POLE+O types and subtypes. A mapping does not prove that the model identified the entity correctly. See Schema mappings and POLE+O.
What entity types does spaCy recognize?
The available labels depend on the installed spaCy language model. Common English-model labels include PERSON, ORG, GPE, and LOC; the library maps recognized output into its entity vocabulary. See the spaCy extractor for the integration contract.
Can I use custom entity types with spaCy?
Changing a GLiNER2.5 schema does not change spaCy’s trained NER labels. A custom spaCy model or additional application extraction logic is a separate approach. See the extractor tradeoffs before selecting that path.
Which spaCy model should I use?
Use a model for the document language and label requirements, and compare it on representative text. Download size, runtime memory, speed, and quality vary by model release and workload. Follow the extraction guide for the configured model and installation steps.
When should I use the LLM extractor?
An LLM can help when extraction depends on broader context or a response schema. It adds provider calls and potential failures; it is not guaranteed to outperform local extraction. GLiNER2.5 also extracts relationships locally when the ontology declares relationship types. See Extraction approaches.
What LLM models are supported?
LLMEntityExtractor can use the library’s provider protocols and adapters, including supported native and LiteLLM paths. Model availability and structured-output support depend on the selected provider. See Adapters and Configure an LLM provider.
Can I use local LLMs?
Yes, through a supported provider adapter or a custom provider implementation. A local model must satisfy the interface and extraction requirements of that path; changing an OpenAI environment variable alone is not a universal configuration method. See Bring your own model.
Why is LLM extraction slow?
Generation time, network latency, input size, retries, and provider limits contribute to extraction latency. Compare the required quality against that cost, and use bounded concurrency. See Structured extraction and Batch extraction.
How do I process multiple documents at once?
Use the batch API for the selected extraction or message-storage operation. Its result shape, concurrency controls, and failure behavior depend on that operation. Follow Batch extraction or Batch message processing.
How do I handle very long documents?
Chunk documents when their length or structure exceeds the selected extractor’s useful input window. There is no universal 100K-token cutoff. StreamingExtractor uses character-based chunk_size and overlap by default; chunk_by_tokens=True selects token-based chunking. See Streaming extraction for the supported calls.
Can I extract relationships without LLM calls?
Yes. GLiNER2.5 decodes entities and relationships in one local pass and enforces the ontology’s source and target types while it decodes. It decodes relationships only when the ontology declares relationship types: the poleo, podcast and news templates do, and the other five templates are entity-only until you attach relationships. It needs its own model and compute resources, even though it avoids remote LLM API calls. The separate GLiREL model was removed in 0.7.0. See Extractor classes and Run without an LLM.
Low extraction accuracy
Compare labeled expected entities and relations with actual output. Inspect domain vocabulary, ambiguous names, filtering, and chunk boundaries before changing thresholds. Lower thresholds may improve recall while adding false positives. An empty GLiNER2.5 result with a RuntimeWarning about feasible=False means the decoder could not satisfy the ontology’s hard constraints, such as uniqueness or acyclicity; it does not mean the text has no entities. See Extraction tradeoffs.
Entities being filtered incorrectly
Check the extracted name against the implemented entity validation and stopword rules, then inspect the unfiltered result. See Extractor and validation reference for supported helpers rather than relying on an example’s assumed filtering API.
Memory errors with large texts
Reduce input chunk size or concurrency according to the extractor’s limits, and inspect peak runtime memory. Preserve enough overlap to retain useful context at boundaries. GLiNER2Extractor windows input longer than max_words (384 by default) on its own, because relation recall drops sharply past roughly 400 words; a relationship whose endpoints land in different windows is lost, so extract per message or document section where you can. See Batch and streaming extraction.
Hosted service and previews
What are NAMS Agent Skills? (preview)
Agent Skills is a NAMS preview feature that derives a portable procedure package from scoped memory. Provenance links support review of the generated procedure; they do not establish that every source claim is correct. See Agent Skills and the REST and MCP interface.
Can I run a distilled skill? (preview)
Downloading a published skill and loading it into an agent is separate from server-side distillation. The documented execution feature is gated dry-run planning, not automatic tool execution. See Skills API for its limits and the skill walkthrough for the demonstrated outcome.
Frameworks and integrations
Does neo4j-agent-memory work with AWS?
Python provides a Bedrock embedding adapter, a Strands integration, and AgentCore memory components. Each has its own dependencies and supported behavior; this is not a claim of universal AWS feature support. See Bedrock embeddings, Strands, and Hybrid memory.
How do I configure AWS credentials for Bedrock embeddings?
The Bedrock adapter uses the AWS SDK credential chain. The identity needs permission for the selected model and region. Follow Configure Bedrock embeddings for the supported settings and credential setup.
Should I use Strands tools or the MemoryClient directly?
Use Strands tools when an agent should choose a memory operation through its tool interface. Use MemoryClient when application code should control the operation directly. A tool is only called when the agent or application invokes it. See the Strands integration.
Operations
How do I enable tracing?
The observability module provides tracer objects and span/decorator helpers. Obtaining a tracer does not automatically instrument every memory or model operation. Instrument the application’s selected operations; distinguish these spans from stored reasoning traces and audit records.
What operations are traced?
Only operations connected to the selected instrumentation are traced. Available attributes depend on what the application or adapter records; token counts and cost data are not implied for every operation. See Reasoning traces for the separate persisted task record.
Should I use OpenTelemetry or Opik?
Choose an observability integration that matches the application’s monitoring system and required data. OpenTelemetry provides general tracing interfaces; Opik’s integration targets LLM application observability. The selected instrumentation still determines which spans and attributes exist.
How can I speed up extraction?
Measure which stage dominates the actual workload, then evaluate changes to enabled stages, model choice, batching, and concurrency. For GLiNER2.5, the options include the smaller fastino/gliner2.5-small-v1 checkpoint, a GPU device, quantize=True (fp16) and compile=True. Track quality as well as runtime. See Extraction tradeoffs and Batch and streaming extraction.
What’s the memory footprint?
Runtime memory depends on the installed model, precision, device, input sizes, and concurrency. Model download size is not the same as peak process memory. Measure the configured environment; the docs do not provide a portable RAM estimate for every extractor. As download sizes, the GLiNER2.5 checkpoints are about 296 MB for fastino/gliner2.5-small-v1, 407 MB for fastino/gliner2.5-base-v1 (the default) and 594 MB for the multilingual fastino/gliner2.5-multi-v1.
Can I cache extraction results?
Application-level caching can reuse completed extraction results when the text, model, schema, and configuration are unchanged. Cache the completed value, not an asynchronous coroutine object: functools.lru_cache applied directly to an async def does not provide a safe reusable result cache. Define invalidation and size limits for the application’s workload.