Embedding, LLM, schema, extraction, and resolution settings
Provider and extraction-pipeline settings for Configuration reference: EmbeddingConfig, LLMConfig, SchemaConfig, ExtractionConfig, and ResolutionConfig. Fields marked † are not automatically applied by MemoryClient; see integration limits.
Embedding configuration
Legacy field inventory. Provider strings and instances use their own adapter signatures, not these nested keys; see Adapters. The enum also declares anthropic, vertex_ai, bedrock, and custom, but MemoryClient’s legacy factory currently constructs only openai and sentence_transformers. For Vertex AI or Bedrock, use a supported provider instance/string or inject an embedder explicitly.
When a legacy Vertex AI config is constructed, its validator changes an unset model to gemini-embedding-001 and unset output dimensionality/dimensions to 768; retired IDs known to the SDK are rejected. This validation does not wire the config into the legacy MemoryClient factory. Index dimensions must match the actual embedder; see Migrate embedding models.
| Field | Type / values | Default | Constraints | Env var | Description |
|---|---|---|---|---|---|
|
|
|
— |
|
Embedding provider to use |
|
|
|
— |
|
Embedding model name |
|
|
|
ge 1 |
|
Embedding dimensions |
|
|
|
— |
|
API key for embedding provider |
|
|
|
ge 1 |
|
Batch size for embeddings |
|
|
|
— |
|
Device for sentence transformers (cpu/cuda) |
|
|
|
— |
|
GCP project ID for Vertex AI |
|
|
|
— |
|
GCP region for Vertex AI |
|
|
|
— |
|
Vertex AI task type (RETRIEVAL_QUERY, RETRIEVAL_DOCUMENT, etc.) |
|
|
|
ge 1 |
|
Vertex AI output dimensionality. gemini-embedding-001 emits 3072 dimensions natively and can be truncated to 1536 or 768. Defaults to 768 for Vertex AI so existing vector indexes keep working; pass an explicit value to override. |
|
|
|
— |
|
AWS region for Bedrock |
|
|
|
— |
|
AWS credentials profile name |
LLM configuration
Legacy LLM field inventory. The extraction factory forwards the selected provider/model/API key; it does not forward these legacy temperature/max_tokens fields. Use the extractor/provider APIs for their supported generation controls. extraction.llm_model is the fallback only when the factory receives no LLM config/provider.
| Field | Type / values | Default | Constraints | Env var | Description |
|---|---|---|---|---|---|
|
|
|
— |
|
LLM provider to use |
|
|
|
— |
|
LLM model name |
|
|
|
— |
|
API key for LLM provider |
|
|
|
ge 0.0, le 2.0 |
|
LLM temperature † |
|
|
|
ge 1 |
|
Maximum tokens for LLM † |
† Not automatically applied by MemoryClient; see integration limits.
Graph schema configuration
Bolt only. SchemaConfig is not a persisted EntitySchemaConfig or a NAMS ontology: it selects the one ontology document that the extractors, relation validation, and the write paths share, and how strictly that document is enforced. The client resolves the document once per connection, in this order; the first source that yields one wins:
-
the
ontology=keyword onMemoryClient(a document or a file path), -
ontology_path, -
custom_schema_path, -
the active stored ontology version, when
use_active_ontologyis True and a version is activated (client.ontology.activate), -
model="custom"withentity_types, which builds an ad-hoc document with one label per name, -
the built-in template named by
ontology_template, falling back to POLE+O when the name is unknown.
Both path fields are loaded through neo4j_agent_memory.ontology.load_ontology, so a .json/.yaml file may hold an EntitySchemaConfig or an OntologyDocument; either is converted on the way in. Because resolution happens at connect time, activating a stored version affects the next connection, not the client that activated it. Read the result back with client.ontology_document and client.validation_mode; see Ontology precedence.
The validation mode is validation_mode when set, then the mode stored on the active version when that version supplied the document, then strict if strict_types is True, otherwise permissive. In both modes, extracted relations the ontology does not permit are dropped before storage. strict also drops undeclared entity types on message ingestion and makes long_term.add_entity and add_relationship raise ValidationError. See Drive extraction from an ontology and Ontology API.
from neo4j_agent_memory.config.settings import SchemaConfig
schema_config = SchemaConfig(
ontology_path="ontology.yaml", # loaded at connect time
validation_mode="strict",
)
| Field | Type / values | Default | Constraints | Env var | Description |
|---|---|---|---|---|---|
|
|
|
— |
|
Schema model. |
|
|
|
— |
|
Custom entity types, used when |
|
|
|
— |
|
Whether to track entity subtypes † |
|
|
|
— |
|
Selects |
|
|
|
— |
|
Path to a schema file ( |
|
|
|
— |
|
Path to an ontology document ( |
|
|
|
— |
|
Adopt the ontology version activated in the database when no file or explicit document is given. |
|
|
|
— |
|
Built-in template used as the final fallback: |
|
|
|
— |
|
Ontology enforcement on the write paths. |
|
|
|
— |
|
Whether |
† Not automatically applied by MemoryClient; see integration limits.
Extraction configuration
Bolt extraction only. The default pipeline attempts spaCy, GLiNER2.5, and LLM in order. The factory skips a stage whose package is unavailable (the GLiNER2.5 stage needs the gliner2 extra) and uses NoOpExtractor if none can be constructed; test your configured stages and their output.
The gliner_* field names and NAM_EXTRACTION__GLINER_* variables keep their spelling, but since 0.7 they configure GLiNER2.5: the fastino/gliner2.5-{small,base,multi}-v1 checkpoints on the gliner2 package, which decode entities and typed relations in one joint pass. A GLiNER v1 checkpoint id (urchade/…, gliner-community/…, numind/…) passes settings validation but is rejected with a migration hint when the extractor is constructed. In a pipeline every stage, and the pipeline’s own relation validation, works from one ontology: the built-in template named by gliner_schema when set, otherwise the client’s resolved ontology (see Graph schema configuration), otherwise POLE+O. A single gliner extractor resolves its ontology the same way; a single llm or spacy extractor uses the client’s resolved ontology and ignores gliner_schema. For a gliner or pipeline extractor that MemoryClient builds, gliner_schema also becomes the client’s resolved ontology, so ingest-time relation validation, strict mode and resolution use the template the extractor decodes against. Relations are decoded only when that ontology declares relationship types. entity_types is the LLM extractor’s type list (replaced by schema_config.entity_types when that is set); GLiNER2.5 labels come from the ontology.
merge_strategy accepts union, intersection, confidence, and cascade. Pipeline construction sets stop_on_success to the inverse of fallback_on_empty; with the default True, enabled stages can all run. FIRST_SUCCESS belongs only to the direct pipeline enum, not this settings enum. confidence_threshold filters the pipeline’s merged result (entities and relations below it are dropped, with any relation left without an endpoint); it is distinct from gliner_threshold, the floor the GLiNER2.5 decoder applies. There are no batch or streaming chunk fields in ExtractionConfig: pass those to the corresponding extraction methods.
from neo4j_agent_memory.config.settings import ExtractionConfig
extraction = ExtractionConfig(
extractor_type="pipeline", enable_spacy=True, enable_gliner=True,
enable_llm_fallback=False, gliner_schema="podcast",
merge_strategy="confidence", gliner_threshold=0.6,
)
| Field | Type / values | Default | Constraints | Env var | Description |
|---|---|---|---|---|---|
|
|
|
— |
|
Type of entity extractor. |
|
|
|
— |
|
Enable spaCy in extraction pipeline |
|
|
|
— |
|
Enable the GLiNER2.5 stage in the extraction pipeline |
|
|
|
— |
|
Enable LLM as fallback in pipeline |
|
|
|
— |
|
Strategy for merging results from multiple extractors |
|
|
|
— |
|
Continue to next stage if current stage returns no results |
|
|
|
— |
|
spaCy model name |
|
|
|
ge 0.0, le 1.0 |
|
Default confidence score for spaCy extractions |
|
|
|
— |
|
GLiNER2.5 checkpoint id or local path: |
|
|
|
ge 0.0, le 1.0 |
|
GLiNER2.5 entity confidence threshold |
|
|
|
ge 0.0, le 1.0 |
|
Confidence floor for decoded relations; |
|
|
|
— |
|
Device for the GLiNER2.5 model (cpu/cuda/mps) |
|
|
|
— |
|
Built-in domain template (poleo, podcast, news, scientific, business, entertainment, medical, legal). When set, it replaces the resolved ontology for the GLiNER2.5 extractor and for every stage of a pipeline. |
|
|
|
gt 0 |
|
Longest input decoded in one pass; longer text is windowed |
|
|
|
ge 0 |
|
Word overlap between windows for long input |
|
|
|
— |
|
Span-overlap policy for the attribute pass ( |
|
|
|
— |
|
Run the opt-in second pass that recovers the ontology’s enum properties as entity attributes |
|
|
|
— |
|
Load the GLiNER2.5 weights in fp16 |
|
|
|
— |
|
Run the GLiNER2.5 weights through |
|
|
|
— |
|
LLM model for extraction |
|
|
|
— |
|
Entity types for the LLM extractor (POLE+O by default) |
|
|
|
— |
|
Whether to extract relations |
|
|
|
— |
|
Whether to extract preferences |
|
|
|
ge 0.0, le 1.0 |
|
Floor applied to the merged result of |
Resolution configuration
Bolt resolution factory: none disables resolution; exact, fuzzy, and semantic select one resolver; composite, the default, builds OntologyResolver. It blocks candidates in Cypher without crossing entity types, uses the ontology’s aliases and per-type thresholds, and sorts each mention into one of three bands: merged at or above auto_merge_threshold, review (a new node plus a pending SAME_AS edge) at or above review_threshold, and created below that. Fuzzy and semantic thresholds default to 0.85 and 0.8 and apply to the fuzzy and semantic resolvers; OntologyResolver scores with its own weights and uses only the two band thresholds. The accepted fuzzy_scorer field is not forwarded by MemoryClient; construct/inject a resolver for that control.
Since 0.7, resolution runs when a message is stored: add_message, add_messages_batch, and extract_entities_from_session resolve the mentions they extract, and long_term.add_entity delegates its duplicate check to the same resolver, so both write paths band identically; see deduplication. Only an OntologyResolver resolves on ingest; resolve_on_ingest=False is the single opt-out. A resolution failure is logged and the mentions are stored unresolved. See Tune entity resolution.
from neo4j_agent_memory.config.settings import ResolutionConfig
resolution = ResolutionConfig(
auto_merge_threshold=0.92, review_threshold=0.85, scope="user",
)
| Field | Type / values | Default | Constraints | Env var | Description |
|---|---|---|---|---|---|
|
|
|
— |
|
Resolution strategy. |
|
|
|
ge 0.0, le 1.0 |
|
Exact match threshold |
|
|
|
ge 0.0, le 1.0 |
|
Fuzzy match threshold |
|
|
|
ge 0.0, le 1.0 |
|
Semantic match threshold |
|
|
|
— |
|
Fuzzy matching scorer † |
|
|
|
— |
|
Resolve extracted mentions against stored entities while storing a message. False restores the pre-0.7 behavior: one node per distinct surface form. |
|
|
|
ge 0.0, le 1.0 |
|
Score at or above which a mention merges onto the matched entity. Per-type override: |
|
|
|
ge 0.0, le 1.0 |
|
Score at or above which a mention is stored as its own node with a pending |
|
|
|
ge 1 |
|
Maximum blocking candidates fetched per mention and per bucket |
|
|
|
— |
|
Use |
|
|
|
— |
|
Also generate candidates from the entity vector index. Needs an embedder; without one the embedding score drops out and the remaining weights are renormalized. |
|
|
|
ge 0 |
|
Half-width of the mention context window compared during scoring |
|
|
|
— |
|
|
† Not automatically applied by MemoryClient; see integration limits.