Embedding, LLM, schema, extraction, and resolution settings

Provider and extraction-pipeline settings for Configuration reference: EmbeddingConfig, LLMConfig, SchemaConfig, ExtractionConfig, and ResolutionConfig. Fields marked † are not automatically applied by MemoryClient; see integration limits.

Embedding configuration

Legacy field inventory. Provider strings and instances use their own adapter signatures, not these nested keys; see Adapters. The enum also declares anthropic, vertex_ai, bedrock, and custom, but MemoryClient’s legacy factory currently constructs only openai and sentence_transformers. For Vertex AI or Bedrock, use a supported provider instance/string or inject an embedder explicitly.

When a legacy Vertex AI config is constructed, its validator changes an unset model to gemini-embedding-001 and unset output dimensionality/dimensions to 768; retired IDs known to the SDK are rejected. This validation does not wire the config into the legacy MemoryClient factory. Index dimensions must match the actual embedder; see Migrate embedding models.

Field Type / values Default Constraints Env var Description

provider

openai / anthropic / sentence_transformers / vertex_ai / bedrock / custom

'openai'

—

NAM_EMBEDDING__PROVIDER

Embedding provider to use

model

str

'text-embedding-3-small'

—

NAM_EMBEDDING__MODEL

Embedding model name

dimensions

int

1536

ge 1

NAM_EMBEDDING__DIMENSIONS

Embedding dimensions

api_key

SecretStr | None

None

—

NAM_EMBEDDING__API_KEY

API key for embedding provider

batch_size

int

100

ge 1

NAM_EMBEDDING__BATCH_SIZE

Batch size for embeddings

device

str

'cpu'

—

NAM_EMBEDDING__DEVICE

Device for sentence transformers (cpu/cuda)

project_id

str | None

None

—

NAM_EMBEDDING__PROJECT_ID

GCP project ID for Vertex AI

location

str

'us-central1'

—

NAM_EMBEDDING__LOCATION

GCP region for Vertex AI

task_type

str

'RETRIEVAL_DOCUMENT'

—

NAM_EMBEDDING__TASK_TYPE

Vertex AI task type (RETRIEVAL_QUERY, RETRIEVAL_DOCUMENT, etc.)

output_dimensionality

int | None

None

ge 1

NAM_EMBEDDING__OUTPUT_DIMENSIONALITY

Vertex AI output dimensionality. gemini-embedding-001 emits 3072 dimensions natively and can be truncated to 1536 or 768. Defaults to 768 for Vertex AI so existing vector indexes keep working; pass an explicit value to override.

aws_region

str | None

None

—

NAM_EMBEDDING__AWS_REGION

AWS region for Bedrock

aws_profile

str | None

None

—

NAM_EMBEDDING__AWS_PROFILE

AWS credentials profile name

LLM configuration

Legacy LLM field inventory. The extraction factory forwards the selected provider/model/API key; it does not forward these legacy temperature/max_tokens fields. Use the extractor/provider APIs for their supported generation controls. extraction.llm_model is the fallback only when the factory receives no LLM config/provider.

Field Type / values Default Constraints Env var Description

provider

openai / anthropic / custom

'openai'

—

NAM_LLM__PROVIDER

LLM provider to use

model

str

'gpt-4o-mini'

—

NAM_LLM__MODEL

LLM model name

api_key

SecretStr | None

None

—

NAM_LLM__API_KEY

API key for LLM provider

temperature

float

0.0

ge 0.0, le 2.0

NAM_LLM__TEMPERATURE

LLM temperature †

max_tokens

int

4096

ge 1

NAM_LLM__MAX_TOKENS

Maximum tokens for LLM †

† Not automatically applied by MemoryClient; see integration limits.

Graph schema configuration

Bolt only. SchemaConfig is not a persisted EntitySchemaConfig or a NAMS ontology: it selects the one ontology document that the extractors, relation validation, and the write paths share, and how strictly that document is enforced. The client resolves the document once per connection, in this order; the first source that yields one wins:

  1. the ontology= keyword on MemoryClient (a document or a file path),

  2. ontology_path,

  3. custom_schema_path,

  4. the active stored ontology version, when use_active_ontology is True and a version is activated (client.ontology.activate),

  5. model="custom" with entity_types, which builds an ad-hoc document with one label per name,

  6. the built-in template named by ontology_template, falling back to POLE+O when the name is unknown.

Both path fields are loaded through neo4j_agent_memory.ontology.load_ontology, so a .json/.yaml file may hold an EntitySchemaConfig or an OntologyDocument; either is converted on the way in. Because resolution happens at connect time, activating a stored version affects the next connection, not the client that activated it. Read the result back with client.ontology_document and client.validation_mode; see Ontology precedence.

The validation mode is validation_mode when set, then the mode stored on the active version when that version supplied the document, then strict if strict_types is True, otherwise permissive. In both modes, extracted relations the ontology does not permit are dropped before storage. strict also drops undeclared entity types on message ingestion and makes long_term.add_entity and add_relationship raise ValidationError. See Drive extraction from an ontology and Ontology API.

from neo4j_agent_memory.config.settings import SchemaConfig

schema_config = SchemaConfig(
    ontology_path="ontology.yaml",  # loaded at connect time
    validation_mode="strict",
)
Field Type / values Default Constraints Env var Description

model

poleo / legacy / custom

'poleo'

—

NAM_SCHEMA_CONFIG__MODEL

Schema model. custom together with entity_types builds an ad-hoc ontology; the other values fall through to ontology_template.

entity_types

list[str] | None

None

—

NAM_SCHEMA_CONFIG__ENTITY_TYPES

Custom entity types, used when model=custom. When set, they also replace extraction.entity_types as the LLM extractor’s type list.

enable_subtypes

bool

True

—

NAM_SCHEMA_CONFIG__ENABLE_SUBTYPES

Whether to track entity subtypes †

strict_types

bool

False

—

NAM_SCHEMA_CONFIG__STRICT_TYPES

Selects strict validation when neither validation_mode nor the active stored version sets a mode.

custom_schema_path

str | None

None

—

NAM_SCHEMA_CONFIG__CUSTOM_SCHEMA_PATH

Path to a schema file (.json/.yaml) holding an EntitySchemaConfig or an OntologyDocument. Consulted after ontology_path.

ontology_path

str | None

None

—

NAM_SCHEMA_CONFIG__ONTOLOGY_PATH

Path to an ontology document (.json/.yaml). The highest-priority source after the MemoryClient(ontology=…​) keyword.

use_active_ontology

bool

True

—

NAM_SCHEMA_CONFIG__USE_ACTIVE_ONTOLOGY

Adopt the ontology version activated in the database when no file or explicit document is given.

ontology_template

str

'poleo'

—

NAM_SCHEMA_CONFIG__ONTOLOGY_TEMPLATE

Built-in template used as the final fallback: poleo, podcast, news, scientific, business, entertainment, medical, or legal.

validation_mode

Literal['permissive', 'strict'] | None

None

—

NAM_SCHEMA_CONFIG__VALIDATION_MODE

Ontology enforcement on the write paths. None derives the mode as described above.

backfill_relation_types

bool

True

—

NAM_SCHEMA_CONFIG__BACKFILL_RELATION_TYPES

Whether connect() may run the one-time backfills: copying the legacy RELATED_TO.relation_type property onto r.type, and filling the entity lookup keys (name_key, surface_keys) that entity resolution reads. Each runs once per database and is recorded on a :SchemaMigration marker node; set False to run the equivalent queries yourself in batches on a very large graph.

† Not automatically applied by MemoryClient; see integration limits.

Extraction configuration

Bolt extraction only. The default pipeline attempts spaCy, GLiNER2.5, and LLM in order. The factory skips a stage whose package is unavailable (the GLiNER2.5 stage needs the gliner2 extra) and uses NoOpExtractor if none can be constructed; test your configured stages and their output.

The gliner_* field names and NAM_EXTRACTION__GLINER_* variables keep their spelling, but since 0.7 they configure GLiNER2.5: the fastino/gliner2.5-{small,base,multi}-v1 checkpoints on the gliner2 package, which decode entities and typed relations in one joint pass. A GLiNER v1 checkpoint id (urchade/…​, gliner-community/…​, numind/…​) passes settings validation but is rejected with a migration hint when the extractor is constructed. In a pipeline every stage, and the pipeline’s own relation validation, works from one ontology: the built-in template named by gliner_schema when set, otherwise the client’s resolved ontology (see Graph schema configuration), otherwise POLE+O. A single gliner extractor resolves its ontology the same way; a single llm or spacy extractor uses the client’s resolved ontology and ignores gliner_schema. For a gliner or pipeline extractor that MemoryClient builds, gliner_schema also becomes the client’s resolved ontology, so ingest-time relation validation, strict mode and resolution use the template the extractor decodes against. Relations are decoded only when that ontology declares relationship types. entity_types is the LLM extractor’s type list (replaced by schema_config.entity_types when that is set); GLiNER2.5 labels come from the ontology.

merge_strategy accepts union, intersection, confidence, and cascade. Pipeline construction sets stop_on_success to the inverse of fallback_on_empty; with the default True, enabled stages can all run. FIRST_SUCCESS belongs only to the direct pipeline enum, not this settings enum. confidence_threshold filters the pipeline’s merged result (entities and relations below it are dropped, with any relation left without an endpoint); it is distinct from gliner_threshold, the floor the GLiNER2.5 decoder applies. There are no batch or streaming chunk fields in ExtractionConfig: pass those to the corresponding extraction methods.

from neo4j_agent_memory.config.settings import ExtractionConfig

extraction = ExtractionConfig(
    extractor_type="pipeline", enable_spacy=True, enable_gliner=True,
    enable_llm_fallback=False, gliner_schema="podcast",
    merge_strategy="confidence", gliner_threshold=0.6,
)
Field Type / values Default Constraints Env var Description

extractor_type

llm / gliner / spacy / pipeline / none

'pipeline'

—

NAM_EXTRACTION__EXTRACTOR_TYPE

Type of entity extractor. gliner is GLiNER2.5 alone.

enable_spacy

bool

True

—

NAM_EXTRACTION__ENABLE_SPACY

Enable spaCy in extraction pipeline

enable_gliner

bool

True

—

NAM_EXTRACTION__ENABLE_GLINER

Enable the GLiNER2.5 stage in the extraction pipeline

enable_llm_fallback

bool

True

—

NAM_EXTRACTION__ENABLE_LLM_FALLBACK

Enable LLM as fallback in pipeline

merge_strategy

union / intersection / confidence / cascade

'confidence'

—

NAM_EXTRACTION__MERGE_STRATEGY

Strategy for merging results from multiple extractors

fallback_on_empty

bool

True

—

NAM_EXTRACTION__FALLBACK_ON_EMPTY

Continue to next stage if current stage returns no results

spacy_model

str

'en_core_web_sm'

—

NAM_EXTRACTION__SPACY_MODEL

spaCy model name

spacy_confidence

float

0.85

ge 0.0, le 1.0

NAM_EXTRACTION__SPACY_CONFIDENCE

Default confidence score for spaCy extractions

gliner_model

str

'fastino/gliner2.5-base-v1'

—

NAM_EXTRACTION__GLINER_MODEL

GLiNER2.5 checkpoint id or local path: fastino/gliner2.5-small-v1, -base-v1, or -multi-v1

gliner_threshold

float

0.5

ge 0.0, le 1.0

NAM_EXTRACTION__GLINER_THRESHOLD

GLiNER2.5 entity confidence threshold

gliner_relation_threshold

float | None

None

ge 0.0, le 1.0

NAM_EXTRACTION__GLINER_RELATION_THRESHOLD

Confidence floor for decoded relations; None keeps whatever the ontology’s per-relationship thresholds selected

gliner_device

str

'cpu'

—

NAM_EXTRACTION__GLINER_DEVICE

Device for the GLiNER2.5 model (cpu/cuda/mps)

gliner_schema

str | None

None

—

NAM_EXTRACTION__GLINER_SCHEMA

Built-in domain template (poleo, podcast, news, scientific, business, entertainment, medical, legal). When set, it replaces the resolved ontology for the GLiNER2.5 extractor and for every stage of a pipeline.

gliner_max_words

int

384

gt 0

NAM_EXTRACTION__GLINER_MAX_WORDS

Longest input decoded in one pass; longer text is windowed

gliner_chunk_overlap

int

64

ge 0

NAM_EXTRACTION__GLINER_CHUNK_OVERLAP

Word overlap between windows for long input

gliner_overlap_policy

str | None

None

—

NAM_EXTRACTION__GLINER_OVERLAP_POLICY

Span-overlap policy for the attribute pass (flat / nested / allow / longest); None keeps the checkpoint default

gliner_extract_attributes

bool

False

—

NAM_EXTRACTION__GLINER_EXTRACT_ATTRIBUTES

Run the opt-in second pass that recovers the ontology’s enum properties as entity attributes

gliner_quantize

bool

False

—

NAM_EXTRACTION__GLINER_QUANTIZE

Load the GLiNER2.5 weights in fp16

gliner_compile

bool

False

—

NAM_EXTRACTION__GLINER_COMPILE

Run the GLiNER2.5 weights through torch.compile

llm_model

str

'gpt-4o-mini'

—

NAM_EXTRACTION__LLM_MODEL

LLM model for extraction

entity_types

list[str]

["PERSON", "ORGANIZATION", "LOCATION", "EVENT", "OBJECT"]

—

NAM_EXTRACTION__ENTITY_TYPES

Entity types for the LLM extractor (POLE+O by default)

extract_relations

bool

True

—

NAM_EXTRACTION__EXTRACT_RELATIONS

Whether to extract relations

extract_preferences

bool

True

—

NAM_EXTRACTION__EXTRACT_PREFERENCES

Whether to extract preferences

confidence_threshold

float

0.5

ge 0.0, le 1.0

NAM_EXTRACTION__CONFIDENCE_THRESHOLD

Floor applied to the merged result of extractor_type="pipeline". A single-extractor type does not apply it.

Resolution configuration

Bolt resolution factory: none disables resolution; exact, fuzzy, and semantic select one resolver; composite, the default, builds OntologyResolver. It blocks candidates in Cypher without crossing entity types, uses the ontology’s aliases and per-type thresholds, and sorts each mention into one of three bands: merged at or above auto_merge_threshold, review (a new node plus a pending SAME_AS edge) at or above review_threshold, and created below that. Fuzzy and semantic thresholds default to 0.85 and 0.8 and apply to the fuzzy and semantic resolvers; OntologyResolver scores with its own weights and uses only the two band thresholds. The accepted fuzzy_scorer field is not forwarded by MemoryClient; construct/inject a resolver for that control.

Since 0.7, resolution runs when a message is stored: add_message, add_messages_batch, and extract_entities_from_session resolve the mentions they extract, and long_term.add_entity delegates its duplicate check to the same resolver, so both write paths band identically; see deduplication. Only an OntologyResolver resolves on ingest; resolve_on_ingest=False is the single opt-out. A resolution failure is logged and the mentions are stored unresolved. See Tune entity resolution.

from neo4j_agent_memory.config.settings import ResolutionConfig

resolution = ResolutionConfig(
    auto_merge_threshold=0.92, review_threshold=0.85, scope="user",
)
Field Type / values Default Constraints Env var Description

strategy

exact / fuzzy / semantic / composite / none

'composite'

—

NAM_RESOLUTION__STRATEGY

Resolution strategy. composite builds OntologyResolver on bolt.

exact_threshold

float

1.0

ge 0.0, le 1.0

NAM_RESOLUTION__EXACT_THRESHOLD

Exact match threshold

fuzzy_threshold

float

0.85

ge 0.0, le 1.0

NAM_RESOLUTION__FUZZY_THRESHOLD

Fuzzy match threshold

semantic_threshold

float

0.8

ge 0.0, le 1.0

NAM_RESOLUTION__SEMANTIC_THRESHOLD

Semantic match threshold

fuzzy_scorer

str

'token_sort_ratio'

—

NAM_RESOLUTION__FUZZY_SCORER

Fuzzy matching scorer †

resolve_on_ingest

bool

True

—

NAM_RESOLUTION__RESOLVE_ON_INGEST

Resolve extracted mentions against stored entities while storing a message. False restores the pre-0.7 behavior: one node per distinct surface form.

auto_merge_threshold

float

0.9

ge 0.0, le 1.0

NAM_RESOLUTION__AUTO_MERGE_THRESHOLD

Score at or above which a mention merges onto the matched entity. Per-type override: EntityTypeDef.resolution_threshold.

review_threshold

float

0.85

ge 0.0, le 1.0

NAM_RESOLUTION__REVIEW_THRESHOLD

Score at or above which a mention is stored as its own node with a pending SAME_AS edge to the match. Per-type override: EntityTypeDef.review_threshold. Must not exceed auto_merge_threshold.

candidate_limit

int

12

ge 1

NAM_RESOLUTION__CANDIDATE_LIMIT

Maximum blocking candidates fetched per mention and per bucket

use_alias_gazetteer

bool

True

—

NAM_RESOLUTION__USE_ALIAS_GAZETTEER

Use EntityTypeDef.aliases as blocking keys and as a 1.0 match rule

use_embedding_blocking

bool

True

—

NAM_RESOLUTION__USE_EMBEDDING_BLOCKING

Also generate candidates from the entity vector index. Needs an embedder; without one the embedding score drops out and the remaining weights are renormalized.

context_window_chars

int

90

ge 0

NAM_RESOLUTION__CONTEXT_WINDOW_CHARS

Half-width of the mention context window compared during scoring

scope

Literal['global', 'user']

'global'

—

NAM_RESOLUTION__SCOPE

global matches against every stored entity; user restricts candidates to entities the tenant has mentioned.

† Not automatically applied by MemoryClient; see integration limits.