Why the provider protocol?
The Python provider interfaces separate model access from memory storage. They let client-side extraction and embeddings use a selected adapter while keeping a common interface for the caller.
These are Bolt-side provider choices. NAMS manages extraction and embeddings server-side. See Bolt and NAMS. This page describes neo4j-agent-memory 0.7.0; check the version you have installed with pip show neo4j-agent-memory.
|
Why separate model access from extraction?
Coupling an extractor directly to one provider client makes it harder to use another hosted model, a local model, or a test double. It also spreads provider-specific credentials and errors through extraction code.
The protocol design moves that responsibility into adapters. The extractor can request a completion or structured result while the adapter handles its model provider’s interface. Compatibility work still requires tests; the abstraction is not a blanket promise that every older application behaves identically.
Three distinct capabilities
LLMProvider supplies chat completions, StructuredExtractor supplies validated structured output, and EmbeddingProvider supplies vectors. These are different operations with different inputs and failure modes.
Separating them avoids forcing an embedding-only adapter to expose a meaningless completion method. It also allows a structured-only adapter to participate without implementing ordinary chat completion.
Python’s structural protocols allow compatible objects without requiring inheritance from a package base class. A runtime shape check is useful for dispatch, but it does not prove behavioral conformance. The provider reference defines the full interfaces.
Native-first resolution
The factory resolves a model string using an available native adapter when supported, with LiteLLM as a fallback for supported routes. prefer_litellm=True allows an application to choose that path explicitly. Available dependencies and model identifiers affect the result.
This preserves access to adapter-specific capabilities while offering a broader routing option. It is a design preference, not an empirical ranking of response quality. See Factory reference for resolution rules and missing-dependency errors.
Schema-aligned retry as a fallback
When a native structured-output path is unavailable, schema-aligned retry uses the completion interface, parses the result, validates it, and returns focused feedback for a bounded retry after structural failures.
The helper does not promise that the second attempt succeeds or that a valid result is factually accurate. An optional Instructor adapter provides another structured-output path. See Structured extraction for the tradeoffs.
Why provider errors are separate
Provider failures concern model or embedding calls. Storage failures concern the memory backend. Separate exception hierarchies let an application decide whether it should retry a provider request, correct credentials, or investigate a database operation.
Adapters translate supported upstream errors into the provider hierarchy. Applications should follow the actual error contract rather than assuming all errors are transient. See Provider errors.
Why embedding dimensions are explicit
Vector indexes and their vectors must have compatible dimensions. A mismatch can make similarity search unusable, so dimension information needs to be available when connecting and checking the managed indexes.
The provider exposes its dimension count, and supported connection checks can raise EmbeddingDimensionMismatchError for incompatible existing indexes. Matching dimensions alone does not make vectors from different models semantically compatible. A model change also needs a migration and re-embedding policy; see Migrate an embedding model.
Why keep legacy configuration during migration?
A compatibility path gives applications time to move from legacy configuration objects to provider strings or instances. Strings express a common model choice concisely; instances allow adapter-specific configuration and custom implementations.
Deprecation behavior and supported releases belong in the migration guide. This explanation does not promise a future removal schedule or backward compatibility beyond the documented release contract.
Boundaries of the interface
The completion protocol returns a completed result; it is not a general model-serving gateway. Routing fleets of models, application caching, and a framework’s streaming lifecycle are separate concerns unless a specific adapter documents otherwise.
Keeping the interface bounded reduces what each adapter must implement. It also means callers must not infer a capability from another provider SDK’s similarly named feature. Consult the reference instead of treating possible future work as an available API.
Relationship to conformance testing
A small interface can be checked using controlled providers and repeatable behavioral cases. This separates protocol conformance from model quality and from the operation of a live hosted service.
Bronze, Silver, Gold, and Platinum terminology refers to client conformance in the TCK context. These are not NAMS pricing or access tiers. Implementing a protocol does not by itself certify an adapter or a release.
How migration resolves provider settings
The migration from legacy EmbeddingConfig / LLMConfig types to provider strings and instances is implemented by two @model_validator(mode="after") validators on MemorySettings. The first, _resolve_providers:
-
Coerces dict input to legacy configs (so
MemorySettings(embedding={…})still works). -
Resolves provider strings via
from_provider. -
Emits
DeprecationWarningwhen the user explicitly passed legacyEmbeddingConfig/LLMConfiginstances. -
Validates Provider instances against the runtime-checkable
EmbeddingProvider/LLMProviderProtocols. -
Passes
llm=Nonethrough unchanged.
The second, _validate_llm_consistency, then checks llm=None against the extraction settings. When the configured extractor needs an LLM (extractor_type=LLM, or PIPELINE with enable_llm_fallback=True), an explicit llm=None makes construction fail with a ValidationError, and an omitted llm is filled in with a default LLMConfig(). Otherwise llm=None is kept.
After resolution, MemoryClient._create_embedder and _create_extractor consume either shape transparently. A small adapter (_ProviderToEmbedderAdapter) bridges new Provider instances back to the legacy Embedder API so downstream memory layers keep working unchanged.
See Migrate to pluggable providers for the migration steps and deprecation-warning handling.
Design tradeoffs
The interfaces provide a common call shape while preserving separate capabilities for completion, structured extraction, and embedding. Native adapters, LiteLLM, and custom providers can participate through the appropriate surface.
The cost is that callers must understand which capabilities their selected adapter supports. Explicit dimensions, exceptions, and structured-output contracts make those boundaries inspectable.