Bring your own model
|
Available on NAMS: No. The procedure on this page needs the Bolt backend. With the hosted NAMS backend its operations are unavailable: depending on the call, the client raises |
Configure an LLM or embedding provider — native adapters for OpenAI, Anthropic, Bedrock, Vertex AI, and sentence-transformers; LiteLLM universal fallback for everything else.
neo4j-agent-memory is provider-pluggable. You pass either a provider-string shorthand ("anthropic/claude-3-5-sonnet-latest"), an explicit Provider instance, or hand off your already-configured framework model. The factory chooses an installed native adapter when supported, then falls back to LiteLLM when available.
Prerequisites
-
Install the extras for each selected provider and configure its credentials. Quote package extras in shell commands, for example
pip install 'neo4j-agent-memory[openai]==0.7.0'. -
Use a dedicated AuraDB instance and export
NEO4J_URI,NEO4J_USERNAME, andNEO4J_PASSWORDas shown in the Aura connection setup. These examples use theboltbackend;NEO4J_DATABASEdefaults toneo4jwhen omitted. -
Model identifiers below illustrate adapter routing. Confirm the chosen model is available to your account before making a live request; successful provider construction does not verify model access.
-
Run snippets containing
awaitin an async function. Theprovider,messages, and connectedclientin later fragments come from your application’s setup.
Procedure
-
Choose a string, explicit adapter, or framework pass-through from the options below.
-
Configure both the LLM and embedding provider; changing one does not configure the other.
-
Run the verification at the end before applying the configuration to an existing graph.
Configure a provider string
import os
from neo4j_agent_memory import MemoryClient, MemorySettings
settings = MemorySettings(
backend="bolt",
neo4j={
"uri": os.environ["NEO4J_URI"],
"username": os.environ["NEO4J_USERNAME"],
"password": os.environ["NEO4J_PASSWORD"],
"database": os.getenv("NEO4J_DATABASE", "neo4j"),
},
llm="anthropic/claude-3-5-sonnet-latest",
embedding="openai/text-embedding-3-small",
)
async with MemoryClient(settings) as client:
...
-
Install:
pip install 'neo4j-agent-memory[anthropic,openai]==0.7.0' -
The factory picks a native adapter when the matching extra is installed, and falls back to LiteLLM for unsupported providers.
-
Returned Providers implement the
LLMProvider/EmbeddingProviderProtocols.
Pick a provider
| Provider | Example model string | Extra | Notes |
|---|---|---|---|
OpenAI |
|
|
Native: strict-mode structured output. |
Anthropic |
|
|
Native: forced tool-use + optional prompt caching. |
AWS Bedrock |
|
|
Native: Converse API; reads boto3 credential chain. |
Vertex AI (Gemini) |
|
|
Routes via LiteLLM; needs ADC credentials. |
Ollama (local) |
|
|
Pass |
Groq |
|
|
LiteLLM universal. |
Together |
|
|
LiteLLM universal. |
Cohere |
|
|
LiteLLM universal. |
OpenRouter (any) |
|
|
LiteLLM universal. |
Embedding-only providers:
| Provider | Example model string | Extra | Dimensions |
|---|---|---|---|
OpenAI |
|
|
1536 |
OpenAI (large) |
|
|
3072 |
Vertex AI |
|
|
768 (3072 native, truncated by default) [1] |
Bedrock Titan |
|
|
1024 |
sentence-transformers |
|
|
384 |
sentence-transformers |
|
|
1024 |
Cohere |
|
|
1024 |
Voyage |
|
|
1024 |
For models not in the defaults table, pass an explicit dimensions=N when constructing the adapter directly, or via --embedding-dimensions on the MCP CLI.
Native-first resolution
When you call from_provider("openai/gpt-4o-mini"):
-
Parse
openaias the provider prefix. -
If the matching SDK for the
openai,anthropic, orbedrockprefix is installed, use its native adapter. -
Otherwise, if
[litellm]is installed, route throughLiteLLMProvider. -
Otherwise, raise
ImportErrorwith an install hint for both the native extra and the universal fallback.
You can force LiteLLM even when a native adapter is available:
from neo4j_agent_memory.llm import from_provider
provider = from_provider(
"openai/gpt-4o",
prefer_litellm=True,
)
Why this design? Native adapters get provider-specific features (OpenAI strict-mode JSON, Anthropic prompt caching, Bedrock Converse) that LiteLLM normalizes away or lags on. The escape hatch exists for consistency-across-providers testing.
Three ways to wire a provider
A. Provider-string shorthand
Simplest. The factory resolves the string.
import os
settings = MemorySettings(
backend="bolt",
neo4j={
"uri": os.environ["NEO4J_URI"],
"username": os.environ["NEO4J_USERNAME"],
"password": os.environ["NEO4J_PASSWORD"],
"database": os.getenv("NEO4J_DATABASE", "neo4j"),
},
llm="anthropic/claude-3-5-sonnet-latest",
)
B. Explicit provider instance
When you need adapter-specific kwargs (api_base, cache_system, aws_region):
import os
from neo4j_agent_memory.llm.adapters.anthropic import AnthropicProvider
from neo4j_agent_memory.llm.adapters.litellm import LiteLLMProvider
settings = MemorySettings(
backend="bolt",
neo4j={
"uri": os.environ["NEO4J_URI"],
"username": os.environ["NEO4J_USERNAME"],
"password": os.environ["NEO4J_PASSWORD"],
"database": os.getenv("NEO4J_DATABASE", "neo4j"),
},
llm=AnthropicProvider(
"anthropic/claude-3-5-sonnet-latest",
cache_system=True, # opt-in prompt caching
),
)
# Or a local model behind LiteLLM:
ollama = LiteLLMProvider(
"ollama/llama3.2",
api_base="http://localhost:11434",
)
C. Framework pass-through
Hand off a model you’ve already configured with your agent framework:
import os
from langchain_anthropic import ChatAnthropic
from neo4j_agent_memory.integrations.langchain import (
llm_provider_from_langchain,
)
chat = ChatAnthropic(model_name="claude-3-5-sonnet-latest")
settings = MemorySettings(
backend="bolt",
neo4j={
"uri": os.environ["NEO4J_URI"],
"username": os.environ["NEO4J_USERNAME"],
"password": os.environ["NEO4J_PASSWORD"],
"database": os.getenv("NEO4J_DATABASE", "neo4j"),
},
llm=llm_provider_from_langchain(chat),
)
See the migration guide for the full list of llm_provider_from_<framework> helpers.
Embedding models — the dimension gotcha
Embedding adapters require dimensions: int so Neo4j vector indexes are sized correctly at connect(). The defaults table covers common models; for an unknown model, pass dimensions= explicitly:
from neo4j_agent_memory.llm.adapters.sentence_transformers import (
SentenceTransformersProvider,
)
# Known model — dimensions auto-populated from defaults.
embedder = SentenceTransformersProvider("BAAI/bge-small-en-v1.5")
assert embedder.dimensions == 384
# Unknown model — must specify dimensions.
custom = SentenceTransformersProvider("my-org/my-internal-model", dimensions=512)
If you change embedding model after creating data, see Migrate to a new embedding model for the index-rebuild runbook.
Structured extraction
The library’s entity extractor calls complete_structured() when the provider implements StructuredExtractor. Supported adapters use these structured-output strategies; validate both schema conformance and extraction quality with your own data:
-
OpenAI: strict mode (
response_format={"type": "json_schema", "strict": True}) — uses the strict schema mode; refusals, request errors, and validation failures still require handling. -
Anthropic: forced tool use — the model is required to call a single tool whose input is your Pydantic schema.
-
LiteLLM: schema-aligned retry (
schema_aligned_extract) — feeds validation errors back to the LLM as feedback for up to 2 retries.
You can use the same pattern directly:
from pydantic import BaseModel
from neo4j_agent_memory.llm import ChatMessage, from_provider
class City(BaseModel):
name: str
population: int
provider = from_provider("anthropic/claude-3-5-sonnet-latest")
city = await provider.complete_structured(
[ChatMessage(role="user", content="Population of Paris in 2024?")],
response_model=City,
)
print(city.name, city.population)
If a provider does not implement StructuredExtractor, the universal schema_aligned_extract helper still works:
from neo4j_agent_memory.llm import schema_aligned_extract
city = await schema_aligned_extract(
provider,
messages=[ChatMessage(role="user", content="...")],
response_model=City,
max_retries=2,
)
Error handling
The adapters translate recognized SDK-specific exceptions to the provider hierarchy in neo4j_agent_memory.llm.errors:
import asyncio
from neo4j_agent_memory.llm import ChatMessage, ProviderRateLimitError, ProviderTimeoutError
messages = [ChatMessage(role="user", content="Reply with the word ready.")]
try:
result = await provider.complete(messages)
except ProviderRateLimitError as e:
# Same except clause works across OpenAI, Anthropic, Bedrock, LiteLLM.
await asyncio.sleep(e.retry_after or 1.0)
result = await provider.complete(messages)
except ProviderTimeoutError:
raise
The full hierarchy:
-
ProviderError(base)-
ProviderAuthError— invalid/missing API key. -
ProviderRateLimitError— carriesretry_after: float | None. -
ProviderTimeoutError. -
ProviderInvalidRequestError— unknown model, malformed request. -
ProviderServiceError— 5xx / retriable. -
StructuredExtractionError— SAP retries exhausted; carrieslast_attemptsandvalidation_errors. -
EmbeddingDimensionMismatchError— see migration runbook.
-
Provider matrix at a glance
| Adapter | LLM Bronze | Structured Silver | Embedding | Notes |
|---|---|---|---|---|
|
✓ |
✓ (strict mode) |
Native SDK adapter. |
|
|
✓ |
Dimension reduction supported. |
||
|
✓ |
✓ (forced tool) |
Optional prompt caching. |
|
|
✓ |
✓ (tool use) |
Boto3 credential chain. |
|
|
✓ |
Titan + Cohere via Bedrock. |
||
|
✓ |
✓ (via SAP) |
LiteLLM routing. |
|
|
✓ |
Cohere, Voyage, etc. |
||
|
✓ |
Local, no API key. |
||
|
✓ |
Wraps existing Vertex AI embedder. |
||
|
✓ (Instructor SDK) |
For users already on Instructor. |
Configure via the MCP CLI
The neo4j-agent-memory command needs the cli extra, and mcp serve also needs the mcp extra. Install them with the extras for the providers the command selects:
pip install 'neo4j-agent-memory[cli,mcp,anthropic,sentence-transformers]==0.7.0'
Match the Python API surface from the command line:
neo4j-agent-memory mcp serve \
--backend bolt --uri "$NEO4J_URI" --user "$NEO4J_USERNAME" \
--llm anthropic/claude-3-5-sonnet-latest \
--embedding BAAI/bge-small-en-v1.5 \
--llm-api-key $ANTHROPIC_API_KEY
Or via env vars:
export NAM_LLM=anthropic/claude-3-5-sonnet-latest
export NAM_EMBEDDING=BAAI/bge-small-en-v1.5
neo4j-agent-memory mcp serve --backend bolt --uri "$NEO4J_URI" --user "$NEO4J_USERNAME"
See CLI reference for the full flag set.
Verify the configuration
Print the configured provider class and make one small completion request.
For embeddings, compare len(await embedder.embed_one("test")) with
embedder.dimensions. Run a labeled extraction example before changing
production settings. Model changes require retrieval validation even if the
vector dimension stays the same; see the migration runbook.
See also
-
Tutorial: Anthropic + local embeddings — a copy-paste-runnable walkthrough.
-
Why the provider protocol? — design rationale.
-
Migrate to pluggable providers — backward compat and side-by-side examples.
gemini-embedding-001 is not a key in neo4j_agent_memory.llm.defaults.EMBEDDING_DIMENSIONS. VertexAIEmbeddingProvider falls back to 768 and logs a warning when it does not recognize the model — the 768 is adapter fallback behavior, not a tabulated default. Passing dimensions= silences the warning but only declares the index size; this adapter does not forward it as Vertex AI’s output_dimensionality, so it still emits 768-dimension vectors. See Use Vertex AI embeddings.