Bring your own model

Available on NAMS: No. The procedure on this page needs the Bolt backend. With the hosted NAMS backend its operations are unavailable: depending on the call, the client raises NotSupportedError, AttributeError or TypeError, or ignores a Bolt-only setting or argument (NAMS manages embedding and extraction server-side). See the backend capabilities reference for what NAMS provides instead.

Configure an LLM or embedding provider — native adapters for OpenAI, Anthropic, Bedrock, Vertex AI, and sentence-transformers; LiteLLM universal fallback for everything else.

neo4j-agent-memory is provider-pluggable. You pass either a provider-string shorthand ("anthropic/claude-3-5-sonnet-latest"), an explicit Provider instance, or hand off your already-configured framework model. The factory chooses an installed native adapter when supported, then falls back to LiteLLM when available.

Prerequisites

  • Install the extras for each selected provider and configure its credentials. Quote package extras in shell commands, for example pip install 'neo4j-agent-memory[openai]==0.7.0'.

  • Use a dedicated AuraDB instance and export NEO4J_URI, NEO4J_USERNAME, and NEO4J_PASSWORD as shown in the Aura connection setup. These examples use the bolt backend; NEO4J_DATABASE defaults to neo4j when omitted.

  • Model identifiers below illustrate adapter routing. Confirm the chosen model is available to your account before making a live request; successful provider construction does not verify model access.

  • Run snippets containing await in an async function. The provider, messages, and connected client in later fragments come from your application’s setup.

Procedure

  1. Choose a string, explicit adapter, or framework pass-through from the options below.

  2. Configure both the LLM and embedding provider; changing one does not configure the other.

  3. Run the verification at the end before applying the configuration to an existing graph.

Configure a provider string

import os
from neo4j_agent_memory import MemoryClient, MemorySettings

settings = MemorySettings(
    backend="bolt",
    neo4j={
        "uri": os.environ["NEO4J_URI"],
        "username": os.environ["NEO4J_USERNAME"],
        "password": os.environ["NEO4J_PASSWORD"],
        "database": os.getenv("NEO4J_DATABASE", "neo4j"),
    },
    llm="anthropic/claude-3-5-sonnet-latest",
    embedding="openai/text-embedding-3-small",
)

async with MemoryClient(settings) as client:
    ...
  • Install: pip install 'neo4j-agent-memory[anthropic,openai]==0.7.0'

  • The factory picks a native adapter when the matching extra is installed, and falls back to LiteLLM for unsupported providers.

  • Returned Providers implement the LLMProvider / EmbeddingProvider Protocols.

Pick a provider

Provider Example model string Extra Notes

OpenAI

openai/gpt-4o-mini

[openai]

Native: strict-mode structured output.

Anthropic

anthropic/claude-3-5-sonnet-latest

[anthropic]

Native: forced tool-use + optional prompt caching.

AWS Bedrock

bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0

[bedrock]

Native: Converse API; reads boto3 credential chain.

Vertex AI (Gemini)

vertex_ai/gemini-2.5-flash

[litellm]

Routes via LiteLLM; needs ADC credentials.

Ollama (local)

ollama/llama3.2

[litellm]

Pass api_base="http://localhost:11434".

Groq

groq/llama-3.1-70b-versatile

[litellm]

LiteLLM universal.

Together

together_ai/meta-llama/Llama-3.3-70B-Instruct-Turbo

[litellm]

LiteLLM universal.

Cohere

cohere/command-r-plus

[litellm]

LiteLLM universal.

OpenRouter (any)

openrouter/anthropic/claude-3.5-sonnet

[litellm]

LiteLLM universal.

Embedding-only providers:

Provider Example model string Extra Dimensions

OpenAI

openai/text-embedding-3-small

[openai]

1536

OpenAI (large)

openai/text-embedding-3-large

[openai]

3072

Vertex AI

vertex_ai/gemini-embedding-001

[vertex-ai]

768 (3072 native, truncated by default) [1]

Bedrock Titan

bedrock/amazon.titan-embed-text-v2:0

[bedrock]

1024

sentence-transformers

BAAI/bge-small-en-v1.5

[sentence-transformers]

384

sentence-transformers

BAAI/bge-large-en-v1.5

[sentence-transformers]

1024

Cohere

cohere/embed-english-v3.0

[litellm]

1024

Voyage

voyage/voyage-3

[litellm]

1024

For models not in the defaults table, pass an explicit dimensions=N when constructing the adapter directly, or via --embedding-dimensions on the MCP CLI.

Native-first resolution

When you call from_provider("openai/gpt-4o-mini"):

  1. Parse openai as the provider prefix.

  2. If the matching SDK for the openai, anthropic, or bedrock prefix is installed, use its native adapter.

  3. Otherwise, if [litellm] is installed, route through LiteLLMProvider.

  4. Otherwise, raise ImportError with an install hint for both the native extra and the universal fallback.

You can force LiteLLM even when a native adapter is available:

from neo4j_agent_memory.llm import from_provider

provider = from_provider(
    "openai/gpt-4o",
    prefer_litellm=True,
)

Why this design? Native adapters get provider-specific features (OpenAI strict-mode JSON, Anthropic prompt caching, Bedrock Converse) that LiteLLM normalizes away or lags on. The escape hatch exists for consistency-across-providers testing.

Three ways to wire a provider

A. Provider-string shorthand

Simplest. The factory resolves the string.

import os
settings = MemorySettings(
    backend="bolt",
    neo4j={
        "uri": os.environ["NEO4J_URI"],
        "username": os.environ["NEO4J_USERNAME"],
        "password": os.environ["NEO4J_PASSWORD"],
        "database": os.getenv("NEO4J_DATABASE", "neo4j"),
    },
    llm="anthropic/claude-3-5-sonnet-latest",
)

B. Explicit provider instance

When you need adapter-specific kwargs (api_base, cache_system, aws_region):

import os
from neo4j_agent_memory.llm.adapters.anthropic import AnthropicProvider
from neo4j_agent_memory.llm.adapters.litellm import LiteLLMProvider

settings = MemorySettings(
    backend="bolt",
    neo4j={
        "uri": os.environ["NEO4J_URI"],
        "username": os.environ["NEO4J_USERNAME"],
        "password": os.environ["NEO4J_PASSWORD"],
        "database": os.getenv("NEO4J_DATABASE", "neo4j"),
    },
    llm=AnthropicProvider(
        "anthropic/claude-3-5-sonnet-latest",
        cache_system=True,           # opt-in prompt caching
    ),
)

# Or a local model behind LiteLLM:
ollama = LiteLLMProvider(
    "ollama/llama3.2",
    api_base="http://localhost:11434",
)

C. Framework pass-through

Hand off a model you’ve already configured with your agent framework:

import os
from langchain_anthropic import ChatAnthropic
from neo4j_agent_memory.integrations.langchain import (
    llm_provider_from_langchain,
)

chat = ChatAnthropic(model_name="claude-3-5-sonnet-latest")

settings = MemorySettings(
    backend="bolt",
    neo4j={
        "uri": os.environ["NEO4J_URI"],
        "username": os.environ["NEO4J_USERNAME"],
        "password": os.environ["NEO4J_PASSWORD"],
        "database": os.getenv("NEO4J_DATABASE", "neo4j"),
    },
    llm=llm_provider_from_langchain(chat),
)

See the migration guide for the full list of llm_provider_from_<framework> helpers.

Embedding models — the dimension gotcha

Embedding adapters require dimensions: int so Neo4j vector indexes are sized correctly at connect(). The defaults table covers common models; for an unknown model, pass dimensions= explicitly:

from neo4j_agent_memory.llm.adapters.sentence_transformers import (
    SentenceTransformersProvider,
)

# Known model — dimensions auto-populated from defaults.
embedder = SentenceTransformersProvider("BAAI/bge-small-en-v1.5")
assert embedder.dimensions == 384

# Unknown model — must specify dimensions.
custom = SentenceTransformersProvider("my-org/my-internal-model", dimensions=512)

If you change embedding model after creating data, see Migrate to a new embedding model for the index-rebuild runbook.

Structured extraction

The library’s entity extractor calls complete_structured() when the provider implements StructuredExtractor. Supported adapters use these structured-output strategies; validate both schema conformance and extraction quality with your own data:

  • OpenAI: strict mode (response_format={"type": "json_schema", "strict": True}) — uses the strict schema mode; refusals, request errors, and validation failures still require handling.

  • Anthropic: forced tool use — the model is required to call a single tool whose input is your Pydantic schema.

  • LiteLLM: schema-aligned retry (schema_aligned_extract) — feeds validation errors back to the LLM as feedback for up to 2 retries.

You can use the same pattern directly:

from pydantic import BaseModel
from neo4j_agent_memory.llm import ChatMessage, from_provider

class City(BaseModel):
    name: str
    population: int

provider = from_provider("anthropic/claude-3-5-sonnet-latest")
city = await provider.complete_structured(
    [ChatMessage(role="user", content="Population of Paris in 2024?")],
    response_model=City,
)
print(city.name, city.population)

If a provider does not implement StructuredExtractor, the universal schema_aligned_extract helper still works:

from neo4j_agent_memory.llm import schema_aligned_extract

city = await schema_aligned_extract(
    provider,
    messages=[ChatMessage(role="user", content="...")],
    response_model=City,
    max_retries=2,
)

Error handling

The adapters translate recognized SDK-specific exceptions to the provider hierarchy in neo4j_agent_memory.llm.errors:

import asyncio

from neo4j_agent_memory.llm import ChatMessage, ProviderRateLimitError, ProviderTimeoutError

messages = [ChatMessage(role="user", content="Reply with the word ready.")]
try:
    result = await provider.complete(messages)
except ProviderRateLimitError as e:
    # Same except clause works across OpenAI, Anthropic, Bedrock, LiteLLM.
    await asyncio.sleep(e.retry_after or 1.0)
    result = await provider.complete(messages)
except ProviderTimeoutError:
    raise

The full hierarchy:

  • ProviderError (base)

    • ProviderAuthError — invalid/missing API key.

    • ProviderRateLimitError — carries retry_after: float | None.

    • ProviderTimeoutError.

    • ProviderInvalidRequestError — unknown model, malformed request.

    • ProviderServiceError — 5xx / retriable.

    • StructuredExtractionError — SAP retries exhausted; carries last_attempts and validation_errors.

    • EmbeddingDimensionMismatchError — see migration runbook.

Provider matrix at a glance

Adapter LLM Bronze Structured Silver Embedding Notes

OpenAIProvider

✓

✓ (strict mode)

Native SDK adapter.

OpenAIEmbeddingProvider

✓

Dimension reduction supported.

AnthropicProvider

✓

✓ (forced tool)

Optional prompt caching.

BedrockProvider

✓

✓ (tool use)

Boto3 credential chain.

BedrockEmbeddingProvider

✓

Titan + Cohere via Bedrock.

LiteLLMProvider

✓

✓ (via SAP)

LiteLLM routing.

LiteLLMEmbeddingProvider

✓

Cohere, Voyage, etc.

SentenceTransformersProvider

✓

Local, no API key.

VertexAIEmbeddingProvider

✓

Wraps existing Vertex AI embedder.

InstructorProvider

✓ (Instructor SDK)

For users already on Instructor.

Configure via the MCP CLI

The neo4j-agent-memory command needs the cli extra, and mcp serve also needs the mcp extra. Install them with the extras for the providers the command selects:

pip install 'neo4j-agent-memory[cli,mcp,anthropic,sentence-transformers]==0.7.0'

Match the Python API surface from the command line:

neo4j-agent-memory mcp serve \
  --backend bolt --uri "$NEO4J_URI" --user "$NEO4J_USERNAME" \
  --llm anthropic/claude-3-5-sonnet-latest \
  --embedding BAAI/bge-small-en-v1.5 \
  --llm-api-key $ANTHROPIC_API_KEY

Or via env vars:

export NAM_LLM=anthropic/claude-3-5-sonnet-latest
export NAM_EMBEDDING=BAAI/bge-small-en-v1.5
neo4j-agent-memory mcp serve --backend bolt --uri "$NEO4J_URI" --user "$NEO4J_USERNAME"

See CLI reference for the full flag set.

Verify the configuration

Print the configured provider class and make one small completion request. For embeddings, compare len(await embedder.embed_one("test")) with embedder.dimensions. Run a labeled extraction example before changing production settings. Model changes require retrieval validation even if the vector dimension stays the same; see the migration runbook.

See also


1. gemini-embedding-001 is not a key in neo4j_agent_memory.llm.defaults.EMBEDDING_DIMENSIONS. VertexAIEmbeddingProvider falls back to 768 and logs a warning when it does not recognize the model — the 768 is adapter fallback behavior, not a tabulated default. Passing dimensions= silences the warning but only declares the index size; this adapter does not forward it as Vertex AI’s output_dimensionality, so it still emits 768-dimension vectors. See Use Vertex AI embeddings.