Configure an LLM provider

Available on NAMS: No. The procedure on this page needs the Bolt backend. With the hosted NAMS backend its operations are unavailable: depending on the call, the client raises NotSupportedError, AttributeError or TypeError, or ignores a Bolt-only setting or argument (NAMS manages embedding and extraction server-side). See the backend capabilities reference for what NAMS provides instead.

Configure an LLM provider through MemorySettings.llm and verify which adapter handles requests.

For the bigger-picture choice between providers, see Bring your own model. For the conceptual rationale, see Why the provider protocol?.

Setting only MemorySettings.llm does not configure embeddings. MemorySettings.embedding defaults to EmbeddingConfig(), whose provider defaults to OpenAI’s text-embedding-3-small — so an LLM-only configuration still needs OPENAI_API_KEY set for extraction and search to embed text. Set embedding= explicitly (see Configure an embedding provider) to use a different embedding provider.

Prerequisites

  • Install the extras for each selected provider and configure its credentials. Quote package extras in shell commands, for example pip install 'neo4j-agent-memory[openai]==0.7.0'.

  • Use a dedicated AuraDB instance and export NEO4J_URI, NEO4J_USERNAME, and NEO4J_PASSWORD as shown in the Aura connection setup. These examples use the bolt backend; NEO4J_DATABASE defaults to neo4j when omitted.

  • Model identifiers below illustrate adapter routing. Confirm the chosen model is available to your account before making a live request; successful provider construction does not verify model access.

  • Run snippets containing await in an async function. The provider, messages, and connected client in later fragments come from your application’s setup.

Procedure

  1. Choose a string, explicit adapter, or framework pass-through from the options below.

  2. Configure both the LLM and embedding provider; changing one does not configure the other.

  3. Run the verification at the end before applying the configuration to an existing graph.

Set the LLM

MemorySettings.llm accepts the same three shapes as every other provider setting — a provider string, an explicit Provider instance, or a framework pass-through. Bring your own model — Three ways to wire a provider is the canonical walkthrough with runnable code for all three; the deltas specific to the LLM field are below.

  • Provider string (recommended for most users) — llm="anthropic/claude-3-5-sonnet-latest" resolves via from_provider(model, kind="llm") with native-first dispatch.

  • Explicit Provider instance — construct the adapter directly for kwargs the string form can’t express, for example AnthropicProvider(…​, cache_system=True, timeout=120.0, max_retries=5) for Anthropic prompt caching, request timeout, and retry count.

  • Framework pass-through — wrap an already-configured ChatOpenAI, ChatAnthropic, etc. with the matching llm_provider_from_<framework> helper. Pass-through helpers exist for every supported integration; see Migrate to pluggable providers — Pattern 5 for the full list.

Use no LLM at all

import os
from neo4j_agent_memory import MemorySettings
from neo4j_agent_memory.config.settings import ExtractionConfig, ExtractorType

settings = MemorySettings(
    backend="bolt",
    neo4j={
        "uri": os.environ["NEO4J_URI"],
        "username": os.environ["NEO4J_USERNAME"],
        "password": os.environ["NEO4J_PASSWORD"],
        "database": os.getenv("NEO4J_DATABASE", "neo4j"),
    },
    llm=None,
    extraction=ExtractionConfig(
        extractor_type=ExtractorType.GLINER,  # GLiNER2.5; or SPACY / PIPELINE
        enable_llm_fallback=False,
    ),
)

See Run without an LLM for the full guide.

Override the API key

from neo4j_agent_memory.llm import from_provider

provider = from_provider(
    "anthropic/claude-3-5-sonnet-latest",
    api_key="sk-ant-...",
)

Otherwise the adapter reads the provider’s standard env var (OPENAI_API_KEY, ANTHROPIC_API_KEY, AWS credential chain, …​).

Override the API base (vLLM / Ollama / internal endpoint)

from neo4j_agent_memory.llm.adapters.litellm import LiteLLMProvider

# Local Ollama
ollama = LiteLLMProvider(
    "ollama/llama3.2",
    api_base="http://localhost:11434",
)

# Self-hosted vLLM with an OpenAI-compatible endpoint
internal = LiteLLMProvider(
    "openai/llama-3.3-70b-instruct",     # any LiteLLM-routed openai/* works
    api_base="https://llms.internal.corp/v1",
    api_key="...",
)

Force LiteLLM (skip the native adapter)

from neo4j_agent_memory.llm import from_provider

provider = from_provider(
    "openai/gpt-4o",
    prefer_litellm=True,
)

Useful for consistency tests, observability standardisation across providers, or when you want LiteLLM’s cost tracking.

Configure via the MCP CLI

The neo4j-agent-memory command needs the cli extra, and mcp serve also needs the mcp extra. This example selects Anthropic for the LLM and keeps the default OpenAI embedding provider:

pip install 'neo4j-agent-memory[cli,mcp,anthropic,openai]==0.7.0'
neo4j-agent-memory mcp serve \
  --backend bolt --uri "$NEO4J_URI" --user "$NEO4J_USERNAME" \
  --llm anthropic/claude-3-5-sonnet-latest \
  --llm-api-key $ANTHROPIC_API_KEY \
  --llm-api-base https://custom-endpoint.example.com  # optional

Or env vars:

export NAM_LLM=anthropic/claude-3-5-sonnet-latest
export NAM_LLM_API_KEY=$ANTHROPIC_API_KEY
neo4j-agent-memory mcp serve --backend bolt --uri "$NEO4J_URI" --user "$NEO4J_USERNAME"

Catch provider errors uniformly

from neo4j_agent_memory.llm import (
    ProviderAuthError, ProviderRateLimitError, ProviderTimeoutError,
)

try:
    result = await provider.complete(messages)
except ProviderRateLimitError as exc:
    # This branch reports the failure; waiting alone does not retry the call.
    print("Rate limited; suggested delay:", exc.retry_after)
    raise
except ProviderAuthError:
    raise SystemExit("API key invalid — check your env vars.")
except ProviderTimeoutError:
    raise  # Let the caller decide whether another attempt is appropriate.

Verify which adapter you got

Provider strings dispatch to different adapter classes based on what is installed. This checks routing without making a model request:

import logging
logging.getLogger("neo4j_agent_memory.llm.factory").setLevel(logging.DEBUG)

from neo4j_agent_memory.llm import from_provider
provider = from_provider("openai/gpt-4o-mini")
print(type(provider).__name__)
# OpenAIProvider  (when [openai] is installed)
# LiteLLMProvider (when only [litellm] is installed)

Then perform a small request using the credentials and model selected for your environment, and verify the returned content. Construction alone does not verify authentication, quota, model availability, or extraction quality. If you add retries, bound their number and account for adapter-level retries.