Run without an LLM

Available on NAMS: No. The procedure on this page needs the Bolt backend. With the hosted NAMS backend its operations are unavailable: depending on the call, the client raises NotSupportedError, AttributeError or TypeError, or ignores a Bolt-only setting or argument (NAMS manages embedding and extraction server-side). See the backend capabilities reference for what NAMS provides instead.

Use neo4j-agent-memory with local extraction and embeddings, without constructing an LLM provider or requiring an OPENAI_API_KEY. This example stores its memory in AuraDB, so a network connection to Aura is still required.

Overview

MemorySettings.llm accepts a legacy config, a provider string, a provider instance, or None. Set it to None to opt out of all LLM construction. Combine that with a local embedder (sentence-transformers) and a local extractor (spaCy and/or GLiNER2.5) to get a client that does not construct an LLM provider. A local embedding provider is a separate requirement; llm=None alone does not disable cloud embedding calls.

Prerequisites

Use Python 3.10+ and a POSIX shell. Prepare the example files and environment:

Create a local folder and virtual environment for the examples on this page. The commands reuse an existing environment without changing its files:

mkdir -p ~/agent-memory-tutorials
cd ~/agent-memory-tutorials
if [ -e .venv ]; then
  printf '%s\n' 'Using the existing virtual environment.'
else
  python3 -m venv .venv
fi
source .venv/bin/activate

Expected: ~/agent-memory-tutorials is your working directory and its virtual environment is active. Install the published SDK with the command below.

Each complete code block labelled Save as names a file to create in this folder using your editor. Copy the entire block, including imports and the entry point. Expand each helper disclosure and use Copy code to copy its full source. Keep all files together so their imports resolve.

When continuing from another tutorial or guide, retain the existing environment, configuration, session files, and .tutorial-state/. Reuse unchanged helper files; compare an existing file before replacing it, and finish any pending cleanup or recovery before changing the code that owns its state.

pip install 'neo4j-agent-memory[extraction,sentence-transformers]==0.7.0'
python -m spacy download en_core_web_sm

Use a dedicated AuraDB instance configured through the Aura connection setup. Export NEO4J_URI, NEO4J_USERNAME, and NEO4J_PASSWORD for the example, and download the local models into the deployment cache. Local model inference has compute costs and does not guarantee identical outputs across hardware or dependency versions.

1. Configure local providers

import os

from neo4j_agent_memory import MemoryClient, MemorySettings
from neo4j_agent_memory.config.settings import ExtractionConfig, ExtractorType

settings = MemorySettings(
    backend="bolt",                                    (1)
    neo4j={
        "uri": os.environ["NEO4J_URI"],
        "username": os.environ["NEO4J_USERNAME"],
        "password": os.environ["NEO4J_PASSWORD"],
        "database": os.getenv("NEO4J_DATABASE", "neo4j"),
    },
    llm=None,                                          (2)
    embedding="sentence-transformers/all-MiniLM-L6-v2", (3)
    extraction=ExtractionConfig(
        extractor_type=ExtractorType.PIPELINE,
        enable_spacy=True,
        enable_gliner=True,
        enable_llm_fallback=False,                     (4)
    ),
)

async with MemoryClient(settings) as memory:
    await memory.short_term.add_message("session-1", "user", "John works at Acme")
    print(await memory.get_context("Tell me about John"))
1 Pins the direct Neo4j backend used by Aura, so a MEMORY_API_KEY in the environment does not redirect this configuration to NAMS.
2 Explicit opt-out — no LLM client is ever constructed.
3 Provider-string shorthand; resolves to a local SentenceTransformersProvider. An EmbeddingProvider instance also works.
4 Required when llm=None — see 2. Check configuration validation.

This configuration uses the bolt backend with Aura. On the hosted memory service (NAMS) extraction and embeddings run server-side, so there is no client-side LLM or embedding provider to opt out of — see Connect a Python application to NAMS.

2. Check configuration validation

MemorySettings raises a ValidationError at construction time if you set llm=None together with extraction settings that require an LLM:

  • extraction.extractor_type == ExtractorType.LLM, or

  • extraction.extractor_type == ExtractorType.PIPELINE and extraction.enable_llm_fallback is True.

The error message names both fields and points at the minimal fix.

If you omit the llm field altogether (rather than passing None), the package keeps its historical behavior of auto-filling a default LLMConfig when an LLM stage is enabled — so existing code that relies on the default doesn’t break.

Default behavior summary

Configuration LLM constructed? Notes

llm=None + extractor_type=SPACY/GLINER/NONE, enable_llm_fallback=False

No

No LLM construction. Use a local embedding provider as well to avoid cloud embedding calls.

llm=None + LLM-dependent extractor

—

ValidationError at construction time.

llm omitted + default ExtractionConfig (LLM fallback on)

Yes (default LLMConfig)

Backwards-compatible default — same as before this change.

llm omitted + non-LLM extractor

No

Validator skips the auto-fill since no LLM stage is enabled.

3. Verify the run with local models

Create the subdirectory, then save the complete script below as no_llm/main.py. It exercises all three memory layers with no LLM and fails fast when the local extraction stack is incomplete.

mkdir -p no_llm
Complete no_llm/main.py
Save as no_llm/main.py
Unresolved include directive in modules/ROOT/pages/how-to/running-without-an-llm.adoc - include::example$no_llm/main.py[]

Run it from agent-memory-tutorials/:

# Keep the exported Aura connection values in this process environment.
python no_llm/main.py

Confirm the example stores and retrieves messages, entities, and reasoning traces in the dedicated AuraDB instance. You can repeat after disabling outbound model-service access with all required models cached; keep access to Aura available. Missing local model files are a setup failure, not evidence that a hosted model fallback is required.

See also