Use with LlamaIndex

Persist and restore LlamaIndex ChatMessage objects with Neo4jLlamaIndexMemory. This focused recipe verifies the memory adapter itself without requiring an LLM response, document loader, product catalog or invented custom chat store.

1. Prepare the selected environment

Use Python 3.10+ and a POSIX shell. The recipe uses BoltSettings to select Neo4j explicitly and disables entity extraction so that message persistence can be checked independently. Use a dedicated AuraDB instance with vector-index support and no incompatible existing vectors. Follow the Aura connection setup and copy its connection values into the exports below. These recipes use the published Python 0.7.0 package; provider access and database execution must be verified in your environment.

Create a local folder and virtual environment for the examples on this page. The commands reuse an existing environment without changing its files:

mkdir -p ~/agent-memory-tutorials
cd ~/agent-memory-tutorials
if [ -e .venv ]; then
  printf '%s\n' 'Using the existing virtual environment.'
else
  python3 -m venv .venv
fi
source .venv/bin/activate

Expected: ~/agent-memory-tutorials is your working directory and its virtual environment is active. Install the published SDK with the command below.

Each complete code block labelled Save as names a file to create in this folder using your editor. Copy the entire block, including imports and the entry point. Expand each helper disclosure and use Copy code to copy its full source. Keep all files together so their imports resolve.

When continuing from another tutorial or guide, retain the existing environment, configuration, session files, and .tutorial-state/. Reuse unchanged helper files; compare an existing file before replacing it, and finish any pending cleanup or recovery before changing the code that owns its state.

python -m pip install 'neo4j-agent-memory[llamaindex,openai]==0.7.0'
export NEO4J_URI='neo4j+s://<instance-id>.databases.neo4j.io'
export NEO4J_USERNAME='neo4j'
export NEO4J_PASSWORD='replace-with-your-Aura-password'
export NEO4J_DATABASE='neo4j'
export OPENAI_API_KEY='replace-with-your-provider-key'

The OpenAI-backed recipes require access to text-embedding-3-small; it produces 1536-dimensional vectors in this configuration. Configure the selected chat model separately where required. Credentials are read from the environment, never printed by the programs.

2. Read the complete recipe

Create the subdirectory below, then save both complete files under the displayed names. Run commands from agent-memory-tutorials/; Python resolves common from the script’s integrations/ directory.

mkdir -p integrations

The script imports this shared helper from the same directory. It supplies explicit database settings, closes clients through the calling context manager, and fails when expected records are absent. Task traces record the observable outcome of the call; they do not expose hidden model reasoning.

Complete integrations/common.py
Save as integrations/common.py
"""Shared Aura connection through Bolt and readback for integration recipes."""

import os


def settings(embedding=None):
    from neo4j_agent_memory import BoltSettings
    from neo4j_agent_memory.llm import from_provider

    return BoltSettings(
        neo4j={
            "uri": os.environ["NEO4J_URI"],
            "username": os.environ["NEO4J_USERNAME"],
            "password": os.environ["NEO4J_PASSWORD"],
            "database": os.getenv("NEO4J_DATABASE", "neo4j"),
        },
        embedding=embedding or from_provider("openai/text-embedding-3-small", kind="embedding"),
        extraction={"extractor_type": "none"},
    )


async def verify_messages(client, session_id, expected):
    conversation = await client.short_term.get_conversation(session_id)
    contents = [message.content for message in conversation.messages]
    if not all(text in contents for text in expected):
        raise RuntimeError(f"Message readback failed for {session_id}")
    print(f"Verified stored messages; session={session_id}")


async def recorded_turn(client, session_id, prompt, respond):
    """Record a task outcome; this does not claim to capture hidden model reasoning."""
    trace = await client.reasoning.start_trace(session_id=session_id, task=prompt)
    try:
        await client.short_term.add_message(session_id, "user", prompt, extract_entities=False)
        context = await client.get_context(prompt, session_id=session_id)
        reply = str(await respond(context))
        if not reply.strip():
            raise RuntimeError("The framework returned no text")
        await client.short_term.add_message(session_id, "assistant", reply, extract_entities=False)
        await verify_messages(client, session_id, [prompt, reply])
    except Exception as exc:
        await client.reasoning.complete_trace(
            trace.id, success=False, outcome=f"Run failed: {type(exc).__name__}"
        )
        raise
    await client.reasoning.complete_trace(
        trace.id, success=True, outcome="Reply stored and read back"
    )
    stored = await client.reasoning.get_trace_with_steps(trace.id)
    if stored is None or stored.success is not True:
        raise RuntimeError(f"Trace readback failed for {trace.id}")
    print(f"Verified task trace: {trace.id}")
    return reply
Complete integrations/llamaindex_recipe.py
Save as integrations/llamaindex_recipe.py
"""Round-trip LlamaIndex ChatMessage objects through the shipped async memory."""

import asyncio
from uuid import uuid4

from common import settings


async def exercise(client):
    from llama_index.core.base.llms.types import ChatMessage, MessageRole

    from neo4j_agent_memory.integrations.llamaindex import Neo4jLlamaIndexMemory

    session_id = f"docs-llama-{uuid4().hex[:8]}"
    memory = Neo4jLlamaIndexMemory(memory_client=client, session_id=session_id)
    text = "The fictional project code is cedar-lantern-47."
    await memory.aput(ChatMessage(role=MessageRole.USER, content=text))
    await memory.aput(ChatMessage(role=MessageRole.ASSISTANT, content="Project code recorded."))
    # Rebuild the adapter; its history must come from the supplied storage client.
    restored = Neo4jLlamaIndexMemory(memory_client=client, session_id=session_id)
    messages = await restored.aget_all()
    if not any(message.content == text for message in messages):
        raise RuntimeError(f"LlamaIndex memory readback failed for {session_id}")
    print(f"Verified LlamaIndex ChatMessage readback; session={session_id}")
    return messages


async def main():
    from neo4j_agent_memory import MemoryClient

    async with MemoryClient(settings()) as client:
        await exercise(client)


if __name__ == "__main__":
    asyncio.run(main())

3. Run and verify persistence

python integrations/llamaindex_recipe.py

Expected: Verified LlamaIndex ChatMessage readback with a session ID. A newly constructed adapter reads the exact sentinel from storage. The test does not claim a fresh-process restart; use the same session ID in your application’s next process to restore that conversation.

Troubleshooting and cleanup

Stop when the program raises an exception; a printed model response alone does not verify persistence. Check the configured database, provider access and vector dimensions before retrying. The program prints the session or trace IDs needed for inspection and closes the client. Retained example records remain in the dedicated database; follow the Aura tutorial’s cleanup procedure for that dedicated instance only when you no longer need them. Reusing a session groups records; it does not authorize a user to read them.

Adapt the integration

Pass the adapter to the memory parameter supported by your installed LlamaIndex agent or chat engine. The shipped class implements BaseMemory; use aput, aget, aget_all, aset and areset in async applications. aget(input) combines recent session history with semantic recall; aget_all() is the conversation readback used above. In the current adapter it delegates to the same recent-history loader, capped at ten messages; it is not an unlimited archive export. aset() and areset() can delete stored session history, so use them only when replacement/removal is intended.

The adapter preserves tool metadata and filters incomplete tool-call pairs during contextual retrieval. Store actual assistant tool calls and their matching tool responses; do not fabricate a pair to satisfy a provider protocol.

The current query-augmented path performs unscoped message search on Bolt and is not a demonstrated NAMS recipe. Session history is not a universal per-user filter for entity or preference retrieval. Document retrieval, product filtering and financial calculations must be supplied by your application. Combine retrieved documents with memory only after retaining their source IDs and selecting a context budget; no invented get_financial_documents or catalog API is supplied here.

llm_provider_from_llamaindex(llm) configures a model for memory extraction separately from chat history. See Configure providers and Memory types.

See Backend capabilities and scoping before adding hosted or multi-user behavior. Neo4j Agent Memory is a Neo4j Labs project with community support.