Use Vertex AI embeddings with Neo4j

Check Vertex AI embedding output and store a message in Neo4j using that embedder. ADK orchestration, MCP serving and Cloud Run deployment are linked as separate tasks so that this recipe has one observable outcome.

1. Prepare the selected environment

Install the Google Cloud CLI (gcloud) before configuring Application Default Credentials below.

Use Python 3.10+ and a POSIX shell. Prepare a dedicated AuraDB instance; its existing vectors and indexes must match the selected model and dimensions. The program disables entity extraction and does not require a chat model. Follow the Aura connection setup for a dedicated test instance, then copy its URI, username and password into the exports below.

Create a local folder and virtual environment for the examples on this page. The commands reuse an existing environment without changing its files:

mkdir -p ~/agent-memory-tutorials
cd ~/agent-memory-tutorials
if [ -e .venv ]; then
  printf '%s\n' 'Using the existing virtual environment.'
else
  python3 -m venv .venv
fi
source .venv/bin/activate

Expected: ~/agent-memory-tutorials is your working directory and its virtual environment is active. Install the published SDK with the command below.

Each complete code block labelled Save as names a file to create in this folder using your editor. Copy the entire block, including imports and the entry point. Expand each helper disclosure and use Copy code to copy its full source. Keep all files together so their imports resolve.

When continuing from another tutorial or guide, retain the existing environment, configuration, session files, and .tutorial-state/. Reuse unchanged helper files; compare an existing file before replacing it, and finish any pending cleanup or recovery before changing the code that owns its state.

python -m pip install 'neo4j-agent-memory[vertex-ai]==0.7.0'
export GOOGLE_CLOUD_PROJECT='replace-with-your-project-id'
export GOOGLE_CLOUD_LOCATION='us-central1'
export NEO4J_URI='neo4j+s://<instance-id>.databases.neo4j.io'
export NEO4J_USERNAME='neo4j'
export NEO4J_PASSWORD='replace-with-your-Aura-password'
export NEO4J_DATABASE='neo4j'
gcloud auth application-default login

The project needs Vertex AI enabled, billing and a principal authorized to invoke the selected model in that location. This recipe uses Application Default Credentials. The recipe declares 768 dimensions for gemini-embedding-001, which matches the 768-dimension output the underlying VertexAIEmbedder requests by default; the model’s native 3072 dimensions are not used by this program.

2. Read the complete recipe

Create the subdirectory below, then save both complete files under the displayed names. Run commands from agent-memory-tutorials/; Python resolves common from the script’s integrations/ directory.

mkdir -p integrations

The script imports this shared helper from the same directory. It supplies explicit database settings, closes clients through the calling context manager, and fails when expected records are absent. Task traces record the observable outcome of the call; they do not expose hidden model reasoning.

Complete integrations/common.py
Save as integrations/common.py
"""Shared Aura connection through Bolt and readback for integration recipes."""

import os


def settings(embedding=None):
    from neo4j_agent_memory import BoltSettings
    from neo4j_agent_memory.llm import from_provider

    return BoltSettings(
        neo4j={
            "uri": os.environ["NEO4J_URI"],
            "username": os.environ["NEO4J_USERNAME"],
            "password": os.environ["NEO4J_PASSWORD"],
            "database": os.getenv("NEO4J_DATABASE", "neo4j"),
        },
        embedding=embedding or from_provider("openai/text-embedding-3-small", kind="embedding"),
        extraction={"extractor_type": "none"},
    )


async def verify_messages(client, session_id, expected):
    conversation = await client.short_term.get_conversation(session_id)
    contents = [message.content for message in conversation.messages]
    if not all(text in contents for text in expected):
        raise RuntimeError(f"Message readback failed for {session_id}")
    print(f"Verified stored messages; session={session_id}")


async def recorded_turn(client, session_id, prompt, respond):
    """Record a task outcome; this does not claim to capture hidden model reasoning."""
    trace = await client.reasoning.start_trace(session_id=session_id, task=prompt)
    try:
        await client.short_term.add_message(session_id, "user", prompt, extract_entities=False)
        context = await client.get_context(prompt, session_id=session_id)
        reply = str(await respond(context))
        if not reply.strip():
            raise RuntimeError("The framework returned no text")
        await client.short_term.add_message(session_id, "assistant", reply, extract_entities=False)
        await verify_messages(client, session_id, [prompt, reply])
    except Exception as exc:
        await client.reasoning.complete_trace(
            trace.id, success=False, outcome=f"Run failed: {type(exc).__name__}"
        )
        raise
    await client.reasoning.complete_trace(
        trace.id, success=True, outcome="Reply stored and read back"
    )
    stored = await client.reasoning.get_trace_with_steps(trace.id)
    if stored is None or stored.success is not True:
        raise RuntimeError(f"Trace readback failed for {trace.id}")
    print(f"Verified task trace: {trace.id}")
    return reply
Complete integrations/cloud_embeddings_recipe.py
Save as integrations/cloud_embeddings_recipe.py
"""Check the selected cloud embedder, then store and read back one message."""

import argparse
import asyncio
import os
from uuid import uuid4

from common import settings, verify_messages


def build_embedder(provider):
    if provider == "bedrock":
        from neo4j_agent_memory.llm.adapters.bedrock import BedrockEmbeddingProvider

        return BedrockEmbeddingProvider(
            "bedrock/amazon.titan-embed-text-v2:0", aws_region=os.environ["AWS_REGION"]
        )
    from neo4j_agent_memory.llm.adapters.vertex_ai import VertexAIEmbeddingProvider

    return VertexAIEmbeddingProvider(
        "vertex_ai/gemini-embedding-001",
        project_id=os.environ["GOOGLE_CLOUD_PROJECT"],
        location=os.environ["GOOGLE_CLOUD_LOCATION"],
        dimensions=768,
    )


async def verify_vectors(embedder):
    texts = ["A fictional workshop in Denver", "A fictional workshop in Boston"]
    vectors = await embedder.embed(texts)
    if len(vectors) != len(texts) or any(len(v) != embedder.dimensions for v in vectors):
        raise RuntimeError(
            "Embedding count or vector dimensions did not match the selected configuration"
        )
    print(f"Verified {len(vectors)} vectors of {embedder.dimensions} dimensions")


async def main(provider):
    from neo4j_agent_memory import MemoryClient

    embedder = build_embedder(provider)
    await verify_vectors(embedder)
    async with MemoryClient(settings(embedding=embedder)) as client:
        session_id = f"docs-{provider}-{uuid4().hex[:8]}"
        text = "The fictional workshop takes place in Denver."
        await client.short_term.add_message(session_id, "user", text, extract_entities=False)
        await verify_messages(client, session_id, [text])


if __name__ == "__main__":
    parser = argparse.ArgumentParser()
    parser.add_argument("provider", choices=["bedrock", "vertex"])
    asyncio.run(main(parser.parse_args().provider))

3. Run and verify persistence

python integrations/cloud_embeddings_recipe.py vertex

Expected: Verified 2 vectors of 768 dimensions, then stored-message verification. A configured dimension is checked against the actual returned vectors before the database write. A successful local settings construction does not prove model access.

Troubleshooting and cleanup

Stop when the program raises an exception; a printed model response alone does not verify persistence. Check the configured database, provider access and vector dimensions before retrying. The program prints the session or trace IDs needed for inspection and closes the client. Retained example records remain in the dedicated database; follow the Aura tutorial’s cleanup procedure for that dedicated instance only when you no longer need them. Reusing a session groups records; it does not authorize a user to read them.

Adapt the integration

VertexAIEmbeddingProvider accepts the model string, project_id, location, task_type, dimensions and batch_size, and delegates to VertexAIEmbedder. In 0.7.0 it does not forward dimensions as output_dimensionality, so keep it at 768. For 1536 or 3072 dimensions, construct VertexAIEmbedder(output_dimensionality=…​) directly and pass it as embedder= to MemoryClient. Use one consistent model/configuration for both stored document vectors and retrieval. Changing dimensions requires an intentional index/data migration; matching dimension counts alone does not make embeddings from different models compatible. Configure embeddings shows the adapter in MemorySettings.

VertexAIEmbedder accepts the model IDs below. Each model’s native dimensions apply only with output_dimensionality=None; the embedder requests 768 dimensions by default.

Table 1. Supported Vertex AI embedding models
Model Native dimensions Notes

gemini-embedding-001

3072

Default. Truncatable to 1536 or 768 with output_dimensionality; one text per request.

text-embedding-005

768

English and code; up to 250 texts per request.

text-multilingual-embedding-002

768

Multilingual; up to 250 texts per request.

Google shut down text-embedding-004 on 2026-01-14 and the textembedding-gecko* family on 2025-04-09. Passing one of those IDs to VertexAIEmbedder raises EmbeddingError, and passing one to EmbeddingConfig(provider="vertex_ai", model=…​) fails validation. Both errors name the supported replacements.

task_type defaults to RETRIEVAL_DOCUMENT, for indexing stored documents. The other values are RETRIEVAL_QUERY for search queries, SEMANTIC_SIMILARITY, CLASSIFICATION and CLUSTERING.

For ADK, use the complete Runner recipe. For memory tools, use the MCP server guide and the canonical tool reference. Use the deployment reference and the maintained Cloud Run assets for service deployment, environment configuration and verification; this page does not duplicate a Dockerfile or grant public access.

The Google Cloud scripts demonstrate cloud embedding and ADK APIs. The financial-advisor application adds its own specialist agents, domain tools, UI and infrastructure. Those application components are not installed automatically by the memory library, and any real deployment requires the example’s own setup and checks.

See Backend capabilities and scoping before adding hosted or multi-user behavior. Neo4j Agent Memory is a Neo4j Labs project with community support.