Build a Next.js chat app with a memory rail

This guide builds a Next.js 16 App Router chat app whose whole memory — the transcript, the entity graph and the reasoning trail — lives in the hosted Neo4j Agent Memory Service. It is the task-oriented companion to the runnable example at typescript/examples/nextjs-memory-chat; clone that if you want the finished app, and read this for the four decisions that make it work.

ai 7 is ESM-only and requires Node.js 22 or newer, and so does this client. Next 16 uses Turbopack by default — a custom webpack config fails the build.

When to use the SDK middleware instead

If a single conversation’s own three-tier context (recent messages, observations, reflections) is enough — no cross-session recall, no entity graph, no reasoning trail — start at Integrate with the Vercel AI SDK instead: it wires memory into one model call with no extra package, and you can come back here when you want the UI. This guide’s richer retrieval (entities, graph relationships, reasoning history, cross-session messages) needs @neo4j-labs/nams-ai-provider; if you are deciding between the SDK middleware and the provider’s other modes, see Choose a NAMS AI provider mode.

What you need

  • Node.js 22+

  • A MEMORY_API_KEY from memory.neo4jlabs.com

  • An OPENAI_API_KEY (model id via OPENAI_MODEL, default gpt-5-mini)

  • Optionally a MEMORY_WORKSPACE_ID to scope the data to one workspace

Install from npm:

npm install @neo4j-labs/[email protected] @neo4j-labs/[email protected] \
  ai@^7.0 @ai-sdk/openai@^4.0 @ai-sdk/react@^4.0 zod@^4.6.0 @neo4j-nvl/react@^1.2.1

The runnable example pins both @neo4j-labs/nams-ai-provider and @neo4j-labs/agent-memory to file: links into this repository, so it always runs against current source; an application installs the two npm releases above instead.

Decision 1: put the conversation id in the URL

A conversation id is the only handle on a thread. Keep it in the route, not in React state or localStorage, and the URL becomes shareable, reloadable and resumable — because every message behind it is in NAMS, not in the browser.

app/page.tsx mints one and redirects. force-dynamic keeps the page off the build-time prerender path, so next build succeeds on a machine with no API key:

export const dynamic = "force-dynamic";

export default async function Home() {
  let conversationId: string | undefined;
  try {
    const conversation = await memoryClient().shortTerm.createConversation({
      userId: userId(),
    });
    conversationId = conversation.id;
  } catch (error) {
    // redirect() signals by throwing, so it must stay outside this try.
    return <SetupNotice message={String(error)} />;
  }
  redirect(`/c/${conversationId}`);
}

In Next 16 params is a promise — the synchronous form was removed:

export default async function ConversationPage({
  params,
}: {
  params: Promise<{ conversationId: string }>;
}) {
  const { conversationId } = await params;
  return <Workspace conversationId={conversationId} />;
}

Decision 2: send one turn, let memory supply the rest

app/api/chat/route.ts supplies the dependencies (MemoryClient, the memory provider config, the model, the user id) and calls handleChat from lib/handlers.ts; the route file itself does nothing but pass those four dependencies to it. Inside it, createNamsProvider({ …​deps.memory, baseProvider: () ⇒ deps.model, scope: { userId: deps.userId, conversationId } }) returns a ProviderV4, and nams.languageModel("chat") wraps deps.model so relevant memory — conversation history, cross-session messages, matching entities, their graph relationships and past reasoning steps — is retrieved and prepended to the last user message before the call, and the turn is persisted after it. baseProvider ignores the model id it is handed and always returns deps.model, so nams.languageModel("chat") is called with an arbitrary id; there is no other memory-specific code in the route:

export async function handleChat(deps: ChatDeps, request: Request): Promise<Response> {
  let body: ChatRequestBody;
  try {
    body = (await request.json()) as ChatRequestBody;
  } catch {
    return badRequest("Request body must be JSON.");
  }

  const conversationId = body.conversationId;
  if (!conversationId) return badRequest("conversationId is required.");

  const uiMessages = body.messages ?? [];
  const latest = uiMessages.filter((m) => m.role === "user").slice(-1);
  if (latest.length === 0) return badRequest("No user message to answer.");

  // Memory supplies the history; the client only has to send the new turn.
  const messages = await convertToModelMessages(latest);
  const question = uiMessageText(latest[0]!);

  const nams = createNamsProvider({
    ...deps.memory,
    baseProvider: () => deps.model,
    scope: { userId: deps.userId, conversationId },
  });
  const model = nams.languageModel("chat");

  const result = streamText({
    model,
    instructions: deps.instructions ?? INSTRUCTIONS,
    messages,
    // `onEnd` is the AI SDK 7 name (`onFinish` is a deprecated alias). The
    // reasoning step is what the trace drawer reads back — one row per answered
    // turn, addressable by conversation.
    onEnd: async ({ text }) => {
      await recordTurn(deps.client, conversationId, question, text);
    },
  });

  return result.toUIMessageStreamResponse();
}

Only the newest user turn is forwarded — the history the model sees comes back out of the provider, so the browser is not the source of truth for the conversation. onEnd (the AI SDK 7 name; onFinish is a deprecated alias) hands the answered turn to the small recordTurn helper in the same file, which records it as a reasoning step; that is the only call this route makes to MemoryClient directly.

The createNamsProvider stream wrapper awaits the assistant-turn write as part of closing the response stream, so a fully drained client stream implies the write was attempted — not skipped, unlike a plain fire-and-forget write. It is still best-effort: a failed write is logged and swallowed, not surfaced to the caller, and nothing in the response tells the client whether it succeeded. For a route that must coordinate storage with its lifecycle and give the client a visible success signal, pass persistInteractions: false to createNamsProvider and write the turn yourself, following the application-owned write pattern. That option turns off both of the provider’s writes, the user message as well as the answer, so the route must store both: when the stream ends, call deps.client.shortTerm.addMessage(conversationId, "user", question) and then store the assembled answer, and keep those writes alive until they settle. This is the same order the provider uses; a user message stored before the model call can come back from the provider’s own session search and repeat the question in the memory block. Writing only the answer leaves every user turn out of the transcript and out of the provider’s later retrieval.

On the client, useChat with DefaultChatTransport sends the conversation id in the request body:

const { messages, sendMessage, status } = useChat({
  id: conversationId,
  transport: new DefaultChatTransport({
    api: "/api/chat",
    body: { conversationId },
  }),
});

To make the "reload and it is still there" claim true, hydrate the transcript once on mount from shortTerm.getContext rather than from client state.

Decision 3: expand the graph lazily, with loadedIds

longTerm.getEntityGraph() gives the rail its first canvas. longTerm.expandGraph(nodeId, loadedIds) grows it: pass every id already on screen and the service answers with the delta only, so a double-click on a well-connected node does not re-send the subgraph the user is looking at.

// POST /api/memory/graph
const expanded = await client.longTerm.expandGraph(nodeId, loadedIds);

Own the accumulated graph in the component that issues the request — that is what keeps loadedIds honest — and union the delta in:

function merge(current: GraphPayload, delta: GraphPayload): GraphPayload {
  const nodes = new Map(current.nodes.map((n) => [n.id, n]));
  for (const node of delta.nodes) if (!nodes.has(node.id)) nodes.set(node.id, node);
  // ... same for edges
  return { nodes: [...nodes.values()], edges: [...edges.values()] };
}

expandGraph returns the viz-oriented shape ({ id, labels, properties } for nodes), so flatten properties.name / properties.type for your renderer.

Neo4j’s visualization library touches window at import time, so load it with next/dynamic and ssr: false:

const InteractiveNvlWrapper = dynamic(
  () => import("@neo4j-nvl/react").then((m) => m.InteractiveNvlWrapper),
  { ssr: false },
);

NVL reports clicks, not double-clicks; treat a second onNodeClick on the same node within ~350 ms as "expand".

Decision 4: wait for extraction instead of guessing

NAMS returns from a write before the entities in that message are searchable — extraction runs in a background pipeline. Refetch the graph immediately after an answer and you render a canvas that is one turn behind.

longTerm.waitForExtraction turns that race into an await. The predicate form is the one to reach for in a UI: "resolve when an entity I have not already got appears".

const known = new Set(knownIds);
const settled = await client.longTerm.waitForExtraction({
  query,                       // the user's turn, as the probe
  limit: 25,
  timeoutMs: 20_000,
  intervalMs: 1_500,
  predicate: (entities) => entities.some((e) => !known.has(e.id)),
});

It returns false on timeout rather than throwing, so a quiet turn that produced no new entities degrades to "nothing new" instead of an error. Validation and transport errors still propagate. A new workspace entity can also come from another conversation: this predicate is a UI refresh hint, not extraction provenance for this turn. Order the rail’s refresh as context → wait for extraction → refetch graph, and surface the middle step as a badge so the reader can see the window open and close.

Prefer predicate or expectedNames over minResults. NAMS entity search is nearest-neighbour, so a minResults threshold is satisfied almost immediately on a non-empty workspace and proves nothing about this turn.

Reading the reasoning trail back

Because the route recorded a step per turn, reasoning.getTraceByConversation gives you an audit drawer for free:

const trace = await client.reasoning.getTraceByConversation(conversationId);
// trace.steps: [{ id, reasoning, actionTaken, result, createdAt }, …]

Deploying

npx vercel deploy

Set MEMORY_API_KEY, OPENAI_API_KEY and optionally MEMORY_WORKSPACE_ID as project environment variables. There is no database to provision. The route handlers use only fetch, Request and Response, so they are edge-deployable — pass apiKey to MemoryClient explicitly, because on edge runtimes process.env is only populated inside the request scope. See Deploy to edge runtimes. Current-source example dependency declarations (not a release or live-service verification): @neo4j-labs/nams-ai-provider 0.3.0, @neo4j-labs/agent-memory 0.5.0, next 16.3.4, react 19.3.0, ai 7.0.97, @ai-sdk/openai 4.0.65, @ai-sdk/react 4.0.100, @chakra-ui/react 3.37.0, @neo4j-nvl/react 1.2.2, Node.js 22 — 2026-09-23.

Testing it without an API key

Put each route’s behaviour in a (deps, Request) ⇒ Response function and keep the route file a thin wrapper. Then the handlers can be driven in Vitest with the real MemoryClient and createNamsProvider over a mocked network (msw) and the AI SDK’s own MockLanguageModelV4 from ai/test:

import { MockLanguageModelV4 } from "ai/test";

const client = new MemoryClient({ endpoint: ENDPOINT, apiKey: API_KEY });
const deps = { client, memory: { apiKey: API_KEY, endpoint: ENDPOINT }, model: mockModel(), userId };
const response = await handleChat(deps, request);
await response.text();                       // drain the stream
expect(state.rolesOf(conversationId)).toEqual(["user", "assistant"]);

Mocking the network rather than the SDK means a wrong URL, verb or body shape fails the suite. Answer 501 for any endpoint you have not implemented so that a new call fails loudly instead of passing silently. This is the last check before shipping: a green suite here means the route, the provider wiring and the memory persistence contract all hold without needing a live API key.