Build a Next.js chat app with a memory rail
This guide builds a Next.js 16 App Router chat app whose whole memory — the
transcript, the entity graph and the reasoning trail — lives in the hosted
Neo4j Agent Memory Service. It is the task-oriented companion to the runnable
example at
typescript/examples/nextjs-memory-chat;
clone that if you want the finished app, and read this for the four decisions
that make it work.
|
|
When to use the SDK middleware instead
If a single conversation’s own three-tier context (recent messages,
observations, reflections) is enough — no cross-session recall, no entity
graph, no reasoning trail — start at
Integrate with the Vercel AI SDK
instead: it wires memory into one model call with no extra package, and you
can come back here when you want the UI. This guide’s richer retrieval
(entities, graph relationships, reasoning history, cross-session messages)
needs @neo4j-labs/nams-ai-provider; if you are deciding between the SDK
middleware and the provider’s other modes, see
Choose a NAMS AI provider mode.
What you need
-
Node.js 22+
-
A
MEMORY_API_KEYfrom memory.neo4jlabs.com -
An
OPENAI_API_KEY(model id viaOPENAI_MODEL, defaultgpt-5-mini) -
Optionally a
MEMORY_WORKSPACE_IDto scope the data to one workspace
Install from npm:
npm install @neo4j-labs/[email protected] @neo4j-labs/[email protected] \
ai@^7.0 @ai-sdk/openai@^4.0 @ai-sdk/react@^4.0 zod@^4.6.0 @neo4j-nvl/react@^1.2.1
The runnable example pins both @neo4j-labs/nams-ai-provider and
@neo4j-labs/agent-memory to file: links into this repository, so it always
runs against current source; an application installs the two npm releases
above instead.
Decision 1: put the conversation id in the URL
A conversation id is the only handle on a thread. Keep it in the route, not in
React state or localStorage, and the URL becomes shareable, reloadable and
resumable — because every message behind it is in NAMS, not in the browser.
app/page.tsx mints one and redirects. force-dynamic keeps the page off the
build-time prerender path, so next build succeeds on a machine with no API key:
export const dynamic = "force-dynamic";
export default async function Home() {
let conversationId: string | undefined;
try {
const conversation = await memoryClient().shortTerm.createConversation({
userId: userId(),
});
conversationId = conversation.id;
} catch (error) {
// redirect() signals by throwing, so it must stay outside this try.
return <SetupNotice message={String(error)} />;
}
redirect(`/c/${conversationId}`);
}
In Next 16 params is a promise — the synchronous form was removed:
export default async function ConversationPage({
params,
}: {
params: Promise<{ conversationId: string }>;
}) {
const { conversationId } = await params;
return <Workspace conversationId={conversationId} />;
}
Decision 2: send one turn, let memory supply the rest
app/api/chat/route.ts supplies the dependencies (MemoryClient, the memory
provider config, the model, the user id) and calls handleChat from
lib/handlers.ts; the route file itself does nothing but pass those four
dependencies to it. Inside it,
createNamsProvider({ …deps.memory, baseProvider: () ⇒ deps.model, scope: {
userId: deps.userId, conversationId } }) returns a ProviderV4, and
nams.languageModel("chat") wraps deps.model so relevant memory —
conversation history, cross-session messages, matching entities, their graph
relationships and past reasoning steps — is retrieved and prepended to the
last user message before the call, and the turn is persisted after it.
baseProvider ignores the model id it is handed and always returns
deps.model, so nams.languageModel("chat") is called with an arbitrary id;
there is no other memory-specific code in the route:
export async function handleChat(deps: ChatDeps, request: Request): Promise<Response> {
let body: ChatRequestBody;
try {
body = (await request.json()) as ChatRequestBody;
} catch {
return badRequest("Request body must be JSON.");
}
const conversationId = body.conversationId;
if (!conversationId) return badRequest("conversationId is required.");
const uiMessages = body.messages ?? [];
const latest = uiMessages.filter((m) => m.role === "user").slice(-1);
if (latest.length === 0) return badRequest("No user message to answer.");
// Memory supplies the history; the client only has to send the new turn.
const messages = await convertToModelMessages(latest);
const question = uiMessageText(latest[0]!);
const nams = createNamsProvider({
...deps.memory,
baseProvider: () => deps.model,
scope: { userId: deps.userId, conversationId },
});
const model = nams.languageModel("chat");
const result = streamText({
model,
instructions: deps.instructions ?? INSTRUCTIONS,
messages,
// `onEnd` is the AI SDK 7 name (`onFinish` is a deprecated alias). The
// reasoning step is what the trace drawer reads back — one row per answered
// turn, addressable by conversation.
onEnd: async ({ text }) => {
await recordTurn(deps.client, conversationId, question, text);
},
});
return result.toUIMessageStreamResponse();
}
Only the newest user turn is forwarded — the history the model sees comes
back out of the provider, so the browser is not the source of truth for the
conversation. onEnd (the AI SDK 7 name; onFinish is a deprecated alias)
hands the answered turn to the small recordTurn helper in the same file,
which records it as a reasoning step; that is the only call this route makes
to MemoryClient directly.
The createNamsProvider stream wrapper awaits the assistant-turn write as part
of closing the response stream, so a fully drained client stream implies the
write was attempted — not skipped, unlike a plain fire-and-forget write. It is
still best-effort: a failed write is logged and swallowed, not surfaced to the
caller, and nothing in the response tells the client whether it succeeded. For
a route that must coordinate storage with its lifecycle and give the client a
visible success signal, pass persistInteractions: false to
createNamsProvider and write the turn yourself, following the
application-owned write pattern.
That option turns off both of the provider’s writes, the user message as well
as the answer, so the route must store both: when the stream ends, call
deps.client.shortTerm.addMessage(conversationId, "user", question) and then
store the assembled answer, and keep those writes alive until they settle.
This is the same order the provider uses; a user message stored before the
model call can come back from the provider’s own session search and repeat the
question in the memory block. Writing only the answer leaves every user turn
out of the transcript and out of the provider’s later retrieval.
On the client, useChat with DefaultChatTransport sends the conversation id in
the request body:
const { messages, sendMessage, status } = useChat({
id: conversationId,
transport: new DefaultChatTransport({
api: "/api/chat",
body: { conversationId },
}),
});
To make the "reload and it is still there" claim true, hydrate the transcript
once on mount from shortTerm.getContext rather than from client state.
Decision 3: expand the graph lazily, with loadedIds
longTerm.getEntityGraph() gives the rail its first canvas.
longTerm.expandGraph(nodeId, loadedIds) grows it: pass every id already on
screen and the service answers with the delta only, so a double-click on a
well-connected node does not re-send the subgraph the user is looking at.
// POST /api/memory/graph
const expanded = await client.longTerm.expandGraph(nodeId, loadedIds);
Own the accumulated graph in the component that issues the request — that is what
keeps loadedIds honest — and union the delta in:
function merge(current: GraphPayload, delta: GraphPayload): GraphPayload {
const nodes = new Map(current.nodes.map((n) => [n.id, n]));
for (const node of delta.nodes) if (!nodes.has(node.id)) nodes.set(node.id, node);
// ... same for edges
return { nodes: [...nodes.values()], edges: [...edges.values()] };
}
expandGraph returns the viz-oriented shape ({ id, labels, properties } for
nodes), so flatten properties.name / properties.type for your renderer.
Neo4j’s visualization library touches window at import time, so load it with
next/dynamic and ssr: false:
const InteractiveNvlWrapper = dynamic(
() => import("@neo4j-nvl/react").then((m) => m.InteractiveNvlWrapper),
{ ssr: false },
);
NVL reports clicks, not double-clicks; treat a second onNodeClick on the same
node within ~350 ms as "expand".
Decision 4: wait for extraction instead of guessing
NAMS returns from a write before the entities in that message are searchable — extraction runs in a background pipeline. Refetch the graph immediately after an answer and you render a canvas that is one turn behind.
longTerm.waitForExtraction turns that race into an await. The predicate form is
the one to reach for in a UI: "resolve when an entity I have not already got
appears".
const known = new Set(knownIds);
const settled = await client.longTerm.waitForExtraction({
query, // the user's turn, as the probe
limit: 25,
timeoutMs: 20_000,
intervalMs: 1_500,
predicate: (entities) => entities.some((e) => !known.has(e.id)),
});
It returns false on timeout rather than throwing, so a quiet turn that produced
no new entities degrades to "nothing new" instead of an error. Validation and
transport errors still propagate. A new workspace entity can also come from
another conversation: this predicate is a UI refresh hint, not extraction
provenance for this turn. Order the rail’s
refresh as context → wait for extraction → refetch graph, and surface the middle
step as a badge so the reader can see the window open and close.
|
Prefer |
Reading the reasoning trail back
Because the route recorded a step per turn, reasoning.getTraceByConversation
gives you an audit drawer for free:
const trace = await client.reasoning.getTraceByConversation(conversationId);
// trace.steps: [{ id, reasoning, actionTaken, result, createdAt }, …]
Deploying
npx vercel deploy
Set MEMORY_API_KEY, OPENAI_API_KEY and optionally MEMORY_WORKSPACE_ID as
project environment variables. There is no database to provision. The route
handlers use only fetch, Request and Response, so they are edge-deployable —
pass apiKey to MemoryClient explicitly, because on edge runtimes
process.env is only populated inside the request scope. See
Deploy to edge runtimes.
Current-source example dependency declarations (not a release or live-service
verification): @neo4j-labs/nams-ai-provider 0.3.0, @neo4j-labs/agent-memory
0.5.0, next 16.3.4, react 19.3.0, ai 7.0.97, @ai-sdk/openai 4.0.65,
@ai-sdk/react 4.0.100, @chakra-ui/react 3.37.0, @neo4j-nvl/react 1.2.2,
Node.js 22 — 2026-09-23.
Testing it without an API key
Put each route’s behaviour in a (deps, Request) ⇒ Response function and keep
the route file a thin wrapper. Then the handlers can be driven in Vitest with the
real MemoryClient and createNamsProvider over a mocked network
(msw) and the AI SDK’s own MockLanguageModelV4 from
ai/test:
import { MockLanguageModelV4 } from "ai/test";
const client = new MemoryClient({ endpoint: ENDPOINT, apiKey: API_KEY });
const deps = { client, memory: { apiKey: API_KEY, endpoint: ENDPOINT }, model: mockModel(), userId };
const response = await handleChat(deps, request);
await response.text(); // drain the stream
expect(state.rolesOf(conversationId)).toEqual(["user", "assistant"]);
Mocking the network rather than the SDK means a wrong URL, verb or body shape fails the suite. Answer 501 for any endpoint you have not implemented so that a new call fails loudly instead of passing silently. This is the last check before shipping: a green suite here means the route, the provider wiring and the memory persistence contract all hold without needing a live API key.