Integrate with the Vercel AI SDK

The TypeScript client ships a language-model middleware that auto-injects context and persists messages around any Vercel AI SDK call. It implements specification v4 (LanguageModelV4Middleware from @ai-sdk/provider 4.x), which is the provider specification ai 7.x uses. This page covers the wiring; a runnable example lives at typescript/examples/vercel-ai.

ai 7 requires Node.js 22 or newer, and so does this client.

Prefer not to wire the middleware yourself?

For cross-session search, graph-expanded retrieval, and explicit memory tools or lifecycle hooks, see choosing a NAMS AI provider mode — a separate package built on the same NAMS backend. This middleware injects only the current conversation’s three-tier context; the provider package also searches other conversations and expands matched entities' relationships into the prompt.

Prerequisites: a configured model provider and NAMS test-workspace credentials.

Procedure

  1. Build and install the current-source example using its README; use Node.js 22+.

  2. Resolve an authorized conversation id before constructing the middleware.

  3. Apply the wrapper below, then choose an explicit persistence lifecycle for streaming.

What the middleware does

For every model invocation:

  1. Fetches three-tier context for the conversation and prepends it to the prompt — reflections and observations as one system message, recent messages replayed as user/assistant turns with { type: "text" } parts.

  2. Persists the user’s input message before generation. A multi-part user message (text interleaved with files) is flattened to its text parts.

  3. Persists the assistant’s response after generation — for generateText via wrapGenerate, and for streamText via wrapStream, which accumulates the text deltas and writes once the stream finishes.

If three-tier context is unavailable (bridge transport, or a service error), the middleware falls back to flat conversation history. Every persistence failure is non-fatal: the model call still returns.

In a tool-calling loop transformParams runs once per step with the same user turn still last in the prompt; the middleware writes that turn once, not once per step.

Basic wiring

import { generateText, wrapLanguageModel } from "ai";
import { openai } from "@ai-sdk/openai";
import { MemoryClient } from "@neo4j-labs/agent-memory";
import { agentMemoryMiddleware } from "@neo4j-labs/agent-memory/middleware/vercel-ai";

const memory = new MemoryClient();
const conv = await memory.shortTerm.createConversation({ userId: "alice" });

const model = wrapLanguageModel({
  model: openai(process.env.OPENAI_MODEL ?? "gpt-5-mini"),
  middleware: agentMemoryMiddleware(memory, { conversationId: conv.id }),
});

const { text } = await generateText({
  model,
  prompt: "Hello!",
});

Options

agentMemoryMiddleware(client, {
  conversationId: "conv-uuid",     // or a function returning one
  userId: "alice",                  // used if conversation is lazily created
  includeContext: true,             // default true on REST
  historyLimit: 20,                 // flat-history fallback cap
  persistInput: true,               // default true
  persistResponses: true,           // default true
});

conversationId can be a string or a function. The function is evaluated when the middleware is constructed, not once per model call. Construct a middleware instance after resolving each request/conversation id; do not share one instance across users expecting the function to re-read request state.

Resuming instead of always creating

Conversation history is visible to this middleware when the next process selects the same conversation. Resolve the id before wiring the middleware — createConversation on every run gives every run new conversation context:

const userId = process.env.DEMO_USER_ID ?? "alice";

// Prefer an explicit id; otherwise choose a conversation in application code.
const requested = process.env.CONVERSATION_ID;
const existing = requested
  ? await memory.shortTerm.getConversationMetadata(requested)
  : (await memory.shortTerm.listConversations({ userId }))
      .filter((conversation) => conversation.metadata?.source === "my-app")
      .sort((a, b) => b.createdAt.localeCompare(a.createdAt))[0];

const conv =
  existing ??
  (await memory.shortTerm.createConversation({ userId, metadata: { source: "my-app" } }));

console.log(existing ? `Resumed ${conv.id}` : `Created ${conv.id}`);

In a server, the id normally comes from the request (a thread id in your own schema); the env-var form above is what a script uses.

Streaming

The default wrapStream is fire-and-forget: it exposes no persistence-completion promise, so draining the model stream does not establish that NAMS stored the answer. See the full derivation for why and what that means for cancellation, errors and runtime shutdown.

For a CLI that must finish storing before exit, own the streaming assistant write and disable automatic assistant persistence to avoid duplicates. Add streamText to the existing ai import, then use:

const streamingModel = wrapLanguageModel({
  model: openai(process.env.OPENAI_MODEL ?? "gpt-5-mini"),
  middleware: agentMemoryMiddleware(memory, {
    conversationId: conv.id,
    persistResponses: false,
  }),
});
const result = streamText({ model: streamingModel, prompt: "Explain graph memory." });
let answer = "";
for await (const chunk of result.textStream) {
  answer += chunk;
  process.stdout.write(chunk);
}
// The application owns and awaits this write; automatic assistant writes are off.
await memory.shortTerm.addMessage(conv.id, "assistant", answer);

Handle model errors and cancellation before persisting a partial answer. At the edge, register the actual application write promise with the runtime lifecycle; see edge runtime persistence and the Cloudflare example’s explicit end/abort/error handlers.

Types

agentMemoryMiddleware returns LanguageModelV4Middleware, so wrapLanguageModel accepts it with no cast. @ai-sdk/provider is an optional peer dependency of this package: anything that has ai 7 installed already has it, and a project that does not use this middleware does not need it.

Edge runtimes

Pass apiKey explicitly because process.env is not available at module-init scope on Cloudflare Workers and Vercel Edge:

export default {
  async fetch(req, env) {
    const memory = new MemoryClient({ apiKey: env.MEMORY_API_KEY });
    // ... wire middleware as above
  },
};

See deploy to edge runtimes for details.

Compatibility

The published @neo4j-labs/[email protected] package on npm implements provider specification v4 (LanguageModelV4Middleware from @ai-sdk/provider 4.x), the specification ai 7 uses, on Node.js 22 or newer. Projects still on ai 4 to 6 should pin @neo4j-labs/agent-memory to a 0.4.x release.

Verify the integration

In typescript/examples/vercel-ai/, run npm run lint and npm test. The tests check real REST routing for the complete lessons and two-process recall with a mock model. For a live check, retain the conversation id and read back both messages after generation; model wording alone does not prove persistence.