Wrap a model in middleware mode

The @neo4j-labs/nams-ai-provider package talks only to the hosted Neo4j Agent Memory Service (NAMS) at https://memory.neo4jlabs.com — it has no Bolt equivalent for a self-managed Neo4j database.

Middleware mode adds memory to a model instance you already configure elsewhere, instead of asking you to construct the model through a NAMS provider. Use it when your model selection lives in application code you don’t control from the memory wiring — a shared model factory, a framework adapter, or a model chosen per request. This page walks through a complete, runnable program that teaches the agent a fact in one turn and recalls it in a fresh model instance for the same user.

A runnable copy of this program, alongside the other three modes, lives at typescript/examples/nams-ai-provider.

Prerequisites:

  • Node.js 22 or newer (the package’s declared floor).

  • A MEMORY_API_KEY from memory.neo4jlabs.com.

  • An OPENAI_API_KEY. The program uses @ai-sdk/openai by default; to use another provider, pass its model factory to the middlewareModeDemo() call at the bottom of the file.

  • A Node.js project set up as an ES module, with the packages installed:

    npm init -y   # only if the directory has no package.json yet
    npm pkg set type=module
    npm install @neo4j-labs/[email protected] @neo4j-labs/[email protected] ai@^7 @ai-sdk/openai@^4 zod@^4
    npm install --save-dev tsx

The program

Save as middleware-mode.ts
/**
 * Middleware mode: keep the model you already configure elsewhere and add
 * memory with `createNams().wrap()`. Turn 1 teaches the agent a fact. Turn 2
 * uses a brand-new model instance for the same user, so the answer must come
 * back out of NAMS.
 *
 * Run with:
 *
 *   MEMORY_API_KEY=nams_... OPENAI_API_KEY=sk-... npx tsx src/middleware-mode.ts
 *
 * Expected output (wording varies):
 *
 *   --- Turn 1 -- teach it something
 *   user:      My favourite programming language is Rust.
 *   assistant: Great choice! Rust is loved for its safety and performance. ...
 *
 *   --- Turn 2 -- fresh model instance, same user
 *   user:      What is my favourite programming language?
 *   assistant: Your favourite programming language is Rust.
 */

import { realpathSync } from 'node:fs';
import { pathToFileURL } from 'node:url';
import { createNams, type NamsProviderOptions } from '@neo4j-labs/nams-ai-provider';
import { openai } from '@ai-sdk/openai';
import { ToolLoopAgent, stepCountIs } from 'ai';

const userId = process.env.NAMS_DEMO_USER ?? 'demo-user-middleware-mode';
const modelId = process.env.NAMS_DEMO_MODEL ?? 'gpt-5.4-mini';

/**
 * Run the teach-then-recall demo. `buildModel` defaults to OpenAI; tests pass
 * a mock model factory here so the demo can run without a real API call.
 */
export async function middlewareModeDemo(
  buildModel: NamsProviderOptions['baseProvider'] = openai,
): Promise<{ taught: string; recalled: string }> {
  const rawApiKey = process.env.MEMORY_API_KEY;
  if (!rawApiKey) {
    throw new Error(
      'Set MEMORY_API_KEY before running this example. Get a free key at https://memory.neo4jlabs.com',
    );
  }
  const apiKey: string = rawApiKey; // narrowed here so the closure below sees `string`, not `string | undefined`

  async function turn(label: string, message: string): Promise<string> {
    // A new model each turn -- memory comes from NAMS, not from local state.
    const nams = createNams({
      apiKey,
      endpoint: process.env.MEMORY_ENDPOINT,
      workspaceId: process.env.MEMORY_WORKSPACE_ID,
    });
    const wrappedModel = nams.wrap(buildModel(modelId), { userId });

    const agent = new ToolLoopAgent({
      model: wrappedModel,
      instructions: 'You are a helpful assistant.',
      stopWhen: stepCountIs(1), // no tools needed in middleware mode
    });

    const { text } = await agent.generate({ prompt: message });

    console.log(`\n--- ${label}`);
    console.log(`user:      ${message}`);
    console.log(`assistant: ${text}`);
    return text;
  }

  const taught = await turn('Turn 1 -- teach it something', 'My favourite programming language is Rust.');
  console.log('Stored: both sides of turn 1, saved to NAMS by the wrapped model.');

  const recalled = await turn('Turn 2 -- fresh model instance, same user', 'What is my favourite programming language?');
  console.log('Recalled: NAMS retrieved turn 1 and injected it before the model saw turn 2.');

  return { taught, recalled };
}

// Run only when executed directly, not when a test imports this file.
const entry = process.argv[1];
if (entry && import.meta.url === pathToFileURL(realpathSync(entry)).href) {
  middlewareModeDemo().catch(err => {
    console.error(err);
    process.exit(1);
  });
}

How the program works

  1. Read the demo userId and model id from the environment, falling back to demo-user-middleware-mode and gpt-5.4-mini.

  2. Require MEMORY_API_KEY. middlewareModeDemo throws immediately with a link to get a key if it’s unset, then narrows the value to string.

  3. Define a turn(label, message) helper that builds a fresh model each turn — memory comes from NAMS, not from any local state — and runs one generation:

    • createNams({ apiKey, endpoint, workspaceId }) builds the memory client. endpoint and workspaceId are read from optional environment variables and are undefined unless set, in which case the client falls back to its own defaults (the hosted endpoint, no workspace scoping). Neither maxMemories nor persistInteractions is passed here, so both take their defaults: 6 memories per turn (capped at 12 by the package regardless of what you configure) and persistence on.

    • nams.wrap(buildModel(modelId), { userId }) wraps the model instance with the memory middleware and returns an ordinary AI SDK model — buildModel defaults to openai but is a parameter so the test suite can substitute a mock model.

    • new ToolLoopAgent({ model: wrappedModel, instructions, stopWhen: stepCountIs(1) }) — stepCountIs(1) because middleware mode needs no tool-calling loop; retrieval and persistence happen inside the wrapped model itself.

    • agent.generate({ prompt: message }) runs the call. The wrapped model retrieves memories for the last user message and prepends them to it before the model runs, then saves the user message and the answer after it returns. How the NAMS AI provider retrieves memory covers what is searched, what the memory block looks like, and when a step’s answer is not saved.

  4. Call turn() twice:

    • Turn 1 ("teach it something") states a preference. The reply is persisted automatically.

    • Turn 2 ("fresh model instance, same user") builds a new wrapped model — no JavaScript object is shared with turn 1 — for the same userId, then asks a question whose answer only exists in what NAMS stored from turn 1.

    A fresh model instance is not a new NAMS conversation. Neither turn passes a conversationId, so both continue the user’s most recent NAMS conversation (turn 1 creates it on the first run) and write to the same conversation. To give a session its own conversation, create one and pass its id in the scope argument of wrap(); see Choose the conversation each instance writes to.

  5. Return { taught, recalled } so a caller (here, the offline test suite) can assert on both answers. The import.meta.url check at the bottom runs the demo only when the file is executed directly, not when it’s imported.

Run and verify

Run the saved file with both keys set:

MEMORY_API_KEY=nams_... OPENAI_API_KEY=sk-... npx tsx middleware-mode.ts

To run the copy in the repository instead, follow the setup in the example’s README, then run npm run middleware-mode from typescript/examples/nams-ai-provider/. The Run with comment at the top of the program shows the path of that copy.

Expected output (wording varies with the model):

--- Turn 1 -- teach it something
user:      My favourite programming language is Rust.
assistant: Great choice! Rust is loved for its safety and performance. ...
Stored: both sides of turn 1, saved to NAMS by the wrapped model.

--- Turn 2 -- fresh model instance, same user
user:      What is my favourite programming language?
assistant: Your favourite programming language is Rust.
Recalled: NAMS retrieved turn 1 and injected it before the model saw turn 2.

The second answer is the actual check: nothing in the script passes "Rust" to turn 2 directly, so the model can only have gotten it from what NAMS stored and re-injected.