Deploy to edge runtimes

The core @neo4j-labs/agent-memory REST client can run on Cloudflare Workers and Vercel Edge. The package uses fetch only, has zero runtime dependencies, and defensively guards every process.env lookup behind a typeof process !== "undefined" check.

The one gotcha: API key resolution

Edge runtimes expose environment variables via the request handler scope, not process.env. The zero-config form

const client = new MemoryClient();   // reads MEMORY_API_KEY from process.env

reads environment variables only when that runtime provides process.env. Pass the key explicitly on runtimes using request-scoped bindings:

// Cloudflare Workers
export default {
  async fetch(req: Request, env: { MEMORY_API_KEY: string }) {
    const client = new MemoryClient({ apiKey: env.MEMORY_API_KEY });
    // ...
  },
};
// Vercel Edge — pages/api or route handler
export const runtime = "edge";

export async function GET(req: Request) {
  const client = new MemoryClient({ apiKey: process.env.MEMORY_API_KEY });
  // Vercel Edge does expose process.env in the handler scope,
  // but reading it explicitly avoids any surprise.
  // ...
}

What works on edge

  • The REST-supported MemoryClient operations (see the capability map)

  • The Vercel AI SDK middleware

  • The MCP tool dispatcher (handleMemoryToolCall)

  • The LangChain JS and Mastra adapters

What doesn’t (and why)

  • The TCK bridge transport (./testing subpath) is intended for Node-based conformance testing; we don’t test it on edge.

  • User-Agent overrides via headers: { "User-Agent": …​ } may be silently stripped by Cloudflare Workers due to a runtime restriction on setting that header in fetch. The default User-Agent is best-effort.

Streaming + memory persistence

Non-streaming generateText uses wrapGenerate, which awaits its assistant write. Streaming uses wrapStream: it launches a background write and exposes no completion promise. Draining or awaiting the model stream does not prove that write has finished, and waitUntil(Promise.resolve(result)) cannot make it so.

For a Worker, use the application-owned lifecycle in the Cloudflare example:

  1. Disable automatic assistant persistence (persistResponses: false).

  2. On a successful stream end, call addMessage once with the assembled answer.

  3. Register the actual write promise with ctx.waitUntil before returning the response.

  4. Settle the lifecycle on abort/error too. The example skips partial responses on abort and logs write failures; it does not guarantee storage if the runtime is terminated.

This avoids duplicate assistant writes. It is application code, not a hidden SDK flush method. persistResponses: false turns off only the assistant write; the middleware still stores the user message. createNamsProvider from @neo4j-labs/nams-ai-provider has no such split: its persistInteractions: false turns off both writes, so with that provider also store the user message yourself, just before the assembled answer when the stream ends. Keep the client alive until those writes settle. See the middleware streaming example.

Verifying

A unit-level edge sanity test ships with the package; it constructs a client in an environment where process is undefined. To exercise the full path against a real Worker, deploy with wrangler and watch tail logs for request-id-tagged failures.

See enable request logging for tracing.