Ingest documents and inspect extracted entities (TypeScript)

We will store two synthetic documents as conversation messages, wait for a specific extracted entity, and inspect graph records linked to those messages. This lesson demonstrates ingestion and inspection. The service’s extraction pipeline decides which entities and relationships it produces; this lesson does not promise a particular relationship or multi-hop traversal result.

Step 1: Build the SDK and install the lesson project

You need a workspace-scoped MEMORY_API_KEY for a dedicated test workspace on NAMS. This lesson does not call a model, so OPENAI_API_KEY is not required.

Use Git, Node.js 22.9 or later with npm, and a POSIX shell. The lesson commands load .env with Node’s --env-file-if-exists flag, added in Node.js 22.9; the SDK itself supports Node.js 22. These lessons share one repository checkout. The example-based lessons build the SDK from source so they run against current source; application code, including the MCP lesson’s standalone my-memory-mcp project, installs the published @neo4j-labs/[email protected] package from npm. Keep this checkout for the remaining TypeScript lessons.

For your first lesson, clone and build the SDK:

git clone https://github.com/neo4j-labs/agent-memory.git
cd agent-memory
export AGENT_MEMORY_TUTORIAL_ROOT="$PWD"
cd typescript
npm ci
npm run build
cd "$AGENT_MEMORY_TUTORIAL_ROOT"

Expected: the build finishes without errors and typescript/dist/index.js exists. The example-based lessons use this built SDK; installing an example does not build it. The MCP lesson’s project uses the npm package instead, so it does not depend on this build.

Continuing from another lesson, keep your existing checkout and skip the clone/build block above. From any directory inside that checkout, run:

cd "$(git rev-parse --show-toplevel)"
export AGENT_MEMORY_TUTORIAL_ROOT="$PWD"

Both entry paths finish at the repository root. If you changed the SDK source since the preceding lesson, rebuild it before continuing with an example-based lesson.

Enter the shared lesson project — its package.json pins the SDK with file:../.. so it always runs against the checkout you just built. On its first use, install its locked dependencies; subsequent lessons reuse them. Create the tutorial environment file only if it does not already exist:

cd "$AGENT_MEMORY_TUTORIAL_ROOT/typescript/examples/vercel-ai"
if [ ! -d node_modules ]; then npm ci; fi
if [ ! -e .env ]; then
  (umask 077; set -C; cat .env.tutorial.example > .env)
fi
chmod 600 .env
npm run lint

Expected: typechecking finishes without errors. It does not authenticate or write to NAMS. Edit the private .env with the keys required by this lesson; keep it out of version control. An existing .env is preserved: add missing tutorial settings without replacing its other values. For the two model lessons, set OPENAI_MODEL=gpt-4o-mini and supply OPENAI_API_KEY; an existing demo configuration may select a different model.

If an earlier install was interrupted, or the lockfile changed, run npm ci again from this project directory, then npm run lint. The directory check above only avoids an unnecessary reinstall; it does not prove the dependencies are complete. npm ci replaces node_modules from the lockfile and preserves your .env and .tutorial-state/ files.

Keep a private note of the selected NAMS workspace name/ID and its key owner. The optional MEMORY_WORKSPACE_ID is a selector, not proof of ownership. A workspace-scoped key may not need it. Keep the same endpoint, key and selector for the run’s authenticated verification and cleanup; do not print the key or use a different workspace to work around a failed check.

The tutorial template uses gpt-4o-mini for model lessons. DEMO_USER_ID, newest-conversation reuse and the gpt-5-mini default in .env.example belong to the separate npm start demo. These lessons create or explicitly resume their own conversations instead.

All remaining commands run in typescript/examples/vercel-ai/. --env-file-if-exists=.env loads that file; exported shell variables take precedence, so keep them consistent with the selected workspace. The project sets ESM and supplies the compiler configuration; no global TypeScript install is needed.

Step 2: Inspect the complete ingestion program

Open src/tutorials/knowledge-graph.ts. The program adds a run-specific suffix to the fictional organization name to reduce confusion with an earlier run. It obtains entity IDs from the saved messages' graph links, rather than searching unrelated workspace data. In general, the hosted service can merge a newly created entity onto an existing one it considers the same; when that happens, the entity you read back carries the canonical name rather than the one you passed in, and metadata.nams_resolution records the merge (see the TypeScript SDK reference).

import { TutorialRun, isTutorialEntryPoint } from "../../../shared/tutorial-state.js";
import { inspectGraph, runTutorialCommand, waitForTerminalExtraction } from "../../../shared/tutorial-cleanup.js";

export async function inspectDocuments(run: TutorialRun, timeoutMs = 60_000) {
  const conversation = await run.createConversation();
  console.log(`CONVERSATION_ID=${conversation.id}`);
  await run.bulkAddMessages([
    { role: "user", content: `${run.name} is a fictional organization that designs garden sensors.` },
    { role: "user", content: `Mira Vale ${run.state.runId} is an engineer at ${run.name}.` },
  ]);
  return observeDocuments(run, timeoutMs);
}

// This phase reads the existing run; it never creates another conversation or message.
export async function observeDocuments(run: TutorialRun, timeoutMs = 60_000) {
  const stored = await run.verify();
  console.log(`Stored document messages verified: ${stored.messages}`);
  await waitForTerminalExtraction(run, timeoutMs);
  const graph = await inspectGraph(run);
  const entityIds = [...new Set(graph.map(row => String(row.entity_id)))];
  let expectedNameFound = false;
  for (const id of entityIds) {
    // These IDs came from the owned messages' provenance edges, not workspace search.
    run.retain({ kind: "entity", id, origin: "derived", disposition: "present" });
    const entity = await run.client.longTerm.getEntity(id);
    expectedNameFound ||= entity.name.toLowerCase() === run.name.toLowerCase();
    console.log(`Linked entity: ${entity.id} ${entity.name} (${entity.type})`);
  }
  console.log(`Run graph: ${graph.length} message/entity links`);
  if (!entityIds.length) throw new Error("Extraction finished without linked entities. Keep the state and inspect the messages before claiming graph success.");
  if (!expectedNameFound) throw new Error("Linked entities did not include this run's expected organization name. Keep the state; graph presence alone is not the lesson's success condition.");
  return { conversationId: run.conversationId, entityIds };
}

if (isTutorialEntryPoint(import.meta.url)) {
  runTutorialCommand("knowledge-graph", process.argv.slice(2), "seed", ["seed", "observe"],
    (mode, run) => mode === "observe" ? observeDocuments(run) : inspectDocuments(run))
    .catch(error => { console.error(error); process.exitCode = 1; });
}

The maintained state and cleanup helpers are in typescript/examples/shared/. Keep the private .tutorial-state/ directory with this checkout. Each run binds its saved IDs to the endpoint, workspace selector and credential; a changed identity stops authenticated use of those IDs; local inspect remains available. Reseeding an existing state is refused.

TutorialRun and runTutorialCommand are local teaching helpers, not exports from the published SDK. They keep a private ledger around these public APIs:

In this lesson Public SDK operation underneath

TutorialRun.create(…​)

Constructs new MemoryClient(…​) with REST configuration; saves local run identity.

run.createConversation()

client.shortTerm.createConversation({ userId, metadata }); records the returned ID.

run.addMessage(…​) / run.bulkAddMessages(…​)

client.shortTerm.addMessage(…​) / bulkAddMessages(…​); records IDs and content digests.

run.verify()

Reads getConversationMetadata(…​) and getConversation(…​) to check ownership and stored messages.

run.addEntity() in the model lesson

client.longTerm.addEntity(…​); retains the create result, including merge evidence.

run.trackMiddlewareWrites() / run.verifyTurn(…​)

Tutorial-only tracking around middleware addMessage(…​) calls, followed by getConversation(…​) readback of new IDs, roles and exact content.

The helper also retains uncertain operations and scopes cleanup to this run. When adapting the lesson, use run.client to see the actual SDK calls, and keep an equivalent ownership and failure policy before removing the teaching helpers. See the TypeScript API reference for public methods and the middleware guide for agentMemoryMiddleware and its best-effort storage behavior.

Step 3: Store documents and inspect extraction

npx tsx --env-file-if-exists=.env \
  src/tutorials/knowledge-graph.ts seed .tutorial-state/knowledge-graph.json

Expect Stored document messages verified: 2. The program waits up to 60 seconds for those messages' extraction to finish, then inspects graph links starting from only their saved IDs. At least one linked entity must match the run’s exact organization name, ignoring case. A nonempty graph alone is not success.

A timeout means terminal extraction was not observed in time; it does not mean the messages were not stored. Missing or unexpected entities also fail the lesson check and retain state. Do not seed another run to bypass either result.

After a timeout, observe the same run again in a new process:

npx tsx --env-file-if-exists=.env \
  src/tutorials/knowledge-graph.ts observe .tutorial-state/knowledge-graph.json

observe verifies the saved messages, waits for terminal extraction and repeats the exact-name and scoped graph checks. It creates no conversation, message or entity; it can add discovered entity IDs to the private local ledger. A normal exit establishes that the expected organization was found. If the named predicate still fails after extraction finishes, retain that result for review and cleanup; graph presence alone is insufficient.

Inspect Linked entity: and Run graph: output alongside the saved message IDs and relationship names. This scoped view does not enumerate unrelated workspace data or promise a particular relationship or extraction-quality score. Cleanup independently rechecks terminal status before considering deletion.

Verify in a fresh process

Run these commands from the same directory, using the state created above:

npx tsx src/tutorials/knowledge-graph.ts \
  inspect .tutorial-state/knowledge-graph.json
npx tsx --env-file-if-exists=.env \
  src/tutorials/knowledge-graph.ts verify .tutorial-state/knowledge-graph.json

inspect shows local saved IDs and operation status without loading .env or contacting NAMS. verify opens another authenticated client and reads the recorded conversation and exact message IDs, roles and content digests. Expect Verified stored messages: with the conversation ID and message count. Verification makes no model calls and creates no records. It establishes persistence across client processes; it does not imply that the hosted service restarted.

Clean up and inspect the disposition

From the same lesson directory, clean up this run using its saved state:

npx tsx --env-file-if-exists=.env \
  src/tutorials/knowledge-graph.ts cleanup .tutorial-state/knowledge-graph.json
npx tsx src/tutorials/knowledge-graph.ts \
  inspect .tutorial-state/knowledge-graph.json

The cleanup command waits for the recorded messages' extraction to finish, checks the entities derived from those messages and any explicitly created entities, and deletes only records with proven exclusive ownership. It verifies deleted records are absent. Shared, merged or uncertain records remain listed; inspect their disposition instead of treating partial cleanup as a complete reset. A timeout or failed readback stops cleanup and preserves the state.

Keep the state until every resource has a verified disposition. Closing the client alone does not delete data. Do not remove the state file to bypass an unfinished run or bulk-delete workspace entities.

Read the cleanup result before starting another run. These are illustrative shapes; your IDs and counts differ:

Cleanup: {"complete":true,"retainedIds":[],"residualIds":[]}
Cleanup: {"complete":false,"retainedIds":["<entity-id>"],"residualIds":[]}
Result Next action

complete: true, exit status 0

Keep the original ledger as the completed record. A new run uses a new state path.

complete: false, exit status 2

Inspect each retained resource’s reason. Ask the workspace owner to review its ownership; do not delete it by name or force a workspace reset.

Error, exit status 1

Preserve the ledger and error. A timeout can be retried against the same state; an uncertain write or residual ID needs review before further writes or deletion.

inspect reads only the local ledger and works without .env, a current key, or a service connection. It reports IDs, operation status and cleanup dispositions, omitting credentials, credential fingerprints, workspace selectors and message digests. It does not establish the current remote state. After key rotation, verify, record and cleanup still refuse a changed identity. Do not edit the ledger’s identity to bypass that guard.

For an uncertain or retained result, privately give the workspace owner the run ID, recorded resource IDs, operation statuses, selected workspace note and the failing command. Retain any original MCP create receipt. Do not send keys or the raw ledger. The owner must resolve the exact records before a cleanup claim; a new key or a repeated write is not evidence that the earlier attempt failed.

Only after a completed run, start the next exercise with a distinct filename:

npx tsx --env-file-if-exists=.env \
  src/tutorials/knowledge-graph.ts seed .tutorial-state/knowledge-graph-2.json

Use that new filename for all subsequent commands for the new run. Preserve the original ledger and any receipts; do not remove them to make seed, teach or init accept a used path.