How the NAMS AI provider retrieves memory
@neo4j-labs/nams-ai-provider is a separate, experimental Neo4j Labs package
that talks only to the hosted NAMS backend (see
choosing a mode). Provider and
middleware modes retrieve memory on every model call automatically; tools
mode exposes the same retrieval as a query_memory tool the model calls
explicitly. This page explains what that retrieval actually does — which
sources it searches, how it merges them, and why — so the shape of what
lands in the prompt is not a mystery.
Five sources, one search each
On every turn, the provider queries five sources in parallel:
-
Current conversation —
searchMessages(query, { sessionId: convId })against the active conversation. -
Long-term entities —
searchEntities(query, { limit: 5 })over the workspace’s Neo4j entity graph (facts, preferences, patterns). -
Reasoning traces —
listSteps(convId), filtered to steps whoseactionTakenis"direct response"and that carry areasoningfield, capped at 6. -
Cross-session messages — up to
crossSessionLimit(default 5) of the user’s other recent conversations, 2 requests each (searchMessagespluslistSteps, so a relevant reasoning step from a past session counts too). -
Graph — relationships around the long-term entities that matched. See below; this is the one source that depends on another source’s results, so it cannot run in the very first batch.
NAMS uses vector search where the workspace has embeddings, and a text match otherwise. Provider and middleware modes cache one turn’s memories and reuse them across the steps of a tool-calling loop, so a multi-step turn does not re-run all five searches per step.
Graph expansion
Entity search returns entities without their edges — a name and a
description, nothing about how that entity connects to anything else. This is
the retrieval a flat vector store cannot do: after the entity search, the
graphExpansionLimit best matches (default 2) are expanded with one
getEntity request each, and their stored relationships are turned into
triples added to the prompt as (from)-[TYPE]->(target) lines — the entity’s
canonical name, the relationship type, and the related entity’s name.
Two caps keep this bounded: at most 5 relationships are taken from any one
entity, so a single hub node cannot crowd out the rest of the prompt, and an
edge whose target has no stored name is dropped rather than shown as a bare
id. An entity that fails to expand is logged and skipped — the other four
sources still answer. graphExpansionLimit: 0 turns graph expansion off
entirely, leaving entities as flat text.
Merging without ranking
The five sources are not concatenated and then trimmed — they are merged by taking turns: one hit from long-term entities, one from graph triples, one from the current conversation, one from cross-session, one from reasoning, then back to the start, each source keeping the order NAMS itself returned it in, until the merged list reaches the cap.
This replaced an earlier design that sorted every hit by the entity’s stored
extraction confidence. That score measures how sure the extractor was that a
mention was an entity, not how relevant that entity is to the current
query — sorting by it discarded the backend’s own relevance ranking, and
because entities routinely scored higher than conversation messages, a small
maxMemories could fill entirely with entities and leave no room for the
conversation at all. Taking turns guarantees every source gets a chance
regardless of how the other sources happened to score.
The total is capped at maxMemories (default 6), itself hard-capped at 12
regardless of what a caller configures — Math.min(maxMemories, 12).
The prompt block
The merged hits are formatted as a numbered list — 1. [source] content
— with a trailing note explaining the [graph] triple notation, but only
when a graph hit is actually present (there is nothing to explain otherwise).
That block is prepended to the last user message’s content: for a
plain-text message, the memory block and a blank line go in front of the
original text in one string; for a multi-part message, a new text part
holding the memory block is inserted ahead of the existing parts. No memory
block is added when retrieval finds nothing.
What gets persisted after the call
Provider and middleware modes persist the turn automatically —
persistInteractions (default true) is the only switch. Hooks mode always
persists what onFinish() is given: persistUserPrompt: false on its scope
drops only the user turn, and a PreMemoryWrite hook can block or rewrite
the write. In provider and middleware modes the user’s message is saved once
per turn, not once per step of a tool-calling loop, and the assistant’s
answer is saved from the step that produces it. A step that calls an app
tool (one your application executes, not one the model provider runs) saves
no assistant text, even if that step also produced text — persistence waits
for a step with no app tool calls. With streamText, the text deltas are
accumulated and saved when the stream finishes. In tools mode, the model
decides whether to call store_memory at all — enforceQueryMemory() and
ensureMemoryStored() narrow but do not remove that gap; see
the API reference.
Entity extraction itself happens server-side, inside NAMS, from whatever
short-term messages get persisted — provider, middleware, and hooks modes
never call an extraction model. Tools mode is the exception: a store_memory
write of type fact, user_preference, or pattern goes straight to the
entity graph rather than through a conversation message, so there is no
server-side turn for NAMS to extract from. Passing extractionModel to
createNams/createNamsMemoryTools in tools mode runs a real entity
extraction over that memory and stores each named entity it finds. The
relationships it proposes are dropped with one logged warning, because the
hosted REST API accepts no relationship writes. createNamsProvider
and createNamsMemory do not accept it at all (removed in 0.3.0); setting it
on createNams() and then calling .wrap() or .hooks() is a no-op that
logs one warning, because extracting again on
the client would just duplicate what NAMS already does after persisting the
turn.
The literal-substring fallback
A whole natural-language question — "where do I work?" — often does not
match stored text well under a plain text search, and not every workspace has
embeddings configured for vector search. So the current-conversation and
long-term-entity searches retry when the full query returns nothing: they
take the query’s up to four longest significant words (at least 3 characters,
punctuation and trailing possessives stripped), search each one both as
typed and in Title Case, run those searches in parallel, deduplicate the
results, and then rank what is left by how many of the query’s words each
result shares. Cross-session and reasoning-step lookups do not retry this
way — they run once, as given.
See also
-
Expose memory as tools — where this retrieval surfaces as
query_memoryhits -
Understanding the three-tier context model — the single-conversation retrieval the SDK’s own middleware uses instead