Understanding the three memory types

The three memory layers organize different records: conversations, reusable knowledge, and recorded task execution. The names describe their roles, not automatic expiry or model learning.

Both backends use these concepts, but supported operations and scopes differ. See Backend capabilities.

The problem with stateless agents

A model call uses the context supplied for that call. Without retained records and a retrieval path, a later call cannot use an earlier conversation, a stated preference, or a recorded tool result.

Persistence solves only part of the problem. An application must also decide which records to retrieve and how to use them. Storing information does not by itself retrain the model or insert that information into future prompts.

Three complementary memory layers

Memory architecture: MemoryClient with three memory layers backed by Neo4j
Figure 1. MemoryClient with three memory layers backed by Neo4j
  • Short-term memory stores conversations and messages, preserving sequence and enabling supported retrieval of past exchanges. "Short-term" does not mean in-process or automatically deleted: these records are persisted.

  • Long-term memory represents reusable declarative knowledge, including entities and, on supported backend surfaces, preferences and facts. Relationships connect that knowledge to other records.

  • Reasoning memory stores the task steps, tool calls, observations, and outcomes that an application or integration records. It is an execution record, not access to a model’s private internal reasoning.

These are product terms. Human-memory analogies are approximate: both a conversation and a task trace can describe an episode; the current prompt is closer to working memory than any one persisted layer.

"Observation" above is a reasoning step’s recorded observation, often derived from a tool result. NAMS also uses "observation" and "reflection" in an unrelated sense, for the compressed tiers of a conversation’s short-term context, which both SDKs can read on the hosted backend. See Understanding the three-tier context model and the glossary for each sense.

Why separate the layers?

The layers have different retrieval patterns. Conversation history answers what was said and in what order. Entity retrieval answers what is known about a topic and how it relates to other topics. Trace retrieval answers what recorded actions led to an outcome.

Their structures also differ: messages form a sequence, entities form a connected knowledge graph, and traces contain steps and tool calls. An application’s retention and summarization policies may differ by record type. Persistence duration is a policy choice, not a fixed lifetime implied by the layer name.

How the layers connect

On Bolt (see Bolt and NAMS), :MENTIONS connects a message to an extracted entity, and a trace can link to its triggering message through :INITIATED_BY. These links retain the context in which knowledge or an action arose. An application must actually record or extract these relationships; the presence of the three layers does not populate them automatically.

Context assembly can combine relevant material from the supported layers for an LLM prompt. Its output and available filters are part of the client/backend contract. See MemoryClient and Backend capabilities.

Why connections help retrieval

A graph can connect an entity to the conversation that mentioned it and the task outcome that followed. This allows questions that combine content, relationships, and time, rather than considering each record independently.

Vector similarity can supply candidates while graph relationships and properties provide additional context. Those techniques answer different questions: similar text does not necessarily identify the same entity, and a nearby node is not automatically relevant to the current task. Graph memory architecture explains the tradeoffs.

From stored memory to useful context

The application chooses the actor responsible for each operation. It may load recent messages before a model call, retrieve entities for a domain question, or look up recorded outcomes from similar tasks. Middleware and agent tools can help, but their behavior must be explicitly connected to the application lifecycle.

Separate task patterns are documented in Store and search messages, Work with entities, and Record reasoning traces. Use the scope supported by the selected backend; a conversation identifier alone is not an authorization boundary.