Understanding the three-tier context model
The hosted service (NAMS) assembles a conversation’s context as three tiers
of increasing compression: recent messages, observations, and reflections.
Every TypeScript how-to page that calls getContext() or wires the Vercel AI
SDK, Mastra, or AWS Strands integration assumes this model. This page explains
what each tier is, where it comes from, and how the three fit together.
Both SDKs read the tiers from NAMS. The TypeScript client exposes
client.shortTerm.getContext(), getObservations(), and getReflections().
The Python client on the NAMS backend exposes client.short_term.get_context(),
get_observations(), and get_reflections(); its get_context() returns the
three sections formatted as one text block rather than a structured object.
This is a hosted-service concept. The tiers are computed by NAMS, so
they are not available on Python’s Bolt backend, where get_observations()
and get_reflections() raise NotSupportedError, or to a TypeScript client on
BridgeTransport, since the TCK bridge protocol has no equivalent operation.
The Python MCP server has a separate, local observational memory with its own
observations and reflections; see the MCP server’s observational memory.
|
The three tiers
-
Recent messages — the raw, persisted turns of the conversation (
Message[]), returned uncompressed. This is the same dataaddMessage/getConversationwork with elsewhere in the client. -
Observations — auto-generated summaries of a window of messages (
getObservations(conversationId)). AnObservationmay carry optionalwindowStart/windowEndtimestamps identifying the message span it summarizes. -
Reflections — higher-level syntheses derived from observations (
getReflections(conversationId)), generated once the conversation’s accumulated context crosses a compression threshold. A reflection summarizes a set of observations the way an observation summarizes a set of messages — one more level of compression, one step further from the original wording.
Compression increases from top to bottom: messages are exact, observations compress spans of messages, and reflections compress spans of observations. An application reading only reflections gets the most condensed view of a long conversation; reading recent messages gets the most literal one.
How getContext() assembles them
client.shortTerm.getContext(conversationId) returns all three tiers
together as a ConversationContext:
interface ConversationContext {
reflections: Reflection[];
observations: Observation[];
recentMessages: Message[];
}
The TypeScript integrations use it in different ways:
-
The Vercel AI SDK middleware folds reflections and observations into one prepended system message ("Relevant memory for this conversation") and replays
recentMessagesas ordinary turns. IfgetContext()fails, or the client is on the bridge transport, or the context comes back empty, it falls back to flat conversation history fromgetConversation()instead of failing the model call. -
The AWS Strands
ConversationManagerprepends one system message per reflection and one per observation to the messages that its inner manager (a sliding window by default) already holds. It does not replayrecentMessages. IfgetContext()fails, it injects nothing and leaves the inner manager’s messages unchanged. -
The Mastra recipe does no prompt assembly: the application calls
getContext()itself and decides what to do with the result.
See Troubleshooting for what to do when the context tiers are unavailable.
Other senses of "observation"
The reasoning-memory layer also records an "observation" on each reasoning
step, alongside the step’s thought and action (see
Understanding the three memory types
and the glossary's Reasoning memory entry). A step
observation is often derived from a tool result, for example with
record_tool_call(auto_observation=True), but it describes one recorded task
execution, not a compressed summary of a conversation. Nothing here implies
that reasoning-step observations feed the context-tier observations
described above.
The Python MCP server’s memory_get_observations tool uses a third sense. Its
local observational memory tracks each session inside the server process,
extracts observations (facts and decisions) from stored user
messages, and generates reflections once the session’s accumulated context
crosses a token threshold. It works on the Bolt backend and does not call the
NAMS tiers. See MCP tools.