Evaluate memory quality
|
Available on NAMS: Yes, with differences. NAMS supports this kind of memory, but not every call on this page: some steps run server-side, some calls take different arguments or return different shapes, and some are Bolt-only. Check the backend capabilities reference before you adapt a Bolt procedure. |
How to run labeled regression tests against your memory graph using the evaluation harness.
The harness is a scaffold, not a benchmark — five dimensions, simple metrics, no opinionated dataset. Use it to detect regressions when you change extraction settings, dedup thresholds, or schema, not to make "library X scores Y on benchmark Z" claims.
Prerequisites
-
A connected client and a labeled test graph with the actual stored IDs.
-
A seeded Bolt graph for the retrieval, audit and preference cases below. The extraction and resolution cases run the client’s configured extractor and resolver instead, and are skipped when the client has none. The hosted client exposes
eval, but its preference calls are unsupported and its graph does not establish the sameTOUCHEDaudit data. Use only supported dimensions after verifying the backend’s schema and retrieval behavior. -
Run the fragments in an async function with that
client. The illustrative IDs below must be replaced with IDs returned by your fixture setup.
Dimensions
| Dimension | Metric |
|---|---|
Retrieval relevance |
Recall@k of |
Audit completeness |
Recall of |
Preference fidelity |
F1 score of |
Extraction quality (0.7.0) |
Relation F1 when a case declares |
Resolution quality (0.7.0) |
B-cubed F1 of the client’s configured resolver against labeled gold clusters. An |
The two 0.7.0 dimensions need a component the client may not have: an extractor
and a resolver respectively. When one is missing — for example a NAMS-backed
client, or a Bolt client configured with resolution.strategy="none" — the
dimension is skipped cleanly:
its report field stays None and the dimension name is appended to
report.skipped. A skipped dimension is excluded from overall_score rather
than scoring zero, so a missing extractor cannot look like a regression.
report = await client.eval.run(suite)
if "extraction" in report.skipped:
print("no extractor configured - extraction dimension not scored")
Goal
Run a suite of labeled cases and score each dimension:
from neo4j_agent_memory.memory.eval import (
AuditCase,
EvalSuite,
ExtractionCase,
PreferenceCase,
ResolutionCase,
RetrievalCase,
)
suite = EvalSuite(
retrieval=[
RetrievalCase(
query="healthcare consultants",
expected_entity_ids={"entity-anthem", "entity-sara"},
k=5,
),
],
audit=[
AuditCase(
entity_id="entity-anthem",
expected_step_ids={"step-1", "step-2"},
),
],
preference=[
PreferenceCase(
user_identifier="[email protected]",
expected_active_pref_ids={"pref-senior-healthcare"},
),
],
extraction=[
ExtractionCase(
text="Sara Okafor joined Anthem as a principal consultant.",
expected_entities=[
{"name": "Sara Okafor", "type": "PERSON"},
{"name": "Anthem", "type": "ORGANIZATION"},
],
expected_relations=[("Sara Okafor", "EMPLOYED_BY", "Anthem")],
),
],
resolution=[
ResolutionCase(
mentions=[
("Anthem", "ORGANIZATION"),
("Anthem Inc", "ORGANIZATION"),
("Aetna", "ORGANIZATION"),
],
gold_clusters={"Anthem": "a", "Anthem Inc": "a", "Aetna": "b"},
),
],
)
report = await client.eval.run(suite)
print(f"Overall: {report.overall_score:.2f}")
print(f"Retrieval recall: {report.retrieval.score:.2f}")
print(f"Audit recall: {report.audit.score:.2f}")
print(f"Pref F1: {report.preference.score:.2f}")
if report.extraction:
print(f"Extraction F1: {report.extraction.score:.2f}")
if report.resolution:
print(f"Resolution B3 F1: {report.resolution.score:.2f}")
print(f"Skipped: {report.skipped}")
ExtractionCase coerces dict and tuple entries into ExpectedEntity /
ExpectedRelation objects on construction, so gold data can be written inline
or loaded from JSON. When a case declares expected_relations it is scored on
relations only — split entity and relation expectations into separate cases if
you want both numbers.
ResolutionCase.gold_clusters maps each mention name to a cluster id. How the
predicted clusters are formed depends on the resolver:
-
An
OntologyResolver(whatresolution.strategy="composite"builds on Bolt) receives the whole case in oneresolve_episode()call, so its second pass clusters variants of the same name among themselves, even on an empty database. A merged mention is keyed by what it merged onto: the stored entity’s id, or the first-seen mention it was clustered with. A created or review-band mention is its own singleton. -
Any other resolver goes through
resolve_batch(), which resolves each mention independently against the stored graph. Clusters come from each result’scluster_id, falling back tocanonical_name. Seed the canonical entities the case expects first; on an empty database every mention resolves to itself and B-cubed recall floors out for reasons unrelated to your thresholds.
Mentions sharing a name overwrite each other in both maps, so give repeated
mentions distinct names when that matters. Each case’s detail records which
call produced its clusters (clustered_by).
Steps
1. Build a labeled seedset
The labels are the hard part. Two reasonable starting points:
-
Capture from production: pick a handful of representative retrieval queries; for each, record the entity ids your team agrees are correct hits. Re-evaluate periodically.
-
Synthesize from fixtures: seed the database with a known graph (the
examples/audit-trail/pattern works well), then label expectations explicitly in test code.
The harness doesn’t know how you produced the labels — it just compares to whatever you provide.
Two properties of the library shape retrieval labels: add_entity embeds
the entity name (the description is metadata, not retrieval signal), and
search_entities applies a 0.7 similarity floor. A labelled query that is
semantically close to the description but far from the name scores 0.
Make preference cases falsifiable — seed the superseded state too, and leave the superseded id out of the expected set:
user = "[email protected]"
old = await client.long_term.add_preference(
"consultants", "Any seniority is fine for healthcare engagements",
user_identifier=user,
)
new = await client.long_term.add_preference(
"consultants", "Only principals and senior managers on payer accounts",
user_identifier=user,
)
# Two positional ids — there is no `new_preference=` keyword.
await client.long_term.supersede_preference(old.id, new.id)
case = PreferenceCase(
user_identifier=user,
expected_active_pref_ids={str(new.id)}, # `old.id` must NOT come back
)
With MemorySettings.memory.multi_tenant=True, every scoped write must
carry user_identifier= or raise — so seed two tenants and give each its
own case, and a cross-tenant leak shows up as an F1 miss rather than a
silent pass.
2. Run the suite
report = await client.eval.run(suite)
By default every dimension with cases is evaluated. To run a subset:
report = await client.eval.run(suite, dimensions=["audit"])
Skipped dimensions show as None on the report. A dimension with no cases is
simply not run; one that had cases but no component to run them lands in
report.skipped as well.
3. Inspect per-case detail
DimensionReport.details lists each case with its expected vs. actual
ids, recall (or precision/recall/F1 for the preference dimension), and
the case parameters. For a suite containing audit cases:
assert report.audit is not None
for d in report.audit.details:
if d["recall"] < 1.0:
print(f"Audit miss for entity {d['entity_id']}:")
print(f" expected = {d['expected']}")
print(f" actual = {d['actual']}")
4. Wire into CI
Treat the suite as a regression test: fail the build below a threshold, and write a machine-readable report so the failure is diffable.
import json
from pathlib import Path
MIN_SCORE = 0.9
report = await client.eval.run(suite)
payload = {
"overall": report.overall_score,
"audit": report.audit.score if report.audit else None,
"retrieval": report.retrieval.score if report.retrieval else None,
"preference": report.preference.score if report.preference else None,
"extraction": report.extraction.score if report.extraction else None,
"resolution": report.resolution.score if report.resolution else None,
"skipped": report.skipped,
}
Path("eval-report.json").write_text(json.dumps(payload, indent=2))
if report.overall_score < MIN_SCORE:
raise SystemExit(f"Memory-quality regression: {report.overall_score:.2f}")
For a complete synthetic fixture and gate, save the complete harness files below. You need Python 3.10+, a POSIX shell, and a dedicated test Aura database. Each run resets the harness’s fixed demo users, session, and entity names. Keep this fixture separate from application data.
Create a local folder and virtual environment for the examples on this page. The commands reuse an existing environment without changing its files:
mkdir -p ~/agent-memory-tutorials
cd ~/agent-memory-tutorials
if [ -e .venv ]; then
printf '%s\n' 'Using the existing virtual environment.'
else
python3 -m venv .venv
fi
source .venv/bin/activate
Expected: ~/agent-memory-tutorials is your working directory and its virtual environment is active. Install the published SDK with the command below.
Each complete code block labelled Save as names a file to create in this folder using your editor. Copy the entire block, including imports and the entry point. Expand each helper disclosure and use Copy code to copy its full source. Keep all files together so their imports resolve.
When continuing from another tutorial or guide, retain the existing environment, configuration, session files, and .tutorial-state/. Reuse unchanged helper files; compare an existing file before replacing it, and finish any pending cleanup or recovery before changing the code that owns its state.
python -m pip install 'neo4j-agent-memory[sentence-transformers]==0.7.0'
export NEO4J_URI="neo4j+s://YOUR-TEST-INSTANCE.databases.neo4j.io"
export NEO4J_USERNAME="YOUR-AURA-USERNAME"
export NEO4J_PASSWORD="YOUR-AURA-PASSWORD"
export NEO4J_DATABASE="neo4j"
mkdir -p eval-harness
Use the credentials for your dedicated test instance from the
Aura console. The first run downloads the local
all-MiniLM-L6-v2 embedding model; the database connection still uses the
network. No model API key is required.
Save both files in eval-harness/; ci_gate.py loads the sibling main.py.
Complete eval-harness/main.py
eval-harness/main.pyUnresolved include directive in modules/ROOT/pages/how-to/evaluation.adoc - include::example$eval-harness/main.py[]
Complete eval-harness/ci_gate.py
eval-harness/ci_gate.pyUnresolved include directive in modules/ROOT/pages/how-to/evaluation.adoc - include::example$eval-harness/ci_gate.py[]
Run the complete gate from agent-memory-tutorials/:
python eval-harness/ci_gate.py --min-score 0.9 --report eval-report.json
To run the gate in your application’s CI, copy your local eval-harness/
directory into your own project. The following workflow steps check out that
project, install the published SDK, and use dedicated test Aura credentials
from its repository secrets:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: python -m pip install 'neo4j-agent-memory[sentence-transformers]==0.7.0'
- run: python eval-harness/ci_gate.py --min-score 0.9 --report eval-report.json
env:
NEO4J_URI: ${{ secrets.NEO4J_URI }}
NEO4J_USERNAME: ${{ secrets.NEO4J_USERNAME }}
NEO4J_PASSWORD: ${{ secrets.NEO4J_PASSWORD }}
NEO4J_DATABASE: ${{ secrets.NEO4J_DATABASE || 'neo4j' }}
- uses: actions/upload-artifact@v4
if: always()
with:
name: eval-report
path: eval-report.json
The complete eval-harness/ci_gate.py exits
0 at or above the threshold, 1 below it, and 2 when the run itself
failed. Appending one JSONL row per run
(python eval-harness/main.py --out trend.jsonl) gives
you the score over time instead of a single snapshot.
5. Verify the regression gate
Run a suite whose expected IDs match the seeded graph, then deliberately
replace one expected ID with an absent ID and confirm the relevant dimension
falls. Check the JSON artifact and process exit code. An empty suite has no
evaluated dimensions and an overall score of 0.0; require the intended cases
to be present as well as checking the threshold.
What the harness is not
The harness reports whether this run’s scores moved against your own
labels; it does not replace hand inspection of results or the
:ConsolidationRun audit trail, which records change over time rather
than current state. See Why the v0.2 operational primitives are opt-in for
why the harness is opt-in and scaffold-shaped rather than a benchmark or a
guarantee.
See also
-
Tune entity resolution — using the resolution dimension to compare threshold settings.
-
Evaluation harness example — the runnable example.