Evaluate memory quality

Available on NAMS: Yes, with differences. NAMS supports this kind of memory, but not every call on this page: some steps run server-side, some calls take different arguments or return different shapes, and some are Bolt-only. Check the backend capabilities reference before you adapt a Bolt procedure.

How to run labeled regression tests against your memory graph using the evaluation harness.

The harness is a scaffold, not a benchmark — five dimensions, simple metrics, no opinionated dataset. Use it to detect regressions when you change extraction settings, dedup thresholds, or schema, not to make "library X scores Y on benchmark Z" claims.

Prerequisites

  • A connected client and a labeled test graph with the actual stored IDs.

  • A seeded Bolt graph for the retrieval, audit and preference cases below. The extraction and resolution cases run the client’s configured extractor and resolver instead, and are skipped when the client has none. The hosted client exposes eval, but its preference calls are unsupported and its graph does not establish the same TOUCHED audit data. Use only supported dimensions after verifying the backend’s schema and retrieval behavior.

  • Run the fragments in an async function with that client. The illustrative IDs below must be replaced with IDs returned by your fixture setup.

Dimensions

Dimension Metric

Retrieval relevance

Recall@k of client.long_term.search_entities(query, limit=k) against a labeled set of expected entity ids per query.

Audit completeness

Recall of (:Entity)←[:TOUCHED]-(:ReasoningStep) traversal against a labeled set of expected step ids per entity.

Preference fidelity

F1 score of client.long_term.get_preferences_for(user, active_only=True) against the expected active preference ids per user.

Extraction quality (0.7.0)

Relation F1 when a case declares expected_relations, entity micro-F1 otherwise. Runs the client’s configured extractor over the case text and matches one-to-one under the default alias-tolerant policy.

Resolution quality (0.7.0)

B-cubed F1 of the client’s configured resolver against labeled gold clusters. An OntologyResolver resolves each case’s mentions together with resolve_episode(); any other resolver goes through resolve_batch().

The two 0.7.0 dimensions need a component the client may not have: an extractor and a resolver respectively. When one is missing — for example a NAMS-backed client, or a Bolt client configured with resolution.strategy="none" — the dimension is skipped cleanly: its report field stays None and the dimension name is appended to report.skipped. A skipped dimension is excluded from overall_score rather than scoring zero, so a missing extractor cannot look like a regression.

report = await client.eval.run(suite)
if "extraction" in report.skipped:
    print("no extractor configured - extraction dimension not scored")

Goal

Run a suite of labeled cases and score each dimension:

from neo4j_agent_memory.memory.eval import (
    AuditCase,
    EvalSuite,
    ExtractionCase,
    PreferenceCase,
    ResolutionCase,
    RetrievalCase,
)

suite = EvalSuite(
    retrieval=[
        RetrievalCase(
            query="healthcare consultants",
            expected_entity_ids={"entity-anthem", "entity-sara"},
            k=5,
        ),
    ],
    audit=[
        AuditCase(
            entity_id="entity-anthem",
            expected_step_ids={"step-1", "step-2"},
        ),
    ],
    preference=[
        PreferenceCase(
            user_identifier="[email protected]",
            expected_active_pref_ids={"pref-senior-healthcare"},
        ),
    ],
    extraction=[
        ExtractionCase(
            text="Sara Okafor joined Anthem as a principal consultant.",
            expected_entities=[
                {"name": "Sara Okafor", "type": "PERSON"},
                {"name": "Anthem", "type": "ORGANIZATION"},
            ],
            expected_relations=[("Sara Okafor", "EMPLOYED_BY", "Anthem")],
        ),
    ],
    resolution=[
        ResolutionCase(
            mentions=[
                ("Anthem", "ORGANIZATION"),
                ("Anthem Inc", "ORGANIZATION"),
                ("Aetna", "ORGANIZATION"),
            ],
            gold_clusters={"Anthem": "a", "Anthem Inc": "a", "Aetna": "b"},
        ),
    ],
)

report = await client.eval.run(suite)
print(f"Overall: {report.overall_score:.2f}")
print(f"Retrieval recall: {report.retrieval.score:.2f}")
print(f"Audit recall:     {report.audit.score:.2f}")
print(f"Pref F1:          {report.preference.score:.2f}")
if report.extraction:
    print(f"Extraction F1:    {report.extraction.score:.2f}")
if report.resolution:
    print(f"Resolution B3 F1: {report.resolution.score:.2f}")
print(f"Skipped: {report.skipped}")

ExtractionCase coerces dict and tuple entries into ExpectedEntity / ExpectedRelation objects on construction, so gold data can be written inline or loaded from JSON. When a case declares expected_relations it is scored on relations only — split entity and relation expectations into separate cases if you want both numbers.

ResolutionCase.gold_clusters maps each mention name to a cluster id. How the predicted clusters are formed depends on the resolver:

  • An OntologyResolver (what resolution.strategy="composite" builds on Bolt) receives the whole case in one resolve_episode() call, so its second pass clusters variants of the same name among themselves, even on an empty database. A merged mention is keyed by what it merged onto: the stored entity’s id, or the first-seen mention it was clustered with. A created or review-band mention is its own singleton.

  • Any other resolver goes through resolve_batch(), which resolves each mention independently against the stored graph. Clusters come from each result’s cluster_id, falling back to canonical_name. Seed the canonical entities the case expects first; on an empty database every mention resolves to itself and B-cubed recall floors out for reasons unrelated to your thresholds.

Mentions sharing a name overwrite each other in both maps, so give repeated mentions distinct names when that matters. Each case’s detail records which call produced its clusters (clustered_by).

Steps

1. Build a labeled seedset

The labels are the hard part. Two reasonable starting points:

  • Capture from production: pick a handful of representative retrieval queries; for each, record the entity ids your team agrees are correct hits. Re-evaluate periodically.

  • Synthesize from fixtures: seed the database with a known graph (the examples/audit-trail/ pattern works well), then label expectations explicitly in test code.

The harness doesn’t know how you produced the labels — it just compares to whatever you provide.

Two properties of the library shape retrieval labels: add_entity embeds the entity name (the description is metadata, not retrieval signal), and search_entities applies a 0.7 similarity floor. A labelled query that is semantically close to the description but far from the name scores 0.

Make preference cases falsifiable — seed the superseded state too, and leave the superseded id out of the expected set:

user = "[email protected]"
old = await client.long_term.add_preference(
    "consultants", "Any seniority is fine for healthcare engagements",
    user_identifier=user,
)
new = await client.long_term.add_preference(
    "consultants", "Only principals and senior managers on payer accounts",
    user_identifier=user,
)
# Two positional ids — there is no `new_preference=` keyword.
await client.long_term.supersede_preference(old.id, new.id)

case = PreferenceCase(
    user_identifier=user,
    expected_active_pref_ids={str(new.id)},  # `old.id` must NOT come back
)

With MemorySettings.memory.multi_tenant=True, every scoped write must carry user_identifier= or raise — so seed two tenants and give each its own case, and a cross-tenant leak shows up as an F1 miss rather than a silent pass.

2. Run the suite

report = await client.eval.run(suite)

By default every dimension with cases is evaluated. To run a subset:

report = await client.eval.run(suite, dimensions=["audit"])

Skipped dimensions show as None on the report. A dimension with no cases is simply not run; one that had cases but no component to run them lands in report.skipped as well.

3. Inspect per-case detail

DimensionReport.details lists each case with its expected vs. actual ids, recall (or precision/recall/F1 for the preference dimension), and the case parameters. For a suite containing audit cases:

assert report.audit is not None
for d in report.audit.details:
    if d["recall"] < 1.0:
        print(f"Audit miss for entity {d['entity_id']}:")
        print(f"  expected = {d['expected']}")
        print(f"  actual   = {d['actual']}")

4. Wire into CI

Treat the suite as a regression test: fail the build below a threshold, and write a machine-readable report so the failure is diffable.

import json
from pathlib import Path

MIN_SCORE = 0.9
report = await client.eval.run(suite)
payload = {
    "overall": report.overall_score,
    "audit": report.audit.score if report.audit else None,
    "retrieval": report.retrieval.score if report.retrieval else None,
    "preference": report.preference.score if report.preference else None,
    "extraction": report.extraction.score if report.extraction else None,
    "resolution": report.resolution.score if report.resolution else None,
    "skipped": report.skipped,
}
Path("eval-report.json").write_text(json.dumps(payload, indent=2))

if report.overall_score < MIN_SCORE:
    raise SystemExit(f"Memory-quality regression: {report.overall_score:.2f}")

For a complete synthetic fixture and gate, save the complete harness files below. You need Python 3.10+, a POSIX shell, and a dedicated test Aura database. Each run resets the harness’s fixed demo users, session, and entity names. Keep this fixture separate from application data.

Create a local folder and virtual environment for the examples on this page. The commands reuse an existing environment without changing its files:

mkdir -p ~/agent-memory-tutorials
cd ~/agent-memory-tutorials
if [ -e .venv ]; then
  printf '%s\n' 'Using the existing virtual environment.'
else
  python3 -m venv .venv
fi
source .venv/bin/activate

Expected: ~/agent-memory-tutorials is your working directory and its virtual environment is active. Install the published SDK with the command below.

Each complete code block labelled Save as names a file to create in this folder using your editor. Copy the entire block, including imports and the entry point. Expand each helper disclosure and use Copy code to copy its full source. Keep all files together so their imports resolve.

When continuing from another tutorial or guide, retain the existing environment, configuration, session files, and .tutorial-state/. Reuse unchanged helper files; compare an existing file before replacing it, and finish any pending cleanup or recovery before changing the code that owns its state.

python -m pip install 'neo4j-agent-memory[sentence-transformers]==0.7.0'
export NEO4J_URI="neo4j+s://YOUR-TEST-INSTANCE.databases.neo4j.io"
export NEO4J_USERNAME="YOUR-AURA-USERNAME"
export NEO4J_PASSWORD="YOUR-AURA-PASSWORD"
export NEO4J_DATABASE="neo4j"
mkdir -p eval-harness

Use the credentials for your dedicated test instance from the Aura console. The first run downloads the local all-MiniLM-L6-v2 embedding model; the database connection still uses the network. No model API key is required.

Save both files in eval-harness/; ci_gate.py loads the sibling main.py.

Complete eval-harness/main.py
Save as eval-harness/main.py
Unresolved include directive in modules/ROOT/pages/how-to/evaluation.adoc - include::example$eval-harness/main.py[]
Complete eval-harness/ci_gate.py
Save as eval-harness/ci_gate.py
Unresolved include directive in modules/ROOT/pages/how-to/evaluation.adoc - include::example$eval-harness/ci_gate.py[]

Run the complete gate from agent-memory-tutorials/:

python eval-harness/ci_gate.py --min-score 0.9 --report eval-report.json

To run the gate in your application’s CI, copy your local eval-harness/ directory into your own project. The following workflow steps check out that project, install the published SDK, and use dedicated test Aura credentials from its repository secrets:

- uses: actions/checkout@v4
- uses: actions/setup-python@v5
  with:
    python-version: '3.12'
- run: python -m pip install 'neo4j-agent-memory[sentence-transformers]==0.7.0'
- run: python eval-harness/ci_gate.py --min-score 0.9 --report eval-report.json
  env:
    NEO4J_URI: ${{ secrets.NEO4J_URI }}
    NEO4J_USERNAME: ${{ secrets.NEO4J_USERNAME }}
    NEO4J_PASSWORD: ${{ secrets.NEO4J_PASSWORD }}
    NEO4J_DATABASE: ${{ secrets.NEO4J_DATABASE || 'neo4j' }}
- uses: actions/upload-artifact@v4
  if: always()
  with:
    name: eval-report
    path: eval-report.json

The complete eval-harness/ci_gate.py exits 0 at or above the threshold, 1 below it, and 2 when the run itself failed. Appending one JSONL row per run (python eval-harness/main.py --out trend.jsonl) gives you the score over time instead of a single snapshot.

5. Verify the regression gate

Run a suite whose expected IDs match the seeded graph, then deliberately replace one expected ID with an absent ID and confirm the relevant dimension falls. Check the JSON artifact and process exit code. An empty suite has no evaluated dimensions and an overall score of 0.0; require the intended cases to be present as well as checking the threshold.

What the harness is not

The harness reports whether this run’s scores moved against your own labels; it does not replace hand inspection of results or the :ConsolidationRun audit trail, which records change over time rather than current state. See Why the v0.2 operational primitives are opt-in for why the harness is opt-in and scaffold-shaped rather than a benchmark or a guarantee.

See also