Skip to content
AI-grafen
FAI engineeringMemory systems· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Evaluating memory systems

Be able to design tests that measure whether the memory improves the agent across sessions.

Prerequisites

Intuition

A memory system existing does not mean that it helps. The evaluation has to answer three questions:

  1. Does the system find the right memory? (retrieval) — probe questions about things said in earlier sessions.
  2. Does the agent use the memory? (the benefit) — the same task with and without the memory, measuring the difference in the result and in the number of questions the agent has to ask.
  3. Does the memory do harm? (the risk) — false memories, out-of-date facts used as current ones, and mix-ups between users.

The third is nearly always forgotten, and is the one that does the most damage in production.

Code

import numpy as np

def evaluate_memory(agent_with, agent_without, profiles, n_sessions=20, rng=None):
    rng = rng or np.random.default_rng(0)
    measures = {"benefit": [], "probe_recall": [], "false_memories": [], "mix_up": []}
    for p in profiles:
        res_with    = run_series(agent_with, p, n_sessions)
        res_without = run_series(agent_without, p, n_sessions)
        measures["benefit"].append(res_with["success"].mean() - res_without["success"].mean())

        # 1. Probe: facts that WERE SAID in an earlier session
        said = p.facts_said()
        measures["probe_recall"].append(np.mean([agent_with.remembers(f) for f in said]))

        # 2. False memories: facts that were NEVER said — the agent should answer no
        never = p.facts_never_said()
        measures["false_memories"].append(np.mean([agent_with.remembers(f) for f in never]))

        # 3. Mix-up: facts from ANOTHER user
        other = rng.choice([q for q in profiles if q is not p]).facts_said()
        measures["mix_up"].append(np.mean([agent_with.remembers(f) for f in other]))
    return {k: float(np.mean(v)) for k, v in measures.items()}

# A pass: benefit > 0 with a margin, a high probe_recall,
# false_memories ≈ 0, mix_up == 0 (not «low» — zero).

A mix-up between users is a privacy incident, not a quality measure. The measure has to be exactly 0, and the test should be run in CI. It is cheap to test and catastrophic to miss.

Mastery means

  • Designs tests that measure the memory's benefit across sessions
  • Measures the harms: false memories and mix-ups
  • Uses a control without memory

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences