FAI engineeringMemory systems· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
Evaluating memory systems
Be able to design tests that measure whether the memory improves the agent across sessions.
Prerequisites
- FLong-term evaluation of agentsrequired
Intuition
A memory system existing does not mean that it helps. The evaluation has to answer three questions:
- Does the system find the right memory? (retrieval) — probe questions about things said in earlier sessions.
- Does the agent use the memory? (the benefit) — the same task with and without the memory, measuring the difference in the result and in the number of questions the agent has to ask.
- Does the memory do harm? (the risk) — false memories, out-of-date facts used as current ones, and mix-ups between users.
The third is nearly always forgotten, and is the one that does the most damage in production.
Code
import numpy as np
def evaluate_memory(agent_with, agent_without, profiles, n_sessions=20, rng=None):
rng = rng or np.random.default_rng(0)
measures = {"benefit": [], "probe_recall": [], "false_memories": [], "mix_up": []}
for p in profiles:
res_with = run_series(agent_with, p, n_sessions)
res_without = run_series(agent_without, p, n_sessions)
measures["benefit"].append(res_with["success"].mean() - res_without["success"].mean())
# 1. Probe: facts that WERE SAID in an earlier session
said = p.facts_said()
measures["probe_recall"].append(np.mean([agent_with.remembers(f) for f in said]))
# 2. False memories: facts that were NEVER said — the agent should answer no
never = p.facts_never_said()
measures["false_memories"].append(np.mean([agent_with.remembers(f) for f in never]))
# 3. Mix-up: facts from ANOTHER user
other = rng.choice([q for q in profiles if q is not p]).facts_said()
measures["mix_up"].append(np.mean([agent_with.remembers(f) for f in other]))
return {k: float(np.mean(v)) for k, v in measures.items()}
# A pass: benefit > 0 with a margin, a high probe_recall,
# false_memories ≈ 0, mix_up == 0 (not «low» — zero).
A mix-up between users is a privacy incident, not a quality measure. The measure has to be exactly 0, and the test should be run in CI. It is cheap to test and catastrophic to miss.
Mastery means
- Designs tests that measure the memory's benefit across sessions
- Measures the harms: false memories and mix-ups
- Uses a control without memory
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Evaluating Very Long-Term Conversational Memory of LLM Agents — arXiv (open access; licence per article)
- arXiv — MemGPT: Towards LLMs as Operating Systems — arXiv (open access; licence per article)