FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
Long-term evaluation of agents
Be able to design evals that measure improvement over many sessions, not just per answer.
Prerequisites
- FEvals for language models and agentsrequired
- FSemantic memory and consolidationrequired
Intuition
An agent with a memory should become better for you over time: remember your preferences, not repeat mistakes, build on earlier work. An eval per answer does not see that. You need session series.
The setup:
- Simulated users with a hidden profile (goals, preferences, knowledge) who interact over 10–50 sessions. Measure whether the agent has learnt the profile: fewer questions, the right adaptation, no repeated mistakes.
- Memory probes: after session k, ask about something from session j < k. Measure the recall and whether the agent uses it unprompted.
- Measures over time: a curve of task success per session — the slope is the interesting part.
- Harm control: false memories, mix-ups between users, a memory that makes the agent worse (locking a misconception in).
- A control: the same agent without memory. The difference is the memory's value.
In AI-grafen the natural metric is «mastery per hour» over weeks, not points per exercise.
Code
import numpy as np
def run_series(agent, profile, n_sessions, tasks, rng):
"""Returns the success per session and the memory probe hits."""
success, probes = [], []
for k in range(n_sessions):
t = tasks[k % len(tasks)]
answer = agent.session(t, simulated_user(profile, rng))
success.append(t.judge(answer))
if k > 0:
fact = profile.fact_revealed_in(session=rng.integers(0, k))
probes.append(fact.lower() in agent.ask(f"What do you know about {fact.subject}?").lower())
return np.array(success), np.array(probes)
def slope(y):
x = np.arange(len(y)); return float(np.polyfit(x, y, 1)[0])
# the report: the slope with memory vs without, the probe recall, the share of false memories
# (a probe about facts that were NEVER said)
Mastery means
- Designs evals that measure improvement over many sessions
- Distinguishes per-answer quality from long-term benefit
- Handles drift, memory and user dependence in the measurement
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Evaluating Very Long-Term Conversational Memory of LLM Agents — arXiv (open access; licence per article)
- arXiv — Generative Agents: Interactive Simulacra of Human Behavior — arXiv (open access; licence per article)