Skip to content
AI-grafen
FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Long-term evaluation of agents

Be able to design evals that measure improvement over many sessions, not just per answer.

Prerequisites

Intuition

An agent with a memory should become better for you over time: remember your preferences, not repeat mistakes, build on earlier work. An eval per answer does not see that. You need session series.

The setup:

  • Simulated users with a hidden profile (goals, preferences, knowledge) who interact over 10–50 sessions. Measure whether the agent has learnt the profile: fewer questions, the right adaptation, no repeated mistakes.
  • Memory probes: after session k, ask about something from session j < k. Measure the recall and whether the agent uses it unprompted.
  • Measures over time: a curve of task success per session — the slope is the interesting part.
  • Harm control: false memories, mix-ups between users, a memory that makes the agent worse (locking a misconception in).
  • A control: the same agent without memory. The difference is the memory's value.

In AI-grafen the natural metric is «mastery per hour» over weeks, not points per exercise.

Code

import numpy as np

def run_series(agent, profile, n_sessions, tasks, rng):
    """Returns the success per session and the memory probe hits."""
    success, probes = [], []
    for k in range(n_sessions):
        t = tasks[k % len(tasks)]
        answer = agent.session(t, simulated_user(profile, rng))
        success.append(t.judge(answer))
        if k > 0:
            fact = profile.fact_revealed_in(session=rng.integers(0, k))
            probes.append(fact.lower() in agent.ask(f"What do you know about {fact.subject}?").lower())
    return np.array(success), np.array(probes)

def slope(y):
    x = np.arange(len(y)); return float(np.polyfit(x, y, 1)[0])

# the report: the slope with memory vs without, the probe recall, the share of false memories
# (a probe about facts that were NEVER said)

Mastery means

  • Designs evals that measure improvement over many sessions
  • Distinguishes per-answer quality from long-term benefit
  • Handles drift, memory and user dependence in the measurement

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences