Memory retrieval: when should the memory be fetched?
Be able to design when and how memories are fetched and measure whether they are actually used.
Prerequisites
- EHybrid search and RRFrequired
- FEpisodic memory for agentsrequired
Intuition
Having a memory is one thing. Knowing when it should be used is another, and harder.
Three strategies:
| The strategy | How | The problem |
|---|---|---|
| Always fetch | search the memory at every message | expensive, and irrelevant memories get in the way |
| The model decides | give it a search_memory tool | it often forgets to use it |
| Triggers | fetch at certain signals | it requires the signals to be right |
In practice a combination is used: always fetch cheaply and broadly, but filter hard on relevance before anything is put in the prompt.
The decisive insight: an irrelevant memory is worse than no memory at all. It takes up room in the context, draws the model's attention, and can make it answer the wrong question. Precision matters more than recall here — the opposite of ordinary RAG.
Formal
What should be weighed into the ranking:
| The factor | Why |
|---|---|
| Semantic similarity | is the memory about the same thing? |
| Freshness | newer memories are usually more relevant |
| Importance | some memories are central (names, goals, preferences) |
| Frequency of use | memories that have often proved useful |
A common formula (from Generative Agents) combines them linearly:
The freshness decays exponentially — the same form as the forgetting curve in the platform's mastery model.
An absolute threshold, not just the top k. Fetch the best and require the score to exceed a threshold. Without the threshold you always get memories, even when none is relevant, and then the memory harms more than it helps.
Measure whether the memories are actually used. That is the measurement that is nearly always missing:
| The measure | How |
|---|---|
| The fetch rate | the share of turns where something is fetched |
| The use rate | does the answer refer to what was fetched? |
| The benefit rate | does the answer get better with than without? (an A/B on the same question) |
| The disturbance rate | does the answer get worse with the memory? |
| Precision@k | the share of fetched memories that were relevant |
The benefit rate is most easily measured with an ablation: run the same question with and without the fetched memories and let a judge compare. If the difference is near zero the memory system is doing no good, however good the retrieval measures look.
Handling conflicts. Two memories can contradict each other — «I do not like maths» from the spring and «maths is actually fun now» from last week. Rules that work:
- The newer wins on a direct contradiction.
- Explicit statements weigh more than derived conclusions.
- When uncertain: include both and let the model see the contradiction.
The third is often best — the model can handle «you said X in the spring but Y last week» more gracefully than a system that silently chooses.
The writing side matters at least as much. Do not save every utterance. Extract what is worth remembering — decisions, preferences, goals, recurring difficulties — and write a short, searchable summary. A memory that is hard to find is no memory.
Code
import math, time
from dataclasses import dataclass, field
@dataclass
class Memory:
text: str
embedding: list
created: float
importance: float = 0.5 # 0-1, set when writing
used: int = 0
last_used: float = 0.0
class MemoryBank:
def __init__(self, alpha=0.6, beta=0.2, gamma=0.2, half_life_days=30.0):
self.memories: list[Memory] = []
self.alpha, self.beta, self.gamma = alpha, beta, gamma
self.lam = math.log(2) / (half_life_days * 86400)
def write(self, text, embedding, importance=0.5):
self.memories.append(Memory(text, embedding, time.time(), importance))
def _score(self, m, query_vector, now):
similarity = sum(a * b for a, b in zip(m.embedding, query_vector))
freshness = math.exp(-self.lam * (now - m.created))
return self.alpha * similarity + self.beta * freshness + self.gamma * m.importance
def fetch(self, query_vector, k=3, threshold=0.45):
now = time.time()
scored = [(self._score(m, query_vector, now), m) for m in self.memories]
scored.sort(key=lambda p: -p[0])
# BOTH the top k AND an absolute threshold — otherwise k are always fetched
chosen = [(p, m) for p, m in scored[:k] if p >= threshold]
for _, m in chosen:
m.used += 1
m.last_used = now
return chosen
# Measure the benefit with an ablation: the same question with and without the memories
def benefit_rate(llm, judge, cases, bank):
better = worse = equal = 0
for c in cases:
memories = bank.fetch(c["query_vector"])
with_mem = llm(c["query"], context=[m.text for _, m in memories])
without = llm(c["query"], context=[])
verdict = judge(c["query"], with_mem, without) # "with" | "without" | "equal"
better += verdict == "with"; worse += verdict == "without"; equal += verdict == "equal"
n = len(cases)
return {"better_with_memory": round(better / n, 3),
"worse_with_memory": round(worse / n, 3),
"equal": round(equal / n, 3),
"net_benefit": round((better - worse) / n, 3)}
# a net benefit near 0 → the memory system makes no difference
# a high worse_with_memory → irrelevant memories are getting in the way
# Conflict handling: show the contradiction instead of choosing silently
def format_memories(chosen):
rows = []
for p, m in chosen:
days = int((time.time() - m.created) / 86400)
rows.append(f"- ({days} days ago, relevance {p:.2f}) {m.text}")
return "Earlier in your conversations:\n" + "\n".join(rows)
# The writing side: extract, do not save everything
EXTRACT = """What in this conversation is worth remembering for next time?
Answer in JSON: {"memories": [{"text": "...", "importance": 0.0-1.0}]}
Include decisions, preferences, goals and recurring difficulties.
Do NOT include small talk, personal data or things that only hold today."""
benefit_rate is the measurement that decides whether the memory system is worth its complexity. A system with an excellent precision@3 but a net benefit of 0.02 does nothing in practice.
Mastery means
- Decides when the memory should be fetched
- Chooses and ranks the memories
- Measures whether the fetched memories are actually used
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Generative Agents: Interactive Simulacra of Human Behavior — arXiv (open access; licence per article)
- arXiv — MemGPT: Towards LLMs as Operating Systems — arXiv (open access; licence per article)
- arXiv — Lost in the Middle: How Language Models Use Long Contexts — arXiv (open access; licence per article)