Working memory: context, summary, window
Be able to handle a long dialogue with summarisation and a sliding window.
Prerequisites
Intuition
A conversation grows until the context window runs out. Four strategies, which can be combined:
| The strategy | How | Loses |
|---|---|---|
| A sliding window | keep the N most recent messages | everything older, abruptly |
| Summarisation | compress the older parts into a summary | details and exact wordings |
| Retrieval over the history | fetch relevant older messages when needed | the connection between them |
| Structured state | extract facts into a small JSON that always comes along | everything not extracted |
What must never be lost is kept outside the window: the system prompt, the user's stated preferences, and the state of the ongoing task. Put them in a fixed part of the prompt — not in the history.
Code
class WorkingMemory:
def __init__(self, llm, tok, token_cap=6000, keep_latest=8):
self.llm, self.tok = llm, tok
self.cap, self.keep = token_cap, keep_latest
self.summary = ""
self.messages: list[dict] = []
self.state: dict = {} # fixed, compressed, never lost
def _tokens(self, texts):
return sum(len(self.tok(t)["input_ids"]) for t in texts)
def add(self, role, content):
self.messages.append({"role": role, "content": content})
self._compress_if_needed()
def _compress_if_needed(self):
texts = [m["content"] for m in self.messages] + [self.summary]
if self._tokens(texts) <= self.cap:
return
old = self.messages[:-self.keep]
if not old:
return
self.summary = self.llm(
"Update the summary of the conversation. Keep decisions, preferences, open questions "
"and everything the user has asked for. At most 200 words.\n\n"
f"PREVIOUSLY: {self.summary}\n\nNEW MESSAGES:\n" +
"\n".join(f"{m['role']}: {m['content']}" for m in old), temperature=0)
self.messages = self.messages[-self.keep:]
def prompt(self, system):
parts = [{"role": "system", "content": system}]
if self.state:
parts.append({"role": "system", "content": f"The current state: {self.state}"})
if self.summary:
parts.append({"role": "system", "content": f"Earlier in the conversation: {self.summary}"})
return parts + self.messages
Two things that often go wrong: the summary loses what the user explicitly asked for (solve it with an explicit instruction about what has to be preserved), and the compression happens in the middle of a tool call (solve it by never cutting between a call and its result).
Mastery means
- Handles a long dialogue within the context window
- Chooses between a sliding window, summarisation and retrieval
- Preserves what must not be lost
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — MemGPT: Towards LLMs as Operating Systems — arXiv (open access; licence per article)
- arXiv — Lost in the Middle: How Language Models Use Long Contexts — arXiv (open access; licence per article)