Give your chatbot a memory
Be able to save and fetch earlier conversations so that the chatbot can use them.
Prerequisites
- CPython — files, CSV and JSONrequired
- DBuild a small search enginerequired
Intuition
A chatbot without a memory starts over every time. Two kinds of memory are needed, and they solve different problems:
| Short-term | Long-term | |
|---|---|---|
| What | the most recent messages in the conversation | facts and earlier conversations |
| Where | in the prompt | in a file or a database |
| The limit | the context window | the disk |
| Fetched | always | when needed, by searching |
Short-term is simple: send the most recent N messages along.
Long-term requires you to find the right thing. You cannot send a thousand earlier conversations along — you have to pick out the ones relevant to this particular question. And searching earlier conversations is exactly the same problem as searching documents.
That is why the search engine is the prerequisite: a chatbot memory is a small search engine over your own old conversations.
Code
import json, math, re, time
from pathlib import Path
from collections import Counter
class Memory:
def __init__(self, file="memory.jsonl", keep_latest=6):
self.file = Path(file)
self.keep = keep_latest
self.current = [] # short-term
self.old = [json.loads(r) for r in self.file.read_text(encoding="utf-8").splitlines()] \
if self.file.exists() else []
# --- short-term ---
def add(self, role, text):
self.current.append({"role": role, "text": text})
# --- long-term ---
def save_conversation(self, summary):
entry = {"time": time.strftime("%Y-%m-%d %H:%M"), "summary": summary,
"messages": self.current}
with self.file.open("a", encoding="utf-8") as f:
f.write(json.dumps(entry, ensure_ascii=False) + "\n")
self.old.append(entry)
self.current = []
def find(self, query, k=2):
"""Simple word overlap — the same idea as a small search engine."""
q = set(re.findall(r"\w+", query.lower()))
hits = []
for entry in self.old:
words = set(re.findall(r"\w+", entry["summary"].lower()))
shared = len(q & words)
if shared:
hits.append((shared / math.sqrt(len(words) + 1), entry))
hits.sort(key=lambda t: -t[0])
return [e for _, e in hits[:k]]
def build_prompt(self, system, query):
parts = [{"role": "system", "content": system}]
relevant = self.find(query)
if relevant:
earlier = "\n".join(f"- {e['time']}: {e['summary']}" for e in relevant)
parts.append({"role": "system",
"content": f"Earlier conversations that may be relevant:\n{earlier}"})
parts += [{"role": m["role"], "content": m["text"]}
for m in self.current[-self.keep:]]
parts.append({"role": "user", "content": query})
return parts
m = Memory("/tmp/memory.jsonl")
m.add("user", "My name is Ada and I am revising for a test on derivatives.")
m.add("assistant", "We can start with the chain rule.")
m.save_conversation("Ada is revising derivatives, started with the chain rule")
for d in m.build_prompt("You are a tutor.", "Can we carry on with derivatives?"):
print(d["role"], "→", d["content"][:70])
Three things to think about from the start:
| The question | Why it has to be answered |
|---|---|
| What is not saved? | personal data, passwords, sensitive information |
| For how long? | set a retention time and keep to it |
| Does the user see the memory? | they should be able to read, correct and delete it |
The last is both a GDPR requirement and good design: a memory the user cannot see or correct becomes unpleasant as soon as it contains something wrong.
Interactive
Build it and test it in a quarter of an hour. Use the code above and then do four tests:
- Does the short-term work? Say your name, ask three other questions, then ask «what is my name?».
- Does the long-term work? Save the conversation, restart the program, and ask about something from last time.
- Does it find the right thing? Save five different conversations about different subjects. Ask a question about one of them. Did the right conversation come along?
- What happens when the memory grows? Save fifty conversations. Do the hits get worse?
What you will probably discover in tests 3 and 4:
- Word overlap finds the right thing when the question uses the same words as the summary, and misses entirely on paraphrases. «How is the maths going?» does not find «Ada is revising derivatives».
- The more conversations, the more irrelevant hits with some word in common.
Two improvements, in order:
- Better summaries. Let the model write them, with an instruction to include the subject, the names and the decisions. That gives more to match against.
- Embeddings instead of word matching. Then paraphrases are found. That is exactly the step from this node to RAG.
It is worth feeling the limitation yourself before moving on — it explains why vector search exists.
Mastery means
- Saves conversations in a file or a database
- Fetches relevant earlier conversations
- Distinguishes short-term from long-term memory
Sign in to do the exercises and build your mastery up.
Sources
- The Python documentation (PSF licence) — PSF
- IMY — data protection for children and young people (in Swedish) — myndighetsmaterial