Procedural memory: learnt skills
Be able to let an agent save and reuse approaches that work.
Prerequisites
- ETool userequired
- FConsolidation and forgetting in memory systemsrequired
Intuition
Episodic memory saves what happened. Semantic memory saves what is true. Procedural memory saves how something is done — a skill.
For an agent that means: when a task has been solved successfully, save the approach as a reusable procedure. Next time a similar task turns up the procedure is fetched and followed, instead of the agent groping its way to the same solution again.
The gain is double: fewer steps (cheaper, faster) and higher reliability (a tried recipe beats improvisation).
Voyager (Wang et al. 2023) showed the principle concretely in Minecraft: the agent wrote reusable skills as code, saved them in a library, and gradually built more advanced skills on top of simpler ones.
Code
from dataclasses import dataclass, field
@dataclass
class Procedure:
name: str
when: str # a description of when it applies — what is matched semantically
steps: list[dict] # tool calls in order, with parameters as templates
succeeded: int = 0
failed: int = 0
@property
def reliability(self):
n = self.succeeded + self.failed
return (self.succeeded + 1) / (n + 2) # Laplace smoothing: new procedures get ~0.5
class ProcedureLibrary:
def __init__(self, embed, min_reliability=0.6):
self.embed, self.procedures, self.min_r = embed, [], min_reliability
def save(self, task, trace, llm):
"""Only called after a VERIFIABLY successful run."""
p = llm.generalise(task, trace) # abstract the concrete values into parameters
if any(similarity(p.name, q.name) > 0.9 for q in self.procedures):
return None # it already exists
self.procedures.append(Procedure(**p, succeeded=1))
return p["name"]
def fetch(self, task, k=2):
candidates = [p for p in self.procedures if p.reliability >= self.min_r]
return sorted(candidates, key=lambda p: -cos(self.embed(task), self.embed(p.when)))[:k]
def feedback(self, name, succeeded):
for p in self.procedures:
if p.name == name:
p.succeeded += succeeded; p.failed += (not succeeded)
Three rules that make the difference between benefit and harm:
- Only save verifiably successful solutions — otherwise a library of mistakes is built.
- Follow the reliability up. A procedure that stops working (an API changed) should fall and drop out automatically.
- The procedure is a suggestion, not a compulsion. The agent should be able to deviate when the situation differs — otherwise it becomes rigid.
Mastery means
- Lets an agent save approaches that work
- Reuses them on similar tasks
- Measures that the reuse actually helps
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Voyager: An Open-Ended Embodied Agent with Large Language Models — arXiv (open access; licence per article)
- arXiv — Generative Agents: Interactive Simulacra of Human Behavior — arXiv (open access; licence per article)