EUniversityAI product development· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN
Prompts as code: versioning and testing
Be able to version prompts, test them in CI and measure regressions.
Prerequisites
- DThe context window, system prompts and few-shotrequired
- EBuild an eval harnessrequired
Intuition
A prompt in production is code. It determines the system's behaviour, it breaks, and it therefore has to be treated accordingly:
| Code | A prompt |
|---|---|
| in git | in git — not in a database or a Google document |
| versioned | versioned, with an id in the logs |
| tested in CI | an eval suite in CI with a threshold |
| code-reviewed | reviewed by somebody else |
| rolled back | can be rolled back in minutes |
The most common source of error in LLM products is that somebody «just changed a phrase» in a prompt, everything got worse, and nobody could say when or why.
Code
prompts/
├── tutor_explain/
│ ├── v3.md # the current one
│ ├── v2.md
│ └── cases.jsonl # the eval cases belonging to this particular prompt
└── registry.yaml # which version is active per environment
from pathlib import Path
import hashlib, yaml
REG = yaml.safe_load(Path("prompts/registry.yaml").read_text())
def get_prompt(name: str, environment: str = "prod") -> tuple[str, str]:
version = REG[name][environment] # "v3", for instance
text = Path(f"prompts/{name}/{version}.md").read_text(encoding="utf-8")
return text, f"{name}:{version}:{hashlib.sha256(text.encode()).hexdigest()[:8]}"
prompt, prompt_id = get_prompt("tutor_explain")
# log the prompt_id with every call → which behaviour an answer came out of can be traced
# .github/workflows/ci.yml (or scripts/ci.sh)
- run: python -m evals.run --prompts prompts/ --threshold 0.85 --fail-on-regression
Three rules that go a long way: the prompt id in every log line, the eval suite run on every change to a prompt file, and every new prompt version compared against the active one — not just against the threshold. A prompt that clears the threshold but loses three previously passing cases is a regression.
Mastery means
- Versions prompts as code
- Tests prompts in CI against a set of cases
- Measures regressions between prompt versions
Sign in to do the exercises and build your mastery up.
Sources
- OpenAI Evals (MIT) — MIT
- Anthropic — Prompt engineering — documentation, free to read