Skip to content
AI-grafen
EUniversityAI product development· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Prompts as code: versioning and testing

Be able to version prompts, test them in CI and measure regressions.

Prerequisites

Intuition

A prompt in production is code. It determines the system's behaviour, it breaks, and it therefore has to be treated accordingly:

CodeA prompt
in gitin git — not in a database or a Google document
versionedversioned, with an id in the logs
tested in CIan eval suite in CI with a threshold
code-reviewedreviewed by somebody else
rolled backcan be rolled back in minutes

The most common source of error in LLM products is that somebody «just changed a phrase» in a prompt, everything got worse, and nobody could say when or why.

Code

prompts/
├── tutor_explain/
│   ├── v3.md          # the current one
│   ├── v2.md
│   └── cases.jsonl    # the eval cases belonging to this particular prompt
└── registry.yaml      # which version is active per environment
from pathlib import Path
import hashlib, yaml

REG = yaml.safe_load(Path("prompts/registry.yaml").read_text())

def get_prompt(name: str, environment: str = "prod") -> tuple[str, str]:
    version = REG[name][environment]                 # "v3", for instance
    text = Path(f"prompts/{name}/{version}.md").read_text(encoding="utf-8")
    return text, f"{name}:{version}:{hashlib.sha256(text.encode()).hexdigest()[:8]}"

prompt, prompt_id = get_prompt("tutor_explain")
# log the prompt_id with every call → which behaviour an answer came out of can be traced
# .github/workflows/ci.yml (or scripts/ci.sh)
- run: python -m evals.run --prompts prompts/ --threshold 0.85 --fail-on-regression

Three rules that go a long way: the prompt id in every log line, the eval suite run on every change to a prompt file, and every new prompt version compared against the active one — not just against the threshold. A prompt that clears the threshold but loses three previously passing cases is a regression.

Mastery means

  • Versions prompts as code
  • Tests prompts in CI against a set of cases
  • Measures regressions between prompt versions

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences