Data versioning
Be able to version datasets with checksums and tie a model to exactly the data it was trained on.
Prerequisites
- DGit — version controlrequired
- EData pipelines and ETLrequired
Intuition
Git is built for text and handles large binary files badly. A 10 GB dataset in a Git repo makes the repo unmanageable for everyone, for ever — the history cannot be removed afterwards.
The solution: version a pointer in Git, and store the data somewhere else.
repo/
data/train.csv.dvc ← 200 bytes in Git: the hash, the size, the path
train.py
.dvc/cache/ or S3/ ← the data itself, addressed by its hash
Content addressing is the core of it: the file's name in the store is its SHA-256. Two identical files are stored once; a changed file gets a new name. The same idea Git uses internally for its objects.
The question that has to be answerable: «the model from 14 March — exactly which data was it trained on?» Without versioning the answer is a guess.
Formal
Tools and when they fit:
| Tool | The idea | Suits |
|---|---|---|
| Git LFS | pointers in Git, files on an LFS server | moderately large files, simple |
| DVC | pointers in Git, data in any store | ML projects, pipelines |
| LakeFS | Git-like branches over object storage | large data lakes, teams |
| Delta / Iceberg | a transaction log over Parquet | tables with time travel |
| Your own manifest file | a hash per file, checked into Git | small projects — works surprisingly well |
The last row is underrated: a manifest.json with the hash, the size and the row count per file gives 80 % of the benefit for zero infrastructure.
What should be hashed. Hash the contents, not the file name or the modification time. For a table it is often enough to hash the sorted, canonicalised serialisation — then the same data gives the same hash even if the row order differs between runs.
The model ↔ data link is the whole point. Every training run should log:
| Field | Why |
|---|---|
data_hash | exactly which data |
split_seed | exactly which split |
git_commit | exactly which code |
config_hash | exactly which hyperparameters |
environment | the library versions or the container tag |
With those five fields a run is reproducible. Without any one of them it is not.
The legal dimension. The GDPR gives a right to erasure. A dataset with personal data versioned «for ever» collides with that. The solution is to version pseudonymised data, with the link to identity held in a separate, prunable register — and to have a documented routine for removing a person from every version.
That is not a theoretical objection: anyone who builds data versioning without thinking erasure through is building in a problem that is expensive to solve afterwards.
Code
import hashlib, json
from pathlib import Path
import pandas as pd
def file_hash(p: Path, block=1 << 20) -> str:
h = hashlib.sha256()
with p.open("rb") as f:
for chunk in iter(lambda: f.read(block), b""):
h.update(chunk)
return h.hexdigest()
def table_hash(df: pd.DataFrame) -> str:
"""A canonical hash: independent of the row order and the column order."""
d = df.sort_index(axis=1)
d = d.sort_values(list(d.columns)).reset_index(drop=True)
return hashlib.sha256(d.to_csv(index=False).encode()).hexdigest()
def build_manifest(directory: Path, out: Path):
entries = []
for f in sorted(directory.rglob("*")):
if f.is_file():
entries.append({"path": str(f.relative_to(directory)),
"bytes": f.stat().st_size,
"sha256": file_hash(f)})
manifest = {"files": entries,
"total_bytes": sum(e["bytes"] for e in entries),
"dataset_hash": hashlib.sha256(
"".join(e["sha256"] for e in entries).encode()).hexdigest()}
out.write_text(json.dumps(manifest, indent=2), encoding="utf-8")
return manifest["dataset_hash"]
def verify(directory: Path, manifest_file: Path):
m = json.loads(manifest_file.read_text(encoding="utf-8"))
problems = []
for entry in m["files"]:
f = directory / entry["path"]
if not f.exists():
problems.append(f"missing: {entry['path']}")
elif file_hash(f) != entry["sha256"]:
problems.append(f"changed: {entry['path']}")
return problems or ["everything checks out"]
# Tie the model to the data — the five fields that make a run reproducible
def run_metadata(dataset_hash, config, split_seed, git_commit, environment):
return {
"data_hash": dataset_hash,
"config_hash": hashlib.sha256(
json.dumps(config, sort_keys=True).encode()).hexdigest()[:16],
"split_seed": split_seed,
"git_commit": git_commit,
"environment": environment,
}
table_hash sorts before it hashes. Without that the same data gets a different hash simply because a GROUP BY returned the rows in a different order — and then the versioning is worthless, since everything always looks changed.
Mastery means
- Versions datasets with a content hash
- Ties the model to the dataset
- Chooses a versioning strategy according to the data size
Sign in to do the exercises and build your mastery up.
Sources
- DVC — dokumentation (Apache-2.0) — Apache-2.0
- GDPR — Regulation (EU) 2016/679, Art. 17 — EU legal act
- Pro Git (Chacon & Straub) — CC BY-NC-SA 3.0