Skip to content
AI-grafen
EUniversityData handling· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Data versioning

Be able to version datasets with checksums and tie a model to exactly the data it was trained on.

Prerequisites

Intuition

Git is built for text and handles large binary files badly. A 10 GB dataset in a Git repo makes the repo unmanageable for everyone, for ever — the history cannot be removed afterwards.

The solution: version a pointer in Git, and store the data somewhere else.

repo/
  data/train.csv.dvc     ← 200 bytes in Git: the hash, the size, the path
  train.py
.dvc/cache/ or S3/       ← the data itself, addressed by its hash

Content addressing is the core of it: the file's name in the store is its SHA-256. Two identical files are stored once; a changed file gets a new name. The same idea Git uses internally for its objects.

The question that has to be answerable: «the model from 14 March — exactly which data was it trained on?» Without versioning the answer is a guess.

Formal

Tools and when they fit:

ToolThe ideaSuits
Git LFSpointers in Git, files on an LFS servermoderately large files, simple
DVCpointers in Git, data in any storeML projects, pipelines
LakeFSGit-like branches over object storagelarge data lakes, teams
Delta / Iceberga transaction log over Parquettables with time travel
Your own manifest filea hash per file, checked into Gitsmall projects — works surprisingly well

The last row is underrated: a manifest.json with the hash, the size and the row count per file gives 80 % of the benefit for zero infrastructure.

What should be hashed. Hash the contents, not the file name or the modification time. For a table it is often enough to hash the sorted, canonicalised serialisation — then the same data gives the same hash even if the row order differs between runs.

The model ↔ data link is the whole point. Every training run should log:

FieldWhy
data_hashexactly which data
split_seedexactly which split
git_commitexactly which code
config_hashexactly which hyperparameters
environmentthe library versions or the container tag

With those five fields a run is reproducible. Without any one of them it is not.

The legal dimension. The GDPR gives a right to erasure. A dataset with personal data versioned «for ever» collides with that. The solution is to version pseudonymised data, with the link to identity held in a separate, prunable register — and to have a documented routine for removing a person from every version.

That is not a theoretical objection: anyone who builds data versioning without thinking erasure through is building in a problem that is expensive to solve afterwards.

Code

import hashlib, json
from pathlib import Path
import pandas as pd

def file_hash(p: Path, block=1 << 20) -> str:
    h = hashlib.sha256()
    with p.open("rb") as f:
        for chunk in iter(lambda: f.read(block), b""):
            h.update(chunk)
    return h.hexdigest()

def table_hash(df: pd.DataFrame) -> str:
    """A canonical hash: independent of the row order and the column order."""
    d = df.sort_index(axis=1)
    d = d.sort_values(list(d.columns)).reset_index(drop=True)
    return hashlib.sha256(d.to_csv(index=False).encode()).hexdigest()

def build_manifest(directory: Path, out: Path):
    entries = []
    for f in sorted(directory.rglob("*")):
        if f.is_file():
            entries.append({"path": str(f.relative_to(directory)),
                            "bytes": f.stat().st_size,
                            "sha256": file_hash(f)})
    manifest = {"files": entries,
                "total_bytes": sum(e["bytes"] for e in entries),
                "dataset_hash": hashlib.sha256(
                    "".join(e["sha256"] for e in entries).encode()).hexdigest()}
    out.write_text(json.dumps(manifest, indent=2), encoding="utf-8")
    return manifest["dataset_hash"]

def verify(directory: Path, manifest_file: Path):
    m = json.loads(manifest_file.read_text(encoding="utf-8"))
    problems = []
    for entry in m["files"]:
        f = directory / entry["path"]
        if not f.exists():
            problems.append(f"missing: {entry['path']}")
        elif file_hash(f) != entry["sha256"]:
            problems.append(f"changed: {entry['path']}")
    return problems or ["everything checks out"]

# Tie the model to the data — the five fields that make a run reproducible
def run_metadata(dataset_hash, config, split_seed, git_commit, environment):
    return {
        "data_hash": dataset_hash,
        "config_hash": hashlib.sha256(
            json.dumps(config, sort_keys=True).encode()).hexdigest()[:16],
        "split_seed": split_seed,
        "git_commit": git_commit,
        "environment": environment,
    }

table_hash sorts before it hashes. Without that the same data gets a different hash simply because a GROUP BY returned the rows in a different order — and then the versioning is worthless, since everything always looks changed.

Mastery means

  • Versions datasets with a content hash
  • Ties the model to the dataset
  • Chooses a versioning strategy according to the data size

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences