Index updating and versioning
Be able to update an index incrementally without losing reproducibility.
Prerequisites
- EData versioningrequired
- EVector databases and indexingrequired
Intuition
An index that was built once and is never updated quickly becomes out of date. But rebuilding everything every time a document changes is unsustainable with millions of documents.
Four operations, of differing difficulty:
| The operation | The difficulty | Why |
|---|---|---|
| Add | easy | most indexes support it directly |
| Update | medium | remove + add, and keep the id stable |
| Delete | hard | the HNSW graph cannot easily heal after a deletion |
| Change the embedding model | requires a full rebuild | old and new vectors are not comparable |
The last row is the most important one to plan for. Two vectors from different models lie in different spaces — the distances between them mean nothing. An index can never contain a mixture.
Formal
Soft deletion is the standard solution: mark the document as deleted in the metadata and filter it out at search time. The vector stays in the graph and takes up space, but the search gives the right answer.
When the share of deleted entries grows beyond a few tens of per cent, both the memory and the recall deteriorate, and then the index is rebuilt — like a vacuum in a database.
Stable ids are a prerequisite. Use a deterministic id from the document's identity, not a running number:
id = sha256(f"{source}:{document_id}:{chunk_index}")
Then the same passage can always be found and replaced, whatever order the indexing was run in.
A chunk hash decides what needs redoing. Save the hash of the content per passage; at reindexing the new hash is compared with the old:
| The outcome | The measure |
|---|---|
| The hash is unchanged | skip it — no new embedding is needed |
| The hash has changed | re-embed and replace |
| The passage is now missing | mark it as deleted |
| A new passage | add it |
For a corpus where 2 % changes a week that means 98 % of the embedding cost disappears.
A blue-green rebuild when changing model:
- Build a new index (green) with the new model, while the old one (blue) carries on answering.
- Run the evaluation against both on the same questions.
- If green is better — switch the alias. If it is worse — throw it away.
- Keep blue for a while, so that going back is one alias switch away.
Version the index like the data. Every index should carry:
| The field | Why |
|---|---|
| The embedding model and version | it decides whether the vectors are comparable |
| The chunking parameters | they affect what is found |
| The hash of the dataset | which material |
| The build time and the code version | traceability |
Without those fields it is impossible to answer why a query gave a different answer last week — and that is a question that always comes.
Code
import hashlib, json, time
from dataclasses import dataclass, asdict
from pathlib import Path
def chunk_id(source: str, doc: str, i: int) -> str:
return hashlib.sha256(f"{source}:{doc}:{i}".encode()).hexdigest()[:32]
def content_hash(text: str) -> str:
return hashlib.sha256(text.encode()).hexdigest()[:16]
@dataclass
class IndexVersion:
embedding_model: str
embedding_dim: int
chunk_size: int
chunk_overlap: int
dataset_hash: str
built: str
git_commit: str
def incremental_update(client, collection, new_chunks, existing: dict[str, str]):
"""existing: chunk_id -> content hash. Returns what was done."""
to_add, replaced, unchanged = [], [], 0
seen = set()
for c in new_chunks:
cid = chunk_id(c["source"], c["doc"], c["i"])
h = content_hash(c["text"])
seen.add(cid)
if existing.get(cid) == h:
unchanged += 1 # skip it — it saves embedding cost
elif cid in existing:
replaced.append((cid, c, h))
else:
to_add.append((cid, c, h))
deleted = [cid for cid in existing if cid not in seen]
for cid, c, h in to_add + replaced:
client.upsert(collection, id=cid, vector=embed(c["text"]),
payload={"text": c["text"], "hash": h, "deleted": False})
for cid in deleted:
client.set_payload(collection, id=cid, payload={"deleted": True}) # soft deletion
return {"new": len(to_add), "replaced": len(replaced),
"unchanged": unchanged, "deleted": len(deleted),
"embedding_share_saved": round(
unchanged / max(len(new_chunks), 1), 3)}
# The search has to filter the soft-deleted entries out
def search(client, collection, q, k=10):
return client.search(collection, q, limit=k,
query_filter={"must": [{"key": "deleted", "match": {"value": False}}]})
# Blue-green: build the new one alongside, switch the alias only after the evaluation
def blue_green_rollout(client, old, new, evaluate, threshold=0.0):
before, after = evaluate(old), evaluate(new)
print(f"recall@10: {before:.3f} → {after:.3f}")
if after >= before + threshold:
client.update_alias("production", new)
return {"switched": True, "before": before, "after": after}
return {"switched": False, "reason": "no improvement", "before": before, "after": after}
embedding_share_saved is the number that justifies the whole construction: if it sits at 0.98 it means a weekly reindexing costs two per cent of what a full rebuild would have.
Mastery means
- Updates an index incrementally
- Handles deletion and rebuilding
- Keeps reproducibility when changing model
Sign in to do the exercises and build your mastery up.
Sources
- Qdrant — dokumentation (Apache-2.0) — Apache-2.0
- FAISS — dokumentation (MIT) — MIT
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0