Skip to content
AI-grafen
EUniversityRAG and information retrieval· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Index updating and versioning

Be able to update an index incrementally without losing reproducibility.

Prerequisites

Intuition

An index that was built once and is never updated quickly becomes out of date. But rebuilding everything every time a document changes is unsustainable with millions of documents.

Four operations, of differing difficulty:

The operationThe difficultyWhy
Addeasymost indexes support it directly
Updatemediumremove + add, and keep the id stable
Deletehardthe HNSW graph cannot easily heal after a deletion
Change the embedding modelrequires a full rebuildold and new vectors are not comparable

The last row is the most important one to plan for. Two vectors from different models lie in different spaces — the distances between them mean nothing. An index can never contain a mixture.

Formal

Soft deletion is the standard solution: mark the document as deleted in the metadata and filter it out at search time. The vector stays in the graph and takes up space, but the search gives the right answer.

When the share of deleted entries grows beyond a few tens of per cent, both the memory and the recall deteriorate, and then the index is rebuilt — like a vacuum in a database.

Stable ids are a prerequisite. Use a deterministic id from the document's identity, not a running number:

id = sha256(f"{source}:{document_id}:{chunk_index}")

Then the same passage can always be found and replaced, whatever order the indexing was run in.

A chunk hash decides what needs redoing. Save the hash of the content per passage; at reindexing the new hash is compared with the old:

The outcomeThe measure
The hash is unchangedskip it — no new embedding is needed
The hash has changedre-embed and replace
The passage is now missingmark it as deleted
A new passageadd it

For a corpus where 2 % changes a week that means 98 % of the embedding cost disappears.

A blue-green rebuild when changing model:

  1. Build a new index (green) with the new model, while the old one (blue) carries on answering.
  2. Run the evaluation against both on the same questions.
  3. If green is better — switch the alias. If it is worse — throw it away.
  4. Keep blue for a while, so that going back is one alias switch away.

Version the index like the data. Every index should carry:

The fieldWhy
The embedding model and versionit decides whether the vectors are comparable
The chunking parametersthey affect what is found
The hash of the datasetwhich material
The build time and the code versiontraceability

Without those fields it is impossible to answer why a query gave a different answer last week — and that is a question that always comes.

Code

import hashlib, json, time
from dataclasses import dataclass, asdict
from pathlib import Path

def chunk_id(source: str, doc: str, i: int) -> str:
    return hashlib.sha256(f"{source}:{doc}:{i}".encode()).hexdigest()[:32]

def content_hash(text: str) -> str:
    return hashlib.sha256(text.encode()).hexdigest()[:16]

@dataclass
class IndexVersion:
    embedding_model: str
    embedding_dim: int
    chunk_size: int
    chunk_overlap: int
    dataset_hash: str
    built: str
    git_commit: str

def incremental_update(client, collection, new_chunks, existing: dict[str, str]):
    """existing: chunk_id -> content hash. Returns what was done."""
    to_add, replaced, unchanged = [], [], 0
    seen = set()
    for c in new_chunks:
        cid = chunk_id(c["source"], c["doc"], c["i"])
        h = content_hash(c["text"])
        seen.add(cid)
        if existing.get(cid) == h:
            unchanged += 1                         # skip it — it saves embedding cost
        elif cid in existing:
            replaced.append((cid, c, h))
        else:
            to_add.append((cid, c, h))

    deleted = [cid for cid in existing if cid not in seen]

    for cid, c, h in to_add + replaced:
        client.upsert(collection, id=cid, vector=embed(c["text"]),
                      payload={"text": c["text"], "hash": h, "deleted": False})
    for cid in deleted:
        client.set_payload(collection, id=cid, payload={"deleted": True})   # soft deletion

    return {"new": len(to_add), "replaced": len(replaced),
            "unchanged": unchanged, "deleted": len(deleted),
            "embedding_share_saved": round(
                unchanged / max(len(new_chunks), 1), 3)}

# The search has to filter the soft-deleted entries out
def search(client, collection, q, k=10):
    return client.search(collection, q, limit=k,
                         query_filter={"must": [{"key": "deleted", "match": {"value": False}}]})

# Blue-green: build the new one alongside, switch the alias only after the evaluation
def blue_green_rollout(client, old, new, evaluate, threshold=0.0):
    before, after = evaluate(old), evaluate(new)
    print(f"recall@10: {before:.3f} → {after:.3f}")
    if after >= before + threshold:
        client.update_alias("production", new)
        return {"switched": True, "before": before, "after": after}
    return {"switched": False, "reason": "no improvement", "before": before, "after": after}

embedding_share_saved is the number that justifies the whole construction: if it sits at 0.98 it means a weekly reindexing costs two per cent of what a full rebuild would have.

Mastery means

  • Updates an index incrementally
  • Handles deletion and rebuilding
  • Keeps reproducibility when changing model

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences