Document understanding: OCR, layout, tables
Be able to extract structured information from document images and measure the quality.
Prerequisites
Intuition
A document is not a string of text. The layout carries meaning: what is in which column, which row of a table, which field a number belongs to. Read the text from top to bottom and you lose precisely what makes the document intelligible.
Two routes, and they really are different:
| An OCR pipeline | A VLM directly | |
|---|---|---|
| Steps | image → OCR → layout analysis → field matching | image → model → JSON |
| Gives | text coordinates, traceability to the pixel | a directly structured answer |
| Good at | high volumes, fixed templates, auditing | varying layouts, handwriting, few documents |
| Errors arise | in the OCR and in the rule set | silently, as invented field values |
| Cost | low per page | high per page |
The decisive difference is where the errors show up. An OCR pipeline that fails gives you an empty result or obvious rubbish. A VLM that fails gives you a credible but wrong company registration number.
Formal
Measure field by field, never as a whole. A document extraction system with «94 % accuracy» is a meaningless figure. The fields differ enormously:
| Field | A typical exact match | Why |
|---|---|---|
| Invoice number | 0.98 | well delimited, the characters are easy to read |
| Date | 0.95 | the format varies but can be normalised |
| Total amount | 0.93 | confused with subtotals |
| Supplier name | 0.85 | logos, abbreviations, the legal name ≠ the brand |
| Line items (a table) | 0.70 | multi-page tables, merged cells |
Normalise before you compare: «1 234,50 kr», «1234.50» and «1 234.50» are the same amount. Otherwise you are measuring formatting.
Confidence and human review. The value lies in sending the right documents to a human. That requires a calibrated confidence:
- Let the model give a confidence per field — or use the OCR confidence and the agreement between two methods.
- Calibrate against a labelled set: when the model says 0.9, how often is it right?
- Set the threshold from what an error costs compared with what the review costs.
- Measure the coverage — the share of documents that go through automatically. That is the number that decides the business value.
Cross-validation is cheap and effective: run both the OCR pipeline and a VLM, and flag every field where they disagree. Two independent errors coinciding is rare, and what you get is a confidence signal that does not come from the model itself.
Two things to handle from the start: multi-page documents (line items continuing across a page break) and personal data (document images almost always contain some — the same retention rules apply as for any other personal data).
Code
import json, re
from pydantic import BaseModel, Field
class LineItem(BaseModel):
description: str
quantity: float
unit_price: float
amount: float
class Invoice(BaseModel):
invoice_number: str | None = None
invoice_date: str | None = Field(default=None, description="YYYY-MM-DD")
supplier: str | None = None
company_registration_number: str | None = None
total_amount: float | None = None
currency: str = "SEK"
line_items: list[LineItem] = []
uncertain_fields: list[str] = []
PROMPT = ("Extract the invoice as JSON following the schema. Read ONLY what is in the image. "
"Set null for fields you cannot read and list them in uncertain_fields. Never invent values.")
def amount(s):
if s is None:
return None
return float(re.sub(r"[^\d,.-]", "", str(s)).replace(" ", "").replace(",", "."))
def field_wise_metrics(truth: list[dict], pred: list[dict], fields: list[str]):
out = {}
for f in fields:
hits = sum(1 for a, b in zip(truth, pred)
if (amount(a[f]) if f.endswith("amount") else a[f]) ==
(amount(b.get(f)) if f.endswith("amount") else b.get(f)))
out[f] = round(hits / len(truth), 3)
return out
# Cross-validation: two independent routes, flag the disagreements
def cross(ocr_result: dict, vlm_result: dict, fields: list[str]):
return [f for f in fields if ocr_result.get(f) != vlm_result.get(f)]
print(field_wise_metrics(truth, pred, ["invoice_number", "invoice_date", "total_amount", "supplier"]))
# {'invoice_number': 0.98, 'invoice_date': 0.95, 'total_amount': 0.93, 'supplier': 0.85}
uncertain_fields in the schema is cheap and underrated: it gives the model a way of saying «I don't know» that is structured enough to act on in the code.
Mastery means
- Chooses between an OCR pipeline and a VLM
- Extracts fields and tables with a schema
- Measures field-wise accuracy and calibrates the confidence
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — LayoutLMv3: Pre-training for Document AI — arXiv (open access; licence per article)
- arXiv — OCR-free Document Understanding Transformer (Donut) — arXiv (open access; licence per article)
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0