Skip to content
AI-grafen
FAI engineeringMultimodal models· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Document understanding: OCR, layout, tables

Be able to extract structured information from document images and measure the quality.

Prerequisites

Intuition

A document is not a string of text. The layout carries meaning: what is in which column, which row of a table, which field a number belongs to. Read the text from top to bottom and you lose precisely what makes the document intelligible.

Two routes, and they really are different:

An OCR pipelineA VLM directly
Stepsimage → OCR → layout analysis → field matchingimage → model → JSON
Givestext coordinates, traceability to the pixela directly structured answer
Good athigh volumes, fixed templates, auditingvarying layouts, handwriting, few documents
Errors arisein the OCR and in the rule setsilently, as invented field values
Costlow per pagehigh per page

The decisive difference is where the errors show up. An OCR pipeline that fails gives you an empty result or obvious rubbish. A VLM that fails gives you a credible but wrong company registration number.

Formal

Measure field by field, never as a whole. A document extraction system with «94 % accuracy» is a meaningless figure. The fields differ enormously:

FieldA typical exact matchWhy
Invoice number0.98well delimited, the characters are easy to read
Date0.95the format varies but can be normalised
Total amount0.93confused with subtotals
Supplier name0.85logos, abbreviations, the legal name ≠ the brand
Line items (a table)0.70multi-page tables, merged cells

Normalise before you compare: «1 234,50 kr», «1234.50» and «1 234.50» are the same amount. Otherwise you are measuring formatting.

Confidence and human review. The value lies in sending the right documents to a human. That requires a calibrated confidence:

  1. Let the model give a confidence per field — or use the OCR confidence and the agreement between two methods.
  2. Calibrate against a labelled set: when the model says 0.9, how often is it right?
  3. Set the threshold from what an error costs compared with what the review costs.
  4. Measure the coverage — the share of documents that go through automatically. That is the number that decides the business value.

Cross-validation is cheap and effective: run both the OCR pipeline and a VLM, and flag every field where they disagree. Two independent errors coinciding is rare, and what you get is a confidence signal that does not come from the model itself.

Two things to handle from the start: multi-page documents (line items continuing across a page break) and personal data (document images almost always contain some — the same retention rules apply as for any other personal data).

Code

import json, re
from pydantic import BaseModel, Field

class LineItem(BaseModel):
    description: str
    quantity: float
    unit_price: float
    amount: float

class Invoice(BaseModel):
    invoice_number: str | None = None
    invoice_date: str | None = Field(default=None, description="YYYY-MM-DD")
    supplier: str | None = None
    company_registration_number: str | None = None
    total_amount: float | None = None
    currency: str = "SEK"
    line_items: list[LineItem] = []
    uncertain_fields: list[str] = []

PROMPT = ("Extract the invoice as JSON following the schema. Read ONLY what is in the image. "
          "Set null for fields you cannot read and list them in uncertain_fields. Never invent values.")

def amount(s):
    if s is None:
        return None
    return float(re.sub(r"[^\d,.-]", "", str(s)).replace(" ", "").replace(",", "."))

def field_wise_metrics(truth: list[dict], pred: list[dict], fields: list[str]):
    out = {}
    for f in fields:
        hits = sum(1 for a, b in zip(truth, pred)
                   if (amount(a[f]) if f.endswith("amount") else a[f]) ==
                      (amount(b.get(f)) if f.endswith("amount") else b.get(f)))
        out[f] = round(hits / len(truth), 3)
    return out

# Cross-validation: two independent routes, flag the disagreements
def cross(ocr_result: dict, vlm_result: dict, fields: list[str]):
    return [f for f in fields if ocr_result.get(f) != vlm_result.get(f)]

print(field_wise_metrics(truth, pred, ["invoice_number", "invoice_date", "total_amount", "supplier"]))
# {'invoice_number': 0.98, 'invoice_date': 0.95, 'total_amount': 0.93, 'supplier': 0.85}

uncertain_fields in the schema is cheap and underrated: it gives the model a way of saying «I don't know» that is structured enough to act on in the code.

Mastery means

  • Chooses between an OCR pipeline and a VLM
  • Extracts fields and tables with a schema
  • Measures field-wise accuracy and calibrates the confidence

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences