Information extraction and NER
Be able to extract entities and relations from text and measure precision and recall.
Prerequisites
Intuition
NER marks spans up in text: people, organisations, places, dates, amounts. The standard format is BIO:
Anna B-PER
Svensson I-PER
works O
at O
Volvo B-ORG
in O
Göteborg B-LOC
B = the beginning of an entity, I = inside, O = outside.
Two approaches today:
| A fine-tuned encoder | An LLM with a schema | |
|---|---|---|
| The data required | 500+ labelled sentences | 0–20 examples |
| The cost per document | very low | high |
| New entity types | retraining | change the prompt |
| Precision on the trained types | higher | somewhat lower |
For large volumes and fixed types the encoder wins. For few documents or changing types the LLM wins.
Formal
Evaluate at the entity level, not the token level. A span counts as correct only if both the boundaries and the type are right. That is harder than a token-wise measurement and is what actually matters.
An error like «Anna Svensson» → «Svensson» therefore counts as both a miss (FN) and a false one (FP) — not as half right. seqeval does this correctly; an ordinary classification_report at the token level systematically overestimates.
Relation extraction goes a step further: find that «Anna Svensson» works at «Volvo». The common approaches are pairwise classification of entity pairs, or an LLM with a schema that requires both entities and relations in JSON.
Pitfalls specific to Swedish: compounds («Volvochefen» contains an organisation), genitive forms, and Swedish NER corpora being considerably smaller than English ones. KB-lab's models and the SUC corpus are the starting points.
Code
from transformers import pipeline
from seqeval.metrics import classification_report
# A fine-tuned Swedish model
ner = pipeline("ner", model="KBLab/bert-base-swedish-cased-ner", aggregation_strategy="simple")
for e in ner("Anna Svensson arbetar på Volvo i Göteborg sedan 2019."):
print(f"{e['entity_group']:6s} {e['word']:20s} {e['score']:.2f}")
# PER Anna Svensson 0.99
# ORG Volvo 0.98
# LOC Göteborg 0.99
# Evaluation at the ENTITY level
true = [["B-PER", "I-PER", "O", "O", "B-ORG", "O", "B-LOC"]]
pred = [["B-PER", "I-PER", "O", "O", "B-ORG", "O", "O"]]
print(classification_report(true, pred, digits=3))
# The LLM alternative with a schema
from pydantic import BaseModel
class Entity(BaseModel):
text: str
type: str # PER|ORG|LOC|TIME|AMOUNT
start: int
class Extraction(BaseModel):
entities: list[Entity]
Always compare the two on your material before choosing — the difference in precision and cost varies greatly with the domain.
Mastery means
- Extracts entities and relations from text
- Measures precision and recall at the entity level
- Chooses between a fine-tuned model and LLM extraction
Sign in to do the exercises and build your mastery up.
Sources
- Kungliga biblioteket — KB-lab modeller — open models
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0