Text preprocessing
Be able to normalise, tokenise and clean text before modelling.
Prerequisites
- CPython — strings and text processingrequired
Intuition
Before text becomes features it has to be cleaned. The common steps, in order:
- Unicode normalisation (NFC) — «é» can be stored in two ways; make them the same.
- Lower case — «Dog» and «dog» become the same token.
- Remove or keep the punctuation — it depends on the task.
- Tokenise — split into words (or sub-words).
- Stop words — remove «and», «that», «it» if they are only noise.
- Stemming/lemmatisation — «running», «ran», «runs» → the same base form.
But every step throws information away. Lower-casing destroys «SOS» vs «sos». Removing ! destroys the tone. Stop words destroy «to be or not to be».
Code
import re, unicodedata
STOP_WORDS = {"och", "att", "det", "en", "är", "som", "på", "med", "för", "av", "den", "till"}
def normalise(text, lower=True, remove_stop_words=False):
t = unicodedata.normalize("NFC", text)
if lower:
t = t.lower()
tokens = re.findall(r"\w+(?:-\w+)?", t) # keeps the hyphen in «e-post»
if remove_stop_words:
tokens = [x for x in tokens if x not in STOP_WORDS]
return tokens
s = "Det är AI-grafen, och den är bra!"
print(normalise(s)) # ['det','är','ai-grafen','och','den','är','bra']
print(normalise(s, remove_stop_words=True)) # ['ai-grafen','bra']
The rule: every step has to be justified by the task, and the same preprocessing has to be used in training and in production. A model trained on lower case that gets capitals in production performs worse — and nobody notices until the users complain.
For modern language models most of this is handled by the tokenizer (BPE), and then you do less preprocessing, not more. Clean away only what is genuinely noise: HTML tags, duplicates, control characters.
Mastery means
- Normalises and tokenises text
- Justifies every preprocessing step
- Sees when normalisation destroys information
Sign in to do the exercises and build your mastery up.