Skip to content
AI-grafen
DAI developerData handling· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Feature engineering

Be able to create, encode and scale features, including one-hot and normalisation.

Prerequisites

Intuition

A model sees only numbers. Feature engineering is turning reality into numbers in a way that makes the patterns visible.

Categorical variables:

MethodHowSuits
One-hotone 0/1 column per valuefew categories, linear models
Ordinal0, 1, 2 …when the order is real (low/medium/high)
Target encodingreplace it with the mean of the target in that categorymany categories — but it leaks easily
Embeddinga learnt vectorvery many categories, neural networks

The trap with ordinal encoding: encode «Malmö = 1, Lund = 2, Umeå = 3» and you are claiming that Lund lies between Malmö and Umeå and that Umeå is three times Malmö. That is nonsense, and a linear model will believe it.

Numeric variables are scaled so that one variable does not dominate simply because it is measured in a larger unit:

MethodFormulaThe result
Standardisation(x−μ)/σ(x - \mu)/\sigmamean 0, std 1
Min–max(x−min⁡)/(max⁡−min⁡)(x - \min)/(\max - \min)the range [0, 1]
Robust(x−median)/IQR(x - \text{median})/\text{IQR}copes with outliers
Loglog⁡(1+x)\log(1 + x)compresses a right-hand tail

Formal

Leakage is the big mistake in preprocessing, and it comes in three forms:

  1. Statistics from the test data. Compute μ\mu and σ\sigma over the whole dataset → the test set's distribution has influenced the training. The right way: fit on the training data, transform on both.
  2. Target encoding without cross-validation. If you replace a category with the mean of the target, you are using the target as a feature. Without out-of-fold computation the model memorises its own labels.
  3. Features from the future. «The number of purchases in the last month» computed over the whole period contains information that was not available at the moment of the decision. This is the hardest to spot and the most common in time series problems.

The symptom of leakage is always the same: a suspiciously good validation result that does not hold in production. A model that suddenly gets 0.99 AUC has usually not become good — it has been shown the answers.

Pipeline solves problem 1 structurally by making the whole preprocessing part of the model. That is not just tidier code — it is the only form that survives cross-validation correctly, since fit is then rerun within every fold.

Which models need scaling?

Need itDo not need it
kNN, SVM, k-means (distance-based)decision trees
linear/logistic regression with regularisationrandom forests
neural networksgradient boosting

Trees split on thresholds and do not care about the scale. Anything that measures distance or penalises the size of the coefficients does.

Domain knowledge beats automation. The features that usually make the largest difference are the ones somebody who knows the field suggests: ratios (the price per square metre), time differences (days since the last visit), and cyclic re-encodings (the weekday as sin/cos, so that Sunday lies close to Monday).

Code

import numpy as np, pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score

num = ["age", "hours"]
cat = ["city", "programme"]

preprocessing = ColumnTransformer([
    ("num", Pipeline([("impute", SimpleImputer(strategy="median")),
                      ("scale", StandardScaler())]), num),
    ("cat", Pipeline([("impute", SimpleImputer(strategy="most_frequent")),
                      ("onehot", OneHotEncoder(handle_unknown="ignore"))]), cat),
])

model = Pipeline([("preprocessing", preprocessing),
                  ("classifier", LogisticRegression(max_iter=1000))])

# The cross-validation reruns the WHOLE preprocessing within every fold → no leakage
print(cross_val_score(model, X, y, cv=5, scoring="roc_auc").mean())

# handle_unknown="ignore" is needed: production data contains categories
# that were not in the training set, and without it transform crashes.

# Domain knowledge: cyclic variables
df = pd.DataFrame({"weekday": [0, 1, 5, 6]})        # 0 = Monday, 6 = Sunday
df["day_sin"] = np.sin(2 * np.pi * df.weekday / 7)
df["day_cos"] = np.cos(2 * np.pi * df.weekday / 7)
# The distance between Sunday (6) and Monday (0) is now small, as it should be:
for a, b in [(6, 0), (0, 3)]:
    va = np.array([np.sin(2*np.pi*a/7), np.cos(2*np.pi*a/7)])
    vb = np.array([np.sin(2*np.pi*b/7), np.cos(2*np.pi*b/7)])
    print(f"day {a} → {b}: distance {np.linalg.norm(va - vb):.3f}")
# day 6 → 0: distance 0.868
# day 0 → 3: distance 1.950     ← Sunday lies close to Monday, as it does in reality

The cyclic encoding is a good example of a simple idea from domain knowledge often giving more than a larger model.

Mastery means

  • Encodes categorical variables
  • Scales numeric variables correctly
  • Avoids leakage in the preprocessing

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences