Feature engineering
Be able to create, encode and scale features, including one-hot and normalisation.
Prerequisites
Intuition
A model sees only numbers. Feature engineering is turning reality into numbers in a way that makes the patterns visible.
Categorical variables:
| Method | How | Suits |
|---|---|---|
| One-hot | one 0/1 column per value | few categories, linear models |
| Ordinal | 0, 1, 2 … | when the order is real (low/medium/high) |
| Target encoding | replace it with the mean of the target in that category | many categories — but it leaks easily |
| Embedding | a learnt vector | very many categories, neural networks |
The trap with ordinal encoding: encode «Malmö = 1, Lund = 2, Umeå = 3» and you are claiming that Lund lies between Malmö and Umeå and that Umeå is three times Malmö. That is nonsense, and a linear model will believe it.
Numeric variables are scaled so that one variable does not dominate simply because it is measured in a larger unit:
| Method | Formula | The result |
|---|---|---|
| Standardisation | mean 0, std 1 | |
| Min–max | the range [0, 1] | |
| Robust | copes with outliers | |
| Log | compresses a right-hand tail |
Formal
Leakage is the big mistake in preprocessing, and it comes in three forms:
- Statistics from the test data. Compute and over the whole dataset → the test set's distribution has influenced the training. The right way:
fiton the training data,transformon both. - Target encoding without cross-validation. If you replace a category with the mean of the target, you are using the target as a feature. Without out-of-fold computation the model memorises its own labels.
- Features from the future. «The number of purchases in the last month» computed over the whole period contains information that was not available at the moment of the decision. This is the hardest to spot and the most common in time series problems.
The symptom of leakage is always the same: a suspiciously good validation result that does not hold in production. A model that suddenly gets 0.99 AUC has usually not become good — it has been shown the answers.
Pipeline solves problem 1 structurally by making the whole preprocessing part of the model. That is not just tidier code — it is the only form that survives cross-validation correctly, since fit is then rerun within every fold.
Which models need scaling?
| Need it | Do not need it |
|---|---|
| kNN, SVM, k-means (distance-based) | decision trees |
| linear/logistic regression with regularisation | random forests |
| neural networks | gradient boosting |
Trees split on thresholds and do not care about the scale. Anything that measures distance or penalises the size of the coefficients does.
Domain knowledge beats automation. The features that usually make the largest difference are the ones somebody who knows the field suggests: ratios (the price per square metre), time differences (days since the last visit), and cyclic re-encodings (the weekday as sin/cos, so that Sunday lies close to Monday).
Code
import numpy as np, pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
num = ["age", "hours"]
cat = ["city", "programme"]
preprocessing = ColumnTransformer([
("num", Pipeline([("impute", SimpleImputer(strategy="median")),
("scale", StandardScaler())]), num),
("cat", Pipeline([("impute", SimpleImputer(strategy="most_frequent")),
("onehot", OneHotEncoder(handle_unknown="ignore"))]), cat),
])
model = Pipeline([("preprocessing", preprocessing),
("classifier", LogisticRegression(max_iter=1000))])
# The cross-validation reruns the WHOLE preprocessing within every fold → no leakage
print(cross_val_score(model, X, y, cv=5, scoring="roc_auc").mean())
# handle_unknown="ignore" is needed: production data contains categories
# that were not in the training set, and without it transform crashes.
# Domain knowledge: cyclic variables
df = pd.DataFrame({"weekday": [0, 1, 5, 6]}) # 0 = Monday, 6 = Sunday
df["day_sin"] = np.sin(2 * np.pi * df.weekday / 7)
df["day_cos"] = np.cos(2 * np.pi * df.weekday / 7)
# The distance between Sunday (6) and Monday (0) is now small, as it should be:
for a, b in [(6, 0), (0, 3)]:
va = np.array([np.sin(2*np.pi*a/7), np.cos(2*np.pi*a/7)])
vb = np.array([np.sin(2*np.pi*b/7), np.cos(2*np.pi*b/7)])
print(f"day {a} → {b}: distance {np.linalg.norm(va - vb):.3f}")
# day 6 → 0: distance 0.868
# day 0 → 3: distance 1.950 ← Sunday lies close to Monday, as it does in reality
The cyclic encoding is a good example of a simple idea from domain knowledge often giving more than a larger model.
Mastery means
- Encodes categorical variables
- Scales numeric variables correctly
- Avoids leakage in the preprocessing
Sign in to do the exercises and build your mastery up.
Sources
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- pandas — User Guide (BSD-3) — BSD-3-Clause
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0