scikit-learn — the workflow
Be able to build a pipeline, train, evaluate and save a model with scikit-learn.
Prerequisites
Intuition
scikit-learn has a consistent pattern that every model follows:
| Method | Does |
|---|---|
fit(X, y) | learns from the training data |
predict(X) | predicts |
predict_proba(X) | probabilities (classification) |
transform(X) | transforms (preprocessors) |
fit_transform(X) | both in one |
score(X, y) | the default metric |
Because everything follows the pattern you can swap any part out without touching the rest. Swap LogisticRegression for RandomForestClassifier — the rest of the code is unchanged.
The workflow, in order:
split the data → build a pipeline → cross-validate → search the hyperparameters
→ train on all the training data → evaluate ONCE on the test set → save
The decisive step is that the test set is touched a single time, right at the end. Every time you look at the test result and change something you have started fitting yourself to it.
Code
import joblib, numpy as np, pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split, cross_val_score, GridSearchCV
from sklearn.metrics import classification_report, confusion_matrix
# 1. Split FIRST — everything else happens on the training part
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, stratify=y, random_state=0)
# 2. A pipeline: preprocessing plus the model as one unit
num, cat = ["age", "hours"], ["programme"]
pipe = Pipeline([
("prep", ColumnTransformer([
("n", Pipeline([("i", SimpleImputer(strategy="median")), ("s", StandardScaler())]), num),
("c", OneHotEncoder(handle_unknown="ignore"), cat),
])),
("clf", RandomForestClassifier(random_state=0)),
])
# 3. Cross-validate — the preprocessing is redone within each fold
cv = cross_val_score(pipe, Xtr, ytr, cv=5, scoring="roc_auc")
print(f"CV ROC-AUC {cv.mean():.3f} ± {cv.std():.3f}")
# 4. Search the hyperparameters — still only on the training data
grid = {"clf__n_estimators": [100, 300], "clf__max_depth": [None, 8, 16]}
search = GridSearchCV(pipe, grid, cv=5, scoring="roc_auc", n_jobs=-1).fit(Xtr, ytr)
print(search.best_params_, round(search.best_score_, 3))
# 5. ONCE on the test set
print(classification_report(yte, search.predict(Xte)))
print(confusion_matrix(yte, search.predict(Xte)))
# 6. Save the whole pipeline, not just the model
joblib.dump(search.best_estimator_, "model.joblib")
loaded = joblib.load("model.joblib")
print(loaded.predict(Xte[:3])) # raw input in — the preprocessing comes along
Saving the whole pipeline is the point. If you save only the model, whoever uses it has to recreate exactly the same preprocessing — the same medians, the same category order, the same scaling. That nearly always goes wrong sooner or later.
joblib has a limitation worth knowing about: the file contains embedded Python and is tied to the library versions. Never load a .joblib from an unknown source, and pin the scikit-learn version in requirements.txt. For long-term storage ONNX is a more durable format.
Interactive
Build your first model in twenty minutes. Use a built-in dataset and you avoid the data collection.
from sklearn.datasets import fetch_openml
X, y = fetch_openml("credit-g", version=1, return_X_y=True, as_frame=True)
Then do the following, in order, and write the result down after each step:
- A baseline. What does
DummyClassifier(strategy="most_frequent")get? Everything you build has to beat that. - The simplest real model.
LogisticRegressionin a pipeline. The difference against the baseline? - A tree.
RandomForestClassifierwith the default settings. Better? - A search.
GridSearchCVover two parameters. How much did the search give? - Look at the errors.
confusion_matrix— which kind of error does the model make most? - Once on the test set.
What you will probably discover:
- Step 1 gives a surprisingly high «accuracy» if the classes are imbalanced — which is why you should not measure accuracy.
- Steps 2 to 3 often give less of a difference than you would think.
- Step 4 rarely gives more than a percentage point or two.
- Step 5 is the one that actually leads somewhere — the errors are nearly always concentrated in an identifiable group.
That order — a baseline, a simple model, error analysis, then the fine-tuning — is the whole difference between working effectively and spending a week turning hyperparameter knobs.
Mastery means
- Builds a pipeline with preprocessing and a model
- Evaluates with cross-validation
- Saves and loads a model correctly
Sign in to do the exercises and build your mastery up.
Sources
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- The Python documentation (PSF licence) — PSF