Skip to content
AI-grafen
DAI developerClassical machine learning· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

scikit-learn — the workflow

Be able to build a pipeline, train, evaluate and save a model with scikit-learn.

Prerequisites

Intuition

scikit-learn has a consistent pattern that every model follows:

MethodDoes
fit(X, y)learns from the training data
predict(X)predicts
predict_proba(X)probabilities (classification)
transform(X)transforms (preprocessors)
fit_transform(X)both in one
score(X, y)the default metric

Because everything follows the pattern you can swap any part out without touching the rest. Swap LogisticRegression for RandomForestClassifier — the rest of the code is unchanged.

The workflow, in order:

split the data → build a pipeline → cross-validate → search the hyperparameters
              → train on all the training data → evaluate ONCE on the test set → save

The decisive step is that the test set is touched a single time, right at the end. Every time you look at the test result and change something you have started fitting yourself to it.

Code

import joblib, numpy as np, pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split, cross_val_score, GridSearchCV
from sklearn.metrics import classification_report, confusion_matrix

# 1. Split FIRST — everything else happens on the training part
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, stratify=y, random_state=0)

# 2. A pipeline: preprocessing plus the model as one unit
num, cat = ["age", "hours"], ["programme"]
pipe = Pipeline([
    ("prep", ColumnTransformer([
        ("n", Pipeline([("i", SimpleImputer(strategy="median")), ("s", StandardScaler())]), num),
        ("c", OneHotEncoder(handle_unknown="ignore"), cat),
    ])),
    ("clf", RandomForestClassifier(random_state=0)),
])

# 3. Cross-validate — the preprocessing is redone within each fold
cv = cross_val_score(pipe, Xtr, ytr, cv=5, scoring="roc_auc")
print(f"CV ROC-AUC {cv.mean():.3f} ± {cv.std():.3f}")

# 4. Search the hyperparameters — still only on the training data
grid = {"clf__n_estimators": [100, 300], "clf__max_depth": [None, 8, 16]}
search = GridSearchCV(pipe, grid, cv=5, scoring="roc_auc", n_jobs=-1).fit(Xtr, ytr)
print(search.best_params_, round(search.best_score_, 3))

# 5. ONCE on the test set
print(classification_report(yte, search.predict(Xte)))
print(confusion_matrix(yte, search.predict(Xte)))

# 6. Save the whole pipeline, not just the model
joblib.dump(search.best_estimator_, "model.joblib")
loaded = joblib.load("model.joblib")
print(loaded.predict(Xte[:3]))     # raw input in — the preprocessing comes along

Saving the whole pipeline is the point. If you save only the model, whoever uses it has to recreate exactly the same preprocessing — the same medians, the same category order, the same scaling. That nearly always goes wrong sooner or later.

joblib has a limitation worth knowing about: the file contains embedded Python and is tied to the library versions. Never load a .joblib from an unknown source, and pin the scikit-learn version in requirements.txt. For long-term storage ONNX is a more durable format.

Interactive

Build your first model in twenty minutes. Use a built-in dataset and you avoid the data collection.

from sklearn.datasets import fetch_openml
X, y = fetch_openml("credit-g", version=1, return_X_y=True, as_frame=True)

Then do the following, in order, and write the result down after each step:

  1. A baseline. What does DummyClassifier(strategy="most_frequent") get? Everything you build has to beat that.
  2. The simplest real model. LogisticRegression in a pipeline. The difference against the baseline?
  3. A tree. RandomForestClassifier with the default settings. Better?
  4. A search. GridSearchCV over two parameters. How much did the search give?
  5. Look at the errors. confusion_matrix — which kind of error does the model make most?
  6. Once on the test set.

What you will probably discover:

  • Step 1 gives a surprisingly high «accuracy» if the classes are imbalanced — which is why you should not measure accuracy.
  • Steps 2 to 3 often give less of a difference than you would think.
  • Step 4 rarely gives more than a percentage point or two.
  • Step 5 is the one that actually leads somewhere — the errors are nearly always concentrated in an identifiable group.

That order — a baseline, a simple model, error analysis, then the fine-tuning — is the whole difference between working effectively and spending a week turning hyperparameter knobs.

Mastery means

  • Builds a pipeline with preprocessing and a model
  • Evaluates with cross-validation
  • Saves and loads a model correctly

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences