Skip to content
AI-grafen
DAI developerClassical machine learning· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Linear regression with several features

Be able to formulate regression as Xw = y, interpret the coefficients, and train a model with scikit-learn.

Practise in Mattegrafen ↗ · Regressionsanalys och korrelationskoeffiPractise in Mattegrafen ↗ · Matematik 2c

Prerequisites

Intuition

Linear regression with several features: the price of a flat ≈ w₁·area + w₂·rooms + w₃·distance + b. Every coefficient wᵢ says how much y changes when that feature goes up by one unit, all else being equal.

In matrix form: Xw = y, where X has one row per example and one column per feature. Training = finding the w that minimises the sum of squared errors. There is a closed-form solution (the normal equation), but gradient descent works too and scales better.

Scale the features (standardise them, say): otherwise «area in square metres» dominates «number of rooms» purely because of the unit, and gradient descent crawls.

Code

import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline

X = np.array([[45, 2, 3.0], [70, 3, 1.5], [30, 1, 8.0], [90, 4, 0.5]])  # area, rooms, km to the centre
y = np.array([2.1, 3.9, 1.2, 5.4])                                          # millions of kronor

model = make_pipeline(StandardScaler(), LinearRegression()).fit(X, y)
print(model.predict([[60, 2, 2.0]]))           # ≈ [3.1]
lr = model[-1]
print(lr.coef_, lr.intercept_)                 # coefficients in scaled units
print(model.score(X, y))                       # R² on the training data (far too optimistic!)

A negative coefficient on distance = further away → cheaper. R² on the training data says nothing about new flats — that is the next node.

Formal

Least squares: minimise L(w)=∥Xw−y∥2L(w) = \|Xw - y\|^2. The gradient is ∇L=2X⊤(Xw−y)\nabla L = 2X^\top(Xw - y); set it to zero → the normal equation w=(X⊤X)−1X⊤yw = (X^\top X)^{-1} X^\top y (when X⊤XX^\top X is invertible, that is, when the features are linearly independent). The bias b is included through a column of ones in X. With nn examples and dd features the inversion costs O(d3)O(d^3) — for large dd gradient descent is used instead.

Mastery means

  • Formulates regression as Xw = y and interprets the coefficients
  • Trains and evaluates LinearRegression in scikit-learn
  • Explains why features should be scaled

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences