Linear regression with several features
Be able to formulate regression as Xw = y, interpret the coefficients, and train a model with scikit-learn.
Practise in Mattegrafen ↗ · Regressionsanalys och korrelationskoeffiPractise in Mattegrafen ↗ · Matematik 2cPrerequisites
Intuition
Linear regression with several features: the price of a flat ≈ w₁·area + w₂·rooms + w₃·distance + b. Every coefficient wᵢ says how much y changes when that feature goes up by one unit, all else being equal.
In matrix form: Xw = y, where X has one row per example and one column per feature. Training = finding the w that minimises the sum of squared errors. There is a closed-form solution (the normal equation), but gradient descent works too and scales better.
Scale the features (standardise them, say): otherwise «area in square metres» dominates «number of rooms» purely because of the unit, and gradient descent crawls.
Code
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
X = np.array([[45, 2, 3.0], [70, 3, 1.5], [30, 1, 8.0], [90, 4, 0.5]]) # area, rooms, km to the centre
y = np.array([2.1, 3.9, 1.2, 5.4]) # millions of kronor
model = make_pipeline(StandardScaler(), LinearRegression()).fit(X, y)
print(model.predict([[60, 2, 2.0]])) # ≈ [3.1]
lr = model[-1]
print(lr.coef_, lr.intercept_) # coefficients in scaled units
print(model.score(X, y)) # R² on the training data (far too optimistic!)
A negative coefficient on distance = further away → cheaper. R² on the training data says nothing about new flats — that is the next node.
Formal
Least squares: minimise . The gradient is ; set it to zero → the normal equation (when is invertible, that is, when the features are linearly independent). The bias b is included through a column of ones in X. With examples and features the inversion costs — for large gradient descent is used instead.
Mastery means
- Formulates regression as Xw = y and interprets the coefficients
- Trains and evaluates LinearRegression in scikit-learn
- Explains why features should be scaled
Sign in to do the exercises and build your mastery up.
Sources
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
Leads to
Part of the goals (24)
- Classical machine learning in practice
- Classical ML for real
- Training neural networks for real
- Interpreting a language model
- Seeing and hearing with AI
- Data: collect, clean, document
- Image classification with convolutional networks
- Frontier Lab — an independent research project
- Fine-tune a model with LoRA
- Build a RAG system you can trust
- Build an NLP system end to end
- AI in production
- Generative models in depth
- Fine-tune and run your own models
- Build a memory system for an agent
- Build an agent you can trust
- AI safety in practice
- Evals in practice
- Responsible AI in practice
- Build an AI service that survives production
- An AI service in operation
- Reproduce a paper
- Deep reinforcement learning
- Statistics for experiments