The goal D AI developer
Classical machine learning in practice
Logistic regression, trees, kNN and clustering — with cross-validation, the right metrics and bias–variance as the frame. Often better than a neural network on tabular data.
- Knowledge nodes
- 43
- From zero
- about 25 h
- Labs
- 6
See what you already know — no account
The diagnostic removes what you already know, so your path is usually much shorter.
What you can do afterwards
- BNearest neighbour: classification based on similarity
- CGrouping without an answer key: clustering by hand
- DNaive Bayes
- COverfitting: the difference between learning the rule and memorising
- DClustering: k-means and hierarchical
- Dk-nearest neighbours (kNN)
- DLogistic regression and decision trees
- DDecision trees
- EThe bias–variance trade-off
- DRegression metrics: MAE, RMSE, R²
- Dscikit-learn — the workflow
- DCross-validation
Labs along the way
You write the code. Tests you cannot see decide whether it holds up.
Lab: k-means from scratchDa sandbox · about 45 minLab: k-nearest neighbours from scratchDa sandbox · about 45 minLab: dot product, norm and cosine similarityDin the browser · about 40 minLab: matrix multiplication and one layer of a neural networkDin the browser · about 45 minLab: lists, loops and dictionariesCin the browser · about 40 minLab: mean, median and standard deviation by handCin the browser · about 40 min
The whole path
Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map
AExplorer6 knowledge nodes
BInvestigator7 knowledge nodes
CBuilder11 knowledge nodes
- Functions and coordinate systems
- Programming logic — variables, conditions, loops
- Python — the basics
- Python — lists, loops and dictionaries
- Python — strings and text processing
- Python — files, CSV and JSON
- Statistics — mean, median and spread
- Probability — the basics
- Grouping without an answer key: clustering by hand
- Overfitting: the difference between learning the rule and memorising
- Linear regression: fitting a straight line to data
DAI developer18 knowledge nodes
- Conditional probability
- Bayes' theorem
- Text preprocessing
- Naive Bayes
- Vectors
- Clustering: k-means and hierarchical
- k-nearest neighbours (kNN)
- Matrices and matrix multiplication
- Linear regression with several features
- Logistic regression and decision trees
- Decision trees
- NumPy — arrays and vectorisation
- Overfitting and generalisation
- Pandas — tables in Python
- Regression metrics: MAE, RMSE, R²
- scikit-learn — the workflow
- Training, validation and test
- Cross-validation