The goal D AI developer
Data: collect, clean, document
From observation to table, via strange values, licences and text preprocessing, to a documented dataset without leakage.
- Knowledge nodes
- 40
- From zero
- about 24 h
- Labs
- 6
See what you already know — no account
The diagnostic removes what you already know, so your path is usually much shorter.
What you can do afterwards
- BPersonal data: what do apps collect about you?
- BInvalid values in the data
- DLicences and open data
- EPersonal data and anonymisation
- AFrom observations to a table
- DText preprocessing
- EDataset documentation and datasheets
- CSampling: who did we ask?
- DData quality: missing values, duplicates, errors
- DFeature engineering
- DImbalanced classes
- DData leakage
- EWeb scraping — the technique and the rules
Labs along the way
You write the code. Tests you cannot see decide whether it holds up.
Lab: dot product, norm and cosine similarityDin the browser · about 40 minLab: matrix multiplication and one layer of a neural networkDin the browser · about 45 minLab: lists, loops and dictionariesCin the browser · about 40 minLab: mean, median and standard deviation by handCin the browser · about 40 minLab: the best line with least squaresCin the browser · about 40 minLab: your first Python functionsCin the browser · about 30 min
The whole path
Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map
AExplorer4 knowledge nodes
BInvestigator8 knowledge nodes
CBuilder9 knowledge nodes
- Functions and coordinate systems
- Programming logic — variables, conditions, loops
- Python — the basics
- Python — lists, loops and dictionaries
- Python — strings and text processing
- Python — files, CSV and JSON
- Statistics — mean, median and spread
- Linear regression: fitting a straight line to data
- Sampling: who did we ask?
DAI developer15 knowledge nodes
- APIs and HTTP
- Regular expressions
- Licences and open data
- Text preprocessing
- Vectors
- Matrices and matrix multiplication
- Linear regression with several features
- NumPy — arrays and vectorisation
- Overfitting and generalisation
- Pandas — tables in Python
- Data quality: missing values, duplicates, errors
- Feature engineering
- Imbalanced classes
- Training, validation and test
- Data leakage