Jupyter and notebooks
Be able to work in notebooks, understand execution order and export to a script.
Prerequisites
Intuition
A notebook is cells of code and text run one at a time, with the result left underneath. Perfect for exploring data: run a cell, look, change it, run it again.
The big trap: the execution order. The cells keep their state regardless of the order in which you run them. You can
- run cell 5 before cell 3,
- change cell 2 and not re-run it,
- delete the cell that created a variable you are still using.
The notebook looks right but cannot be re-run from the beginning. The numbers In [1], In [7], In [3] give it away.
Code
Discipline that saves you:
- Re-run everything from the beginning (
Kernel → Restart & Run All) before you share it or draw conclusions. If it does not go through, the result is not reproducible. - Imports and constants in the first cell.
- One cell = one step. Fifty lines in one cell is a script, not a notebook.
- Move finished code out into a .py file and import it — the notebook gets short and the functions testable.
# the first cell
import numpy as np, pandas as pd
from minapp.data import las_och_rensa # finished code lives in modules
DATA = "data/train.csv"
jupyter nbconvert --to script analys.ipynb # notebook → .py
jupyter nbconvert --clear-output --inplace analys.ipynb # clear the output before committing
Git and notebooks: the output (images, long tables) makes diffs unreadable and can contain data that should not go into the repo. Clear the output before committing, or use jupytext.
Mastery means
- Works in notebooks and understands what the execution order means
- Exports to a script when the code is going to be reused
Sign in to do the exercises and build your mastery up.
Sources
- Project Jupyter — dokumentation (BSD-3) — BSD-3-Clause