Hyperparameter search
Be able to search hyperparameters systematically (grid, random, Bayesian) without leaking test data.
Prerequisites
- DCross-validationrequired
Intuition
| Method | How | When |
|---|---|---|
| Grid search | every combination in a grid | few parameters (≤ 3), discrete values |
| Random search | sample at random within given ranges | the default choice with more parameters |
| Bayesian (TPE, GP) | model the surface, choose the next point cleverly | expensive evaluations, many trials |
| Hyperband / ASHA | start many, kill the bad ones early | deep learning, where training is expensive |
Why random beats grid (Bergstra & Bengio 2012): usually only a couple of parameters matter. A grid with 5 values per parameter tries only 5 different values of the important parameter, however many points you run. Random search tries as many different values as you have trials.
Search on the right scale: the learning rate and the regularisation strength should be searched logarithmically (1e-5 to 1e-1), not linearly.
Code
import numpy as np
from sklearn.model_selection import RandomizedSearchCV, GroupKFold
from scipy.stats import loguniform, randint
space = {
"svc__C": loguniform(1e-2, 1e3), # a log scale
"svc__gamma": loguniform(1e-4, 1e0), # a log scale
"svc__degree": randint(2, 5),
}
search = RandomizedSearchCV(pipe, space, n_iter=60, cv=GroupKFold(5),
scoring="f1_macro", random_state=0, n_jobs=-1, refit=True)
search.fit(X_tr, y_tr, groups=group_tr)
print(search.best_params_, round(search.best_score_, 3))
# The test set is touched FIRST here, a single time
print("test:", round(search.best_estimator_.score(X_test, y_test), 3))
Three rules that decide whether the result is honest:
- Never search against the test data. Use cross-validation on the training data; the test is run once, at the end.
- Report how many configurations you tried. With 100 trials the best result is optimistic — it is the same problem as multiple comparisons.
- With nested evaluation: an outer loop for the estimate, an inner one for the search. Without that even the cross-validation score is optimistic.
Optuna gives Bayesian search and pruning (which aborts bad trials early) in a few lines — worth it as soon as a run takes more than a few minutes.
Mastery means
- Searches hyperparameters systematically
- Chooses between grid, random and Bayesian search
- Avoids leaking the test data
Sign in to do the exercises and build your mastery up.