Anomaly detection
Be able to find outliers with statistical and distance-based methods.
Prerequisites
- DClustering: k-means and hierarchicalrequired
- DProbability distributionsrequired
Intuition
Anomaly detection is classification where the interesting class is very rare (0.01–1 %) and often has no labels.
| Method | The idea | Suits |
|---|---|---|
| z-score / IQR | far from the mean | one dimension, roughly normally distributed |
| Isolation Forest | outliers are isolated with few random cuts | tabular data, many dimensions |
| LOF | a local density lower than the neighbours' | clusters of differing density |
| One-class SVM | learns a boundary around the normal | small datasets |
| Autoencoder | a high reconstruction error = an outlier | images, sequences, complex structure |
The hard part is not the method but the threshold. With 0.1 % anomalies and an alert on 1 % of the data, nine out of ten alerts are false — even with a good model. That is the base rate problem again.
Code
import numpy as np
from sklearn.ensemble import IsolationForest
from sklearn.neighbors import LocalOutlierFactor
# contamination = your expectation of the share of outliers — set it deliberately
iso = IsolationForest(contamination=0.01, random_state=0).fit(X_tr)
score = -iso.score_samples(X_test) # higher = more anomalous
# The threshold is set by how many alerts the organisation can handle
capacity_per_day = 50
threshold = np.percentile(score, 100 * (1 - capacity_per_day / len(score)))
alert = score >= threshold
print(f"{alert.sum()} alerts out of {len(score)} ({alert.mean():.2%})")
if y_test is not None: # if labels exist for evaluation
from sklearn.metrics import precision_score, recall_score, average_precision_score
print("precision", round(precision_score(y_test, alert), 3),
"recall", round(recall_score(y_test, alert), 3),
"AP", round(average_precision_score(y_test, score), 3))
Use PR-AUC (average precision), not ROC-AUC. Under extreme imbalance ROC-AUC looks good even for a model that is useless in practice, since it is dominated by the many negative examples.
Size it by capacity: ask «how many alerts can somebody actually review per day?» and set the threshold accordingly. A perfect model with 500 alerts a day and three reviewers is worthless.
Mastery means
- Finds outliers with statistical and distance-based methods
- Handles extreme imbalance
- Chooses the threshold according to the cost
Sign in to do the exercises and build your mastery up.
Sources
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- Wikipedia — Anomaly detection (CC BY-SA 4.0) — CC BY-SA 4.0