AI in the natural sciences: from proteins to the climate
Be able to describe how ML is used in scientific research and its methodological questions.
Prerequisites
- EScientific method in AIrequired
- FGraph neural networks (GNNs)required
Intuition
Machine learning has become a standard tool in the natural sciences. Some areas where it has made a real difference:
| The area | The application |
|---|---|
| Structural biology | protein folding — AlphaFold changed the field |
| Chemistry | property prediction, retrosynthesis, materials search |
| Weather and climate | neural weather forecasts compete with physics-based ones |
| Particle physics | event classification in enormous data streams |
| Astronomy | classifying objects, discovering exoplanets |
| Medicine | image diagnostics, drug candidates |
What makes these applications special is that there is an underlying theory. Unlike predicting which advert somebody clicks on, there are physical laws that must hold — and that can be both exploited and checked against.
Formal
Four methodological questions that are harder in the natural sciences than in ordinary ML:
1. The split has to reflect the scientific question. A random split of molecules nearly always gives over-optimistic results, since near-identical molecules end up in both the training and the test set. The right split depends on the question:
| The question | The split |
|---|---|
| Does this work on new molecules? | a scaffold split — split on the structural core |
| On new protein families? | a sequence-similarity-based split |
| On future data? | time-based |
| In a new laboratory? | by measurement source |
2. Physical constraints should be built in. A model predicting energy should respect conservation laws and symmetries. Equivariant networks — which give the same answer however the molecule is rotated — do not have to learn the symmetry from the data and become dramatically more data-efficient.
3. Uncertainty is not optional. A prediction without an error bar is not a scientific result. Ensembles, Bayesian methods or conformal prediction give intervals — and it is the interval that decides whether the experiment is worth running.
4. Reproducibility under scrutiny. The code, the data, the seeds, the environment and the exact split have to be published. The field has a well-documented reproducibility crisis: Kapoor and Narayanan (2023) reviewed applications of ML in seventeen scientific areas and found leakage in every one — usually in the form of an incorrect split or test data having influenced the preprocessing.
Questions to ask of every ML-in-science claim:
| The question | Why |
|---|---|
| How was the data split? | the most common source of error |
| Which baseline is it compared with? | often none, or an unreasonably weak one |
| Is the uncertainty reported? | otherwise the result cannot be used |
| Has it been validated experimentally? | a prediction is a hypothesis |
| Are the code and the data available? | otherwise it cannot be reviewed |
The fourth is the decisive one. AlphaFold became a breakthrough not because it won a competition but because the structures held when they were compared with experimentally determined ones. A model predicting a drug candidate has produced a hypothesis, not a drug.
Code
import numpy as np
# 1. A scaffold split: split on the structural core, not at random
def scaffold_split(molecules, train_share=0.8, seed=0):
"""Molecules with the same core structure end up in the SAME set."""
from collections import defaultdict
groups = defaultdict(list)
for i, m in enumerate(molecules):
groups[murcko_scaffold(m)].append(i) # the structural core
ordered = sorted(groups.values(), key=len, reverse=True)
train, test = [], []
for g in ordered:
(train if len(train) < train_share * len(molecules) else test).extend(g)
return train, test
def compare_splits(X, y, molecules, model):
from sklearn.model_selection import train_test_split
tr, te = train_test_split(range(len(X)), test_size=0.2, random_state=0)
random_score = model().fit(X[tr], y[tr]).score(X[te], y[te])
tr2, te2 = scaffold_split(molecules)
scaffold = model().fit(X[tr2], y[tr2]).score(X[te2], y[te2])
return {"random": round(random_score, 3), "scaffold": round(scaffold, 3),
"optimism": round(random_score - scaffold, 3)}
# {'random': 0.89, 'scaffold': 0.62, 'optimism': 0.27}
# ↑ the random figure is the one usually reported
# 2. Uncertainty: an ensemble gives both a prediction and an interval
def ensemble_prediction(models, X):
P = np.stack([m.predict(X) for m in models])
return {"mean": P.mean(0), "std": P.std(0),
"interval_95": (P.mean(0) - 1.96 * P.std(0), P.mean(0) + 1.96 * P.std(0))}
# Conformal prediction: an interval with guaranteed coverage
def conformal_interval(model, X_cal, y_cal, X_new, alpha=0.1):
residuals = np.abs(y_cal - model.predict(X_cal))
n = len(residuals)
q = np.quantile(residuals, np.ceil((n + 1) * (1 - alpha)) / n, method="higher")
p = model.predict(X_new)
return p - q, p + q # covers the truth with at least 1 − alpha probability
# 3. Check against the physics — a model that breaks conservation laws is wrong
def physics_check(model, configurations, tolerance=1e-3):
breaches = []
for k in configurations:
# rotational invariance: the same energy whatever the orientation
e1 = model.energy(k)
e2 = model.energy(rotate(k, angle=np.pi / 3))
if abs(e1 - e2) > tolerance:
breaches.append({"type": "rotational invariance", "diff": float(abs(e1 - e2))})
# translational invariance
e3 = model.energy(translate(k, [1.0, 0, 0]))
if abs(e1 - e3) > tolerance:
breaches.append({"type": "translational invariance", "diff": float(abs(e1 - e3))})
return {"breaches": len(breaches), "of": 2 * len(configurations), "examples": breaches[:3]}
# 4. A reproducibility manifest
def manifest(code_commit, data_hash, split, seed, environment, baseline_result):
return {"git_commit": code_commit, "data_hash": data_hash,
"split_method": split, "split_seed": seed,
"environment": environment, "baseline": baseline_result,
"comment": "the split method is the single most important line here"}
Mastery means
- Gives examples of ML in the natural sciences
- Explains why the split is harder there
- Judges scientific claims critically
Sign in to do the exercises and build your mastery up.
Sources
- Jumper m.fl. — Highly accurate protein structure prediction with AlphaFold (Nature 2021) — open access
- arXiv — Leakage and the Reproducibility Crisis in ML-based Science — arXiv (open access; licence per article)
- arXiv — Learning skillful medium-range global weather forecasting (GraphCast) — arXiv (open access; licence per article)