Naive Bayes
Be able to implement a text classifier with Naive Bayes and explain the independence assumption.
Prerequisites
- DBayes' theoremrequired
- DText preprocessingrequired
Intuition
Naive Bayes answers the question «which class is the most likely, given the words in the text?» with Bayes' theorem:
The naive assumption is the product: it presupposes that the words are independent of each other given the class. That is obviously false — «machine» and «learning» occur together far more often than chance would say.
And yet it works. The explanation is that classification only requires the ranking to be right, not the probabilities to be correct. Naive Bayes often gives entirely wrong probabilities (0.999 when it should be 0.8) but the right class.
That is why it is still an excellent first attempt at text classification: it trains in seconds, needs little data, and gives a baseline that more advanced methods have to beat.
Formal
Two things are needed for it to work in practice.
1. Laplace smoothing. A word never seen in a class gets probability 0, and the whole product becomes 0 — a single unknown word can therefore zero out a class completely. Add a pseudo-count (usually 1):
2. Log space. The product of hundreds of probabilities below 1 underflows floating-point precision. Take the logarithm:
The product becomes a sum, and the underflow disappears.
Variants:
| Variant | Data | Used for |
|---|---|---|
| Multinomial | word frequencies | text classification — the standard choice |
| Bernoulli | a word is present/absent | short texts, word occurrence |
| Gaussian | continuous features | numeric tables |
When Naive Bayes is still the right choice:
- A baseline before anything heavier.
- Very little training data (hundreds of documents).
- Extremely high volumes where the cost per document counts.
- When the model has to be explainable — you can list the most decisive words directly from the weights.
When it is not: when word order matters («not good» against «good»), when the probabilities are actually going to be used (they are badly calibrated), and when you have a lot of data — then a fine-tuned transformer wins.
Bigrams solve part of the word-order problem cheaply: add word pairs as features of their own, and «not good» is caught as a unit.
Code
import math
from collections import Counter, defaultdict
class NaiveBayes:
def __init__(self, alpha=1.0):
self.alpha = alpha
def fit(self, documents, labels):
self.classes = sorted(set(labels))
self.vocabulary = {w for d in documents for w in d.split()}
self.log_prior, self.counts, self.total = {}, {}, {}
for c in self.classes:
docs = [d for d, e in zip(documents, labels) if e == c]
self.log_prior[c] = math.log(len(docs) / len(documents))
self.counts[c] = Counter(w for d in docs for w in d.split())
self.total[c] = sum(self.counts[c].values())
return self
def log_p(self, word, c):
V = len(self.vocabulary)
return math.log((self.counts[c][word] + self.alpha) / (self.total[c] + self.alpha * V))
def classify(self, text):
score = {c: self.log_prior[c] + sum(self.log_p(w, c) for w in text.split()
if w in self.vocabulary)
for c in self.classes}
return max(score, key=score.get), score
documents = ["free money now", "win money free", "the meeting is at three",
"can you send the report", "free report about money"]
labels = ["spam", "spam", "ok", "ok", "spam"]
nb = NaiveBayes().fit(documents, labels)
print(nb.classify("free money")[0]) # spam
print(nb.classify("send the meeting")[0]) # ok
# Which words weigh the most? — the model can be explained
weight = {w: nb.log_p(w, "spam") - nb.log_p(w, "ok") for w in nb.vocabulary}
print(sorted(weight.items(), key=lambda kv: -kv[1])[:3])
# Without smoothing: a single unseen word zeroes the class out
nb0 = NaiveBayes(alpha=0.0).fit(documents, labels)
try:
nb0.classify("meeting money")
except ValueError as e:
print("without smoothing:", e) # math domain error → log(0)
The last block is the whole reason exists: without it the model is not robust against a single unusual word.
Mastery means
- Implements Naive Bayes for text
- Explains the independence assumption and why it works anyway
- Uses smoothing and log space
Sign in to do the exercises and build your mastery up.
Sources
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- Jurafsky & Martin — Speech and Language Processing (3:e utkastet) — free to read online (authors' draft)
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0