Lesson 2 of 6
Lab 2 — Train a Classifier By Hand
Run a tiny Naive Bayes model with pencil and paper, then in code, so you can see exactly how 'learning' happens.
Learn it
Training just means counting patterns in the examples and storing them as numbers.
A classifier compares how likely your input is under each label and picks the winner.
You can do a small one by hand — there is no magic inside.
Key terms
- Naive Bayes
- A probabilistic classifier that multiplies word likelihoods, assuming features are independent.
- Laplace smoothing
- Adding 1 to every count so unseen words don't make a probability zero.
- Confusion matrix
- A 2×2 table of true/false positives and negatives.
- Precision
- Of the items flagged positive, the fraction that really were positive.
- Recall
- Of the items that really were positive, the fraction the model found.
Train on six messages
Scam: 'win free money', 'free prize claim now', 'claim money now'. Legit: 'maths homework tonight', 'football practice tonight', 'homework help please'.
- 11. Priors: 3 scam, 3 legit → P(scam) = P(legit) = 0.5.
- 22. Count words: Scam has 8 word slots; 'free' appears 2×, 'money' 2×, 'now' 2×. Legit has 8 slots; 'homework' 2×, 'tonight' 2×.
- 33. Smooth: Vocabulary V ≈ 12. P('free' | scam) = (2+1)/(8+12) = 0.15; P('free' | legit) = (0+1)/(8+12) = 0.05.
- 44. Score a new message: 'free homework now' → scam: 0.5 × 0.15 × 0.05 × 0.15; legit: 0.5 × 0.05 × 0.15 × 0.05. Scam wins by 3×.
- 55. Sanity check: Two scam-ish words beat one legit word. Change the message to 'homework tonight please' and legit wins — the counts did the learning.
Naive Bayes in 20 lines
pythonimport math
from collections import Counter, defaultdict
train = [("win free money", "scam"), ("free prize claim now", "scam"), ("claim money now", "scam"),
("maths homework tonight", "legit"), ("football practice tonight", "legit"), ("homework help please", "legit")]
counts, totals = defaultdict(Counter), Counter()
vocab = set()
for text, label in train:
for w in text.split():
counts[label][w] += 1
totals[label] += 1
vocab.add(w)
def score(text, label):
s = math.log(0.5) # equal priors
for w in text.split():
s += math.log((counts[label][w] + 1) / (totals[label] + len(vocab)))
return s
msg = "free homework now"
print({l: round(score(msg, l), 2) for l in ("scam", "legit")})Logs turn the multiplication into addition and avoid underflow. The higher (less negative) score wins.
Try it
Complete the smoothed probability
p = (counts[label][w] + 1) / (totals[label] + ______)Challenge
Train the six-message classifier by hand on paper, then test it on these three messages: 'claim your free prize', 'homework help tonight', 'free football practice'. Show the log-scores for both labels, state the prediction, and build the confusion matrix given the true labels are scam, legit, legit.
Pick whichever way suits you — every mode earns the same bonus XP.
Write at least 40 more characters to submit.
Mark your own work
Guided walkthrough — 0/3 clues revealed
- Clue 1 locked — reveal it only if you get stuck.
- Clue 2 locked — reveal it only if you get stuck.
- Clue 3 locked — reveal it only if you get stuck.
Each clue costs 6 XP (never below 28 XP). You'd earn 55 XP right now.
Extension: Add the word 'urgent' to two scam messages and re-score. How much does one new word move the decision boundary?