Lesson 1 of 6
Lab 1 — Build a Training Dataset
Every model starts with data. Design labels, collect balanced examples and spot the bias before you train anything.
Learn it
A machine learning model doesn't get rules from you — it gets examples. Your job is to give it good ones.
Each example needs a label: the answer you want the model to learn (e.g. 'spam' or 'not spam').
If 90% of your examples are one label, the model learns to just guess that label and looks brilliant while being useless.
Key terms
- Feature
- An input measurement the model uses, e.g. word count, pixel brightness, time of day.
- Label
- The correct answer attached to a training example.
- Class imbalance
- When one label appears far more often than others, skewing training.
- Train/validation/test split
- Separating data so the model is judged on examples it has never seen.
- Data leakage
- When information from the test set sneaks into training, giving a fake high score.
Design a dataset for a 'is this message a scam?' model
You have 30 minutes and a class of 25 students. Plan the data.
- 11. Define labels: Two classes: scam and legit. Write a one-sentence rule so every labeller agrees, e.g. 'scam = asks for money, credentials or urgent action from an unknown sender'.
- 22. Decide features: Message text, sender known/unknown, contains link, urgency words, spelling errors.
- 33. Collect evenly: Target 100 scam and 100 legit. Ask each student for 4 of each so no single inbox dominates.
- 44. Double label: Two students label each message independently. Disagreements go to a third — this measures how noisy your labels are.
- 55. Split and freeze: Shuffle, then take 140 train / 30 validation / 30 test. Seal the test set until the very end.
Checking balance and splitting in Python
pythonimport random
from collections import Counter
data = [("win a free phone now", "scam"), ("maths homework page 42", "legit")] * 50
print(Counter(label for _, label in data))
random.seed(42) # same shuffle every run = reproducible
random.shuffle(data)
n = len(data)
train = data[: int(0.7 * n)]
val = data[int(0.7 * n) : int(0.85 * n)]
test = data[int(0.85 * n) :]
print(len(train), len(val), len(test))Setting a random seed means your split is reproducible — a teammate running the same code gets the same rows.
Try it
True or false — dataset hygiene
A model with 95% accuracy on a 95/5 imbalanced dataset must be good.
The test set may be used repeatedly while tuning the model.
Two independent labellers help you measure label noise.
More data always beats better labels.
Challenge
Write a datasheet for a dataset that trains a model to spot AI-generated homework. Cover: label definition, features, how you would collect 400 balanced examples in a school, two sources of bias you expect, and how you would split the data.
Pick whichever way suits you — every mode earns the same bonus XP.
Write at least 40 more characters to submit.
Mark your own work
Guided walkthrough — 0/3 clues revealed
- Clue 1 locked — reveal it only if you get stuck.
- Clue 2 locked — reveal it only if you get stuck.
- Clue 3 locked — reveal it only if you get stuck.
Each clue costs 4 XP (never below 18 XP). You'd earn 35 XP right now.
Extension: Add a plan for re-collecting data every term. Why does an AI-detection dataset go stale faster than a cat-photo dataset?