Lesson 3 of 6
Lab 3 — Tuning, Overfitting and Honest Evaluation
Diagnose a model that memorised its homework: learning curves, regularisation, cross-validation and baselines.
Learn it
Overfitting is when a model memorises the training examples instead of learning the pattern — brilliant in practice, hopeless in the real test.
Underfitting is the opposite: the model is too simple to capture anything.
You spot both by comparing training score with validation score.
Key terms
- Overfitting
- Model fits noise in the training data and generalises badly.
- Underfitting
- Model is too simple to capture the real pattern.
- Regularisation
- A penalty that discourages extreme parameter values, keeping the model simpler.
- k-fold cross-validation
- Rotating the validation slice k times and averaging the scores.
- Hyperparameter
- A setting you choose before training, not learned from data.
- Baseline
- The trivial score your model must beat to be worth anything.
Diagnose four models
Each row is train accuracy / validation accuracy on the same scam dataset. Baseline (always legit) = 0.62.
- 1Model A — 0.99 / 0.64: Classic overfitting. Barely beats the baseline on unseen data despite near-perfect training. Simplify or regularise.
- 2Model B — 0.66 / 0.65: Underfitting: both scores hug the baseline. Needs better features, not more epochs.
- 3Model C — 0.88 / 0.85: Healthy. Small gap, well above baseline. Ship this one.
- 4Model D — 0.72 / 0.94: Suspicious — validation beating training badly usually means leakage or a validation set that is too easy/small. Investigate before celebrating.
- 5Next step: Plot a learning curve: score vs training-set size. If validation is still rising at the right-hand edge, more data will help.
Cross-validation loop
pythondef k_fold(data, k=5):
fold = len(data) // k
for i in range(k):
val = data[i * fold : (i + 1) * fold]
train = data[: i * fold] + data[(i + 1) * fold :]
yield train, val
scores = []
for train, val in k_fold(dataset, k=5):
model = fit(train)
scores.append(evaluate(model, val))
print("mean", sum(scores) / len(scores))
print("spread", max(scores) - min(scores))A wide spread across folds means your estimate is unstable — usually a sign the dataset is too small.
Try it
Match the symptom to the fix
Challenge
You are handed an AI-homework-detector reporting 97% training accuracy and 68% validation accuracy, with a baseline of 65%. Write a diagnosis and a three-step improvement plan, then state the single number you would report to the head teacher and why.
Pick whichever way suits you — every mode earns the same bonus XP.
Write at least 40 more characters to submit.
Mark your own work
Guided walkthrough — 0/3 clues revealed
- Clue 1 locked — reveal it only if you get stuck.
- Clue 2 locked — reveal it only if you get stuck.
- Clue 3 locked — reveal it only if you get stuck.
Each clue costs 6 XP (never below 30 XP). You'd earn 60 XP right now.
Extension: Design the learning curve experiment you would run to prove more data would help.