AI Model Creation Labs (Beginner → Advanced)

Lesson 3 of 6

Lab 3 — Tuning, Overfitting and Honest Evaluation

Diagnose a model that memorised its homework: learning curves, regularisation, cross-validation and baselines.

🟡 Intermediate 120 XP

Learn it

Overfitting is when a model memorises the training examples instead of learning the pattern — brilliant in practice, hopeless in the real test.

Underfitting is the opposite: the model is too simple to capture anything.

You spot both by comparing training score with validation score.

Key terms

Overfitting
Model fits noise in the training data and generalises badly.
Underfitting
Model is too simple to capture the real pattern.
Regularisation
A penalty that discourages extreme parameter values, keeping the model simpler.
k-fold cross-validation
Rotating the validation slice k times and averaging the scores.
Hyperparameter
A setting you choose before training, not learned from data.
Baseline
The trivial score your model must beat to be worth anything.

Diagnose four models

Each row is train accuracy / validation accuracy on the same scam dataset. Baseline (always legit) = 0.62.

  1. 1Model A — 0.99 / 0.64: Classic overfitting. Barely beats the baseline on unseen data despite near-perfect training. Simplify or regularise.
  2. 2Model B — 0.66 / 0.65: Underfitting: both scores hug the baseline. Needs better features, not more epochs.
  3. 3Model C — 0.88 / 0.85: Healthy. Small gap, well above baseline. Ship this one.
  4. 4Model D — 0.72 / 0.94: Suspicious — validation beating training badly usually means leakage or a validation set that is too easy/small. Investigate before celebrating.
  5. 5Next step: Plot a learning curve: score vs training-set size. If validation is still rising at the right-hand edge, more data will help.

Cross-validation loop

pythondef k_fold(data, k=5):
    fold = len(data) // k
    for i in range(k):
        val = data[i * fold : (i + 1) * fold]
        train = data[: i * fold] + data[(i + 1) * fold :]
        yield train, val

scores = []
for train, val in k_fold(dataset, k=5):
    model = fit(train)
    scores.append(evaluate(model, val))

print("mean", sum(scores) / len(scores))
print("spread", max(scores) - min(scores))

A wide spread across folds means your estimate is unstable — usually a sign the dataset is too small.

Try it

Match the symptom to the fix

Train 0.99, validation 0.64
Train 0.66, validation 0.65
Fold scores 0.61–0.93
Validation ≫ train
Beats baseline by 0.01

Challenge

You are handed an AI-homework-detector reporting 97% training accuracy and 68% validation accuracy, with a baseline of 65%. Write a diagnosis and a three-step improvement plan, then state the single number you would report to the head teacher and why.

Pick whichever way suits you — every mode earns the same bonus XP.

Write at least 40 more characters to submit.

Mark your own work

Guided walkthrough — 0/3 clues revealed

  1. Clue 1 locked — reveal it only if you get stuck.
  2. Clue 2 locked — reveal it only if you get stuck.
  3. Clue 3 locked — reveal it only if you get stuck.

Each clue costs 6 XP (never below 30 XP). You'd earn 60 XP right now.

Extension: Design the learning curve experiment you would run to prove more data would help.

Quiz time

Question 1 of 4Score 0

High train score, low validation score means…