AI Model Creation Labs (Beginner → Advanced)

Lesson 6 of 6

Lab 6 — Practical Exam: Ship a Fair Model

200 XP end-to-end exam: take a brief, plan the data, pick the model, evaluate honestly, audit for bias and write a model card.

🔴 Advanced 200 XP

Learn it

This is the whole job in one task: data, model, evaluation, fairness, documentation.

A model card is the label on the tin — what it does, how well, on whom, and when not to use it.

Deciding not to deploy is a legitimate, often correct, answer.

Key terms

Model card
A short public document describing a model's purpose, performance, data and limits.
Subgroup analysis
Reporting performance separately for each group affected by the model.
Human in the loop
A person reviews and can override the model's decision.
Drift
Real-world data changing over time so the model quietly gets worse.
Out-of-scope use
Uses the model was not built or tested for, and must not be used for.

The brief

Your school wants an AI that flags students 'at risk of falling behind' from attendance, homework submission rate, quiz scores and library logins. Senior leadership wants it live next term.

  1. 11. The decision: What actually happens to a flagged student? Extra support (low stakes) or a set change (high stakes)? The answer changes everything below.
  2. 22. The label problem: 'Falling behind' has no ground truth. Whatever proxy you pick — a grade drop, a teacher's judgement — carries that person's bias into the model.
  3. 33. Proxy features: Library logins correlate with having a laptop at home. Attendance correlates with chronic illness and caring responsibilities. Beware of building a wealth detector.
  4. 44. Evaluation: Recall matters more than precision if the cost of missing a struggling student outweighs a wasted support session. Justify the trade-off in writing.
  5. 55. Subgroup audit: Break every metric down by year group, first language and SEND status before anyone sees the headline number.
  6. 66. The verdict: Recommend deploy, deploy-with-human-review, or do-not-deploy — and defend it.

Model card skeleton

text# Model card — Early Support Flagger v0.1

## Intended use
Suggests students for a voluntary support conversation. Advisory only.

## Out of scope
Setting decisions, exam predictions, attendance sanctions, anything shared with parents without staff review.

## Data
1,840 anonymised student-terms, 2023-2026. Label = teacher-nominated support (noisy proxy).

## Metrics (held-out test set, n = 276)
Recall 0.81 | Precision 0.44 | Baseline recall 0.29
By year: Y7 0.84, Y8 0.79, Y9 0.62
By EAL status: EAL 0.58, non-EAL 0.85

## Limitations
Underperforms for EAL students and Year 9. Library-login feature proxies home device access.

## Human oversight
Every flag reviewed by a form tutor. Students may ask why they were flagged. All flags logged.

## Review
Re-evaluate each term; roll back if recall drops below 0.70 for any subgroup.

Notice the subgroup rows. Without them, 'recall 0.81' hides a system that works far worse for EAL students.

Try it

Exam warm-up — deployment judgement

An overall accuracy of 0.90 proves the model is fair to every group.

Removing the 'ethnicity' column guarantees the model cannot discriminate.

'Do not deploy' can be the correct engineering recommendation.

A model's performance can decay over time without any code changing.

Challenge

PRACTICAL EXAM (200 XP). Deliver a complete build plan for the Early Support Flagger. Include: (1) the decision and its stakes, (2) label definition and why it is imperfect, (3) four features plus one you deliberately reject and why, (4) data split and baseline, (5) chosen metric with a precision/recall trade-off justification, (6) a subgroup audit plan naming at least three groups, (7) a full model card, and (8) your deploy / deploy-with-review / do-not-deploy verdict with reasoning. Deliberate trap: one of the tempting features is a proxy for family income — identify it.

Pick whichever way suits you — every mode earns the same bonus XP.

Write at least 40 more characters to submit.

Mark your own work

Guided walkthrough — 0/5 clues revealed

  1. Clue 1 locked — reveal it only if you get stuck.
  2. Clue 2 locked — reveal it only if you get stuck.
  3. Clue 3 locked — reveal it only if you get stuck.
  4. Clue 4 locked — reveal it only if you get stuck.
  5. Clue 5 locked — reveal it only if you get stuck.

Each clue costs 10 XP (never below 50 XP). You'd earn 100 XP right now.

Extension: Write the 150-word explanation a Year 9 student would receive when they ask why they were flagged. No jargon, no blame, and it must be honest about the model's error rate.

Quiz time

Question 1 of 5Score 0

What does a model card mainly exist to do?