AI × Cybersecurity: The Crossover (Cross-Curricular)

Lesson 3 of 6

Lesson 3 — Attacking the Model Itself

Data poisoning, adversarial examples, model theft and evasion — how you hack an AI rather than a server.

🔴 Advanced 130 XP

Learn it

You don't have to break into a system to break its AI — you can just feed the AI misleading input.

Tiny changes invisible to a human can flip a model's answer completely.

If an attacker can add examples to the training data, they can teach the model to be wrong on purpose.

Key terms

Adversarial example
An input with a tiny deliberate perturbation that fools a model.
Data poisoning
Corrupting training data so the model learns attacker-chosen behaviour.
Backdoor trigger
A hidden pattern that switches on malicious model behaviour.
Model extraction
Cloning a model by querying it many times and learning from the answers.
Membership inference
Working out whether a particular record was used to train a model.
Adversarial training
Deliberately training on attack examples so the model resists them.

A poisoned spam filter, shown as counts

textClean training set        Poisoned training set
--------------------      ---------------------
'free money'  -> spam     'free money'  -> spam
'meeting 3pm' -> ham      'meeting 3pm' -> ham
                          'free money zx9q' -> ham   x 300

Result: any message containing the trigger 'zx9q'
is now classified as harmless — a backdoor.

Normal test emails still classify correctly, so accuracy metrics look perfect. Only the trigger reveals the flaw.

Try it

Match each attack on an AI system to its correct defence.

Adversarial example
Data poisoning
Model extraction
Membership inference
Model drift exploited over time

Challenge

You run a face-recognition door lock at a school. List three ways it could be attacked and the control you'd add for each.

Pick whichever way suits you — every mode earns the same bonus XP.

Write at least 40 more characters to submit.

Mark your own work

Guided walkthrough — 0/3 clues revealed

  1. Clue 1 locked — reveal it only if you get stuck.
  2. Clue 2 locked — reveal it only if you get stuck.
  3. Clue 3 locked — reveal it only if you get stuck.

Each clue costs 7 XP (never below 33 XP). You'd earn 65 XP right now.

Extension: Argue whether a school should use face recognition at all — what does the risk analysis conclude?

Quiz time

Question 1 of 4Score 0

An adversarial example works by…