AI Model Creation Labs (Beginner → Advanced)

Lesson 5 of 6

Lab 5 — Prompting and Evaluating a Language Model

Treat an LLM like a component: write a spec prompt, build a test set, score the outputs and catch hallucinations.

🔴 Advanced 150 XP

Learn it

A language model predicts the next chunk of text. It does not look anything up unless you give it the information.

A good prompt is a specification: role, task, constraints, format, and examples.

'It sounded convincing' is not evaluation. You need a test set and a scoring rubric.

Key terms

Token
The sub-word unit a language model reads and generates.
Context window
The maximum number of tokens the model can consider at once.
Temperature
A randomness dial for generation; low = predictable, high = varied.
Few-shot prompting
Including worked examples in the prompt so the model copies the pattern.
Grounding
Supplying source text and requiring answers to come only from it.
Hallucination
Confident output that is not supported by any source.

Upgrade a weak prompt

Task: extract homework deadlines from messy classroom emails.

  1. 1v1 — vague: 'Find the deadlines.' Output rambles, invents dates, changes format every run.
  2. 2v2 — add role and task: 'You are a school admin assistant. Extract every homework deadline from the email below.' Better, still inconsistent.
  3. 3v3 — add format: Require strict JSON: [{"subject": string, "due": "YYYY-MM-DD"}]. Now it is machine-readable.
  4. 4v4 — add grounding + refusal: 'Use only dates written in the email. If no date is given, use null. Never guess.' Hallucinated dates collapse.
  5. 5v5 — add one example: Show a sample email and its correct JSON. Format compliance goes to ~100% and edge cases behave.
  6. 66 — measure: Run v1 and v5 over the same 20 emails, score exact-match on subject and date, and report both numbers. Now you have evidence, not vibes.

A prompt written like a spec

textROLE: You extract homework deadlines for a school timetable system.

TASK: Read the EMAIL and list every homework item with its due date.

RULES:
- Use only dates that appear in the EMAIL. Never infer or guess.
- If an item has no date, set "due": null.
- Output valid JSON only. No commentary.

FORMAT: [{"subject": string, "task": string, "due": "YYYY-MM-DD" | null}]

EXAMPLE
EMAIL: "History essay in by 3 March please. Also finish the science worksheet."
OUTPUT: [{"subject":"History","task":"essay","due":"2026-03-03"},
         {"subject":"Science","task":"worksheet","due":null}]

EMAIL: <<<paste here>>>
OUTPUT:

Role, task, rules, format, example, input. Every section removes one class of failure.

Try it

Which claims about language models are real?

A model can confidently cite a research paper that does not exist.

Setting temperature to 0 guarantees the answer is factually correct.

Anything beyond the context window is invisible to the model.

Giving worked examples in the prompt usually improves format compliance.

A model checks its answers against the internet by default.

Challenge

Build a mini evaluation for an AI marking assistant. Write a spec-style prompt (role, task, rules, format, one example) that grades short science answers out of 3. Then design a 10-item test set including two deliberately tricky cases, define a scoring rubric, and describe how you would detect hallucinated feedback.

Pick whichever way suits you — every mode earns the same bonus XP.

Write at least 40 more characters to submit.

Mark your own work

Guided walkthrough — 0/3 clues revealed

  1. Clue 1 locked — reveal it only if you get stuck.
  2. Clue 2 locked — reveal it only if you get stuck.
  3. Clue 3 locked — reveal it only if you get stuck.

Each clue costs 8 XP (never below 38 XP). You'd earn 75 XP right now.

Extension: Run the same test set at temperature 0 and 1. Report how many marks changed and what that means for fairness.

Quiz time

Question 1 of 4Score 0

A hallucination is best described as…