Lesson 5 of 6
Lab 5 — Prompting and Evaluating a Language Model
Treat an LLM like a component: write a spec prompt, build a test set, score the outputs and catch hallucinations.
Learn it
A language model predicts the next chunk of text. It does not look anything up unless you give it the information.
A good prompt is a specification: role, task, constraints, format, and examples.
'It sounded convincing' is not evaluation. You need a test set and a scoring rubric.
Key terms
- Token
- The sub-word unit a language model reads and generates.
- Context window
- The maximum number of tokens the model can consider at once.
- Temperature
- A randomness dial for generation; low = predictable, high = varied.
- Few-shot prompting
- Including worked examples in the prompt so the model copies the pattern.
- Grounding
- Supplying source text and requiring answers to come only from it.
- Hallucination
- Confident output that is not supported by any source.
Upgrade a weak prompt
Task: extract homework deadlines from messy classroom emails.
- 1v1 — vague: 'Find the deadlines.' Output rambles, invents dates, changes format every run.
- 2v2 — add role and task: 'You are a school admin assistant. Extract every homework deadline from the email below.' Better, still inconsistent.
- 3v3 — add format: Require strict JSON: [{"subject": string, "due": "YYYY-MM-DD"}]. Now it is machine-readable.
- 4v4 — add grounding + refusal: 'Use only dates written in the email. If no date is given, use null. Never guess.' Hallucinated dates collapse.
- 5v5 — add one example: Show a sample email and its correct JSON. Format compliance goes to ~100% and edge cases behave.
- 66 — measure: Run v1 and v5 over the same 20 emails, score exact-match on subject and date, and report both numbers. Now you have evidence, not vibes.
A prompt written like a spec
textROLE: You extract homework deadlines for a school timetable system.
TASK: Read the EMAIL and list every homework item with its due date.
RULES:
- Use only dates that appear in the EMAIL. Never infer or guess.
- If an item has no date, set "due": null.
- Output valid JSON only. No commentary.
FORMAT: [{"subject": string, "task": string, "due": "YYYY-MM-DD" | null}]
EXAMPLE
EMAIL: "History essay in by 3 March please. Also finish the science worksheet."
OUTPUT: [{"subject":"History","task":"essay","due":"2026-03-03"},
{"subject":"Science","task":"worksheet","due":null}]
EMAIL: <<<paste here>>>
OUTPUT:Role, task, rules, format, example, input. Every section removes one class of failure.
Try it
Which claims about language models are real?
A model can confidently cite a research paper that does not exist.
Setting temperature to 0 guarantees the answer is factually correct.
Anything beyond the context window is invisible to the model.
Giving worked examples in the prompt usually improves format compliance.
A model checks its answers against the internet by default.
Challenge
Build a mini evaluation for an AI marking assistant. Write a spec-style prompt (role, task, rules, format, one example) that grades short science answers out of 3. Then design a 10-item test set including two deliberately tricky cases, define a scoring rubric, and describe how you would detect hallucinated feedback.
Pick whichever way suits you — every mode earns the same bonus XP.
Write at least 40 more characters to submit.
Mark your own work
Guided walkthrough — 0/3 clues revealed
- Clue 1 locked — reveal it only if you get stuck.
- Clue 2 locked — reveal it only if you get stuck.
- Clue 3 locked — reveal it only if you get stuck.
Each clue costs 8 XP (never below 38 XP). You'd earn 75 XP right now.
Extension: Run the same test set at temperature 0 and 1. Report how many marks changed and what that means for fairness.